top of page

Ivo Sage Open-Source Legal AI Challenges the Closed Legal Model Playbook

2 hours ago
12 min read

Ivo has released the first open model from a legal AI company post-trained specifically for long-horizon contract work. The Ivo Sage open-source legal AI release challenges a market built largely around proprietary models, controlled platforms, and closely guarded evaluation methods.

The company introduced Sage on October 1, 2026, during its first Ivo Inscribe customer conference in San Francisco. Developers can download its MIT-licensed adapter weights, inspect the published evaluations, and adapt the model for their own contract workflows.

That openness creates the central tension. Ivo is not claiming Sage beats every frontier model, and its strict task-completion numbers show considerable room for improvement. Instead, it argues that specialized training, organizational context, and transparent testing matter more than simply selecting the largest closed model.

The move puts pressure on legal AI vendors that sell access to proprietary systems without exposing their underlying models. It also gives technical legal teams a testable alternative, although running Sage requires far more infrastructure than downloading an ordinary desktop application.

What the Ivo Sage Open-Source Legal AI Release Includes

Ivo has published a specialized model, its training description, and unusually detailed evaluation records under an MIT license.

The Ivo Sage model is designed for long-horizon agentic legal work. Long-horizon work involves completing a multi-step assignment through repeated document reading, analysis, tool use, drafting, and revision.

That description matters because Sage is not a legal question-answering chatbot. It reads source documents and produces deliverables such as contract redlines, draft agreements, and legal memoranda. A single assignment can involve dozens of tool calls before the model produces its final work.

Ivo did not train a foundation model from the beginning. Sage consists of LoRA adapter weights applied to DeepSeek V4 Flash. LoRA, or low-rank adaptation, changes selected model behavior by training a relatively small set of additional weights.

DeepSeek V4 Flash is a mixture-of-experts model with 284 billion total parameters and 13 billion active parameters. A mixture-of-experts system activates only selected parts of the network for each token, reducing computation relative to using every parameter.

The Sage adapter itself occupies 12.2 GiB. Ivo trained the feed-forward and output projections while leaving the attention weights frozen. The model supports a 262,144-token context window during the reported training and evaluation process.

Ivo says it used reinforcement learning across 1,040 training assignments from the Legal Agent Benchmark. Reinforcement learning adjusts model behavior by rewarding outputs that score well against defined criteria.

The benchmark environment gave Sage six tools for reading, writing, editing, searching, and navigating files. An episode ended when the model stopped using tools or reached specified turn and generation limits.

This setup tries to measure completed legal work rather than isolated answers. A contract assignment can require locating clauses, comparing documents, applying instructions, revising language, and saving a usable deliverable.

Ivo also published task lists, evaluation transcripts, deliverables, judge decisions, and training records. That disclosure lets researchers inspect more than a final leaderboard number, although reproducing the complete experiment still requires substantial computing resources.

The release uses the MIT license, which generally permits commercial use, modification, and redistribution with limited conditions. However, Sage remains dependent on the separate terms governing its DeepSeek base model and related components.

The distinction between open weights and a complete open system is important. Ivo published the trained adapter, evaluation materials, and usage instructions. It did not release a small, self-contained application that any lawyer can immediately run.

The company recommends serving Sage through SGLang on multiple graphics processors. Its sample configuration uses eight-way tensor parallelism, which divides model execution across eight accelerators.

That requirement limits immediate adoption by smaller legal teams. Researchers, legal technology vendors, universities, and enterprises with existing AI infrastructure are better positioned to experiment first.

Still, the release changes what those groups can examine. They can test the same model against internal playbooks, build alternative agent loops, measure failure patterns, and propose improvements without waiting for Ivo.

This creates a public foundation for contract-specific experimentation. It also gives competitors and independent researchers a concrete artifact against which they can test Ivo’s claims.

Why Specialized Contract Training Matters Now

Sage tests whether focused post-training can turn a capable general model into a more efficient contract agent without hiding the evidence.

The central training resource was the Legal Agent Benchmark, developed to measure multi-step legal work. Its assignments cover general legal practice and contract-related tasks rather than simple multiple-choice questions.

Ivo divided the available assignments into training and validation groups. Related contracts and source documents stayed on the same side of the split, reducing the risk that nearly identical materials appeared in both groups.

The reported validation set contained 180 held-out assignments. Of those, 125 covered general legal work and 55 focused on contracts. Ivo ran one episode per task and used three judge passes for each episode.

Sage raised the overall criterion pass rate from 82.4 percent for the base model to 91.5 percent. A criterion pass rate measures the share of individual requirements satisfied across assignments.

The stricter result tells a less flattering story. Sage completed every requirement on only 15.6 percent of all validation tasks, compared with 7.2 percent for the base model.

Its improvement was stronger on the 55 contract assignments. The contract criterion pass rate rose from 70.1 percent to 91.3 percent after post-training.

However, its all-criteria contract pass rate reached only 16.4 percent. That figure means more than four out of five assignments still missed at least one required element under Ivo’s evaluation protocol.

This difference is essential for interpreting the release. A 91.3 percent criterion rate does not mean Sage completed 91.3 percent of contract assignments perfectly.

Legal deliverables can fail because of one missing escalation, an incorrect clause, or a neglected instruction. A high average can therefore coexist with a low rate of fully acceptable work.

Sage also used fewer steps after training. The mean episode length fell from 31.2 turns to 25 turns, while sampled token usage declined by 21 percent.

Those results support Ivo’s mechanism argument. Specialized reinforcement learning apparently improved contract behavior while reducing unnecessary model activity in the tested environment.

Ivo also evaluated Sage on RedlineBench, a separate contract-redlining benchmark excluded from training. The reported score rose from 30.8 for DeepSeek V4 Flash to 45.7 for Sage.

Performance on LegalBench remained nearly unchanged. Sage recorded an average score of 83.0, compared with 83.1 for its base model across 161 tasks.

That stability matters because specialization can damage broader capabilities. Ivo’s results suggest its post-training improved long contract workflows without materially reducing short-form legal reasoning on that benchmark.

The evidence is still company-produced. Ivo selected the configuration, ran the competing models, and published the report. Independent replication remains necessary before buyers treat the rankings as settled.

The benchmark also relies on model judges. An LLM judge scores each deliverable against a rubric, which scales evaluation but introduces another model into the measurement process.

Ivo attempted to reduce that uncertainty through three judge passes and public evaluation records. Those choices improve inspectability, but they do not replace review by qualified lawyers across real matters.

The training design nevertheless targets a recognizable weakness in general models. Contract work requires sustained attention to documents, instructions, exceptions, and dependencies across many steps.

A general chatbot can produce persuasive language while overlooking a provision elsewhere in the agreement. Sage is trained to operate inside a structured workflow where written deliverables, not fluent answers, determine the reward.

That makes the release relevant beyond legal AI. It supports a broader shift from general model selection toward domain-specific post-training, tool design, evaluation, and organizational context.

Open Weights Put Pressure on Proprietary Legal AI

Ivo is betting that model access will become less valuable than the context, workflow, and institutional knowledge surrounding the model.

Legal AI leaders have generally built controlled products around proprietary models, private integrations, and vendor-managed infrastructure. That approach can simplify deployment, support, security reviews, and accountability.

It also prevents customers from examining the underlying model directly. Buyers typically evaluate the complete service through demonstrations, pilot projects, contractual commitments, and vendor-supplied benchmarks.

Sage introduces a different route. A sophisticated legal department can inspect the weights, reproduce parts of the evaluation, and fine-tune the adapter using its own approved data.

That does not eliminate the need for a commercial platform. It changes the location of defensible value.

Min-Kyu Jung, Ivo’s co-founder and chief executive, argues that the next improvement will come from work context rather than larger models. His examples include documents, negotiation playbooks, and a legal team’s established operating methods.

That view aligns with Ivo’s commercial focus. Its contract products apply company-specific positions and workflows rather than treating every agreement as an isolated legal text.

The Sage release turns this product thesis into a public experiment. If users can reproduce useful behavior from accessible weights, closed model access becomes a weaker competitive barrier.

Harvey occupies the other side of this strategic comparison. It created the benchmark that Ivo used for Sage’s post-training, but its commercial advantage centers on a proprietary legal platform.

The comparison should not be reduced to free versus paid software. Harvey and similar vendors provide deployment, integrations, security controls, support, and product interfaces that model weights alone cannot supply.

A downloadable adapter also cannot automatically retrieve the correct playbook, maintain document permissions, or preserve an audit trail. Those surrounding systems often determine whether legal AI can enter production.

Ivo’s release therefore pressures proprietary vendors at the model layer, not across the entire product stack. The model becomes easier to inspect, while workflow execution remains a commercial battleground.

Previous open legal models also make Ivo’s “first” claim narrower than it initially sounds. Equall introduced SaulLM-7B, an MIT-licensed language model tailored for legal text, in 2024.

Other researchers have released legal encoders, reasoning models, retrieval systems, and jurisdiction-specific language models. Open legal AI did not begin with Sage.

Ivo describes Sage as the first open model released by a legal AI company and post-trained for long-horizon contract work. That carefully defined claim distinguishes agentic contract execution from general legal language modeling.

The distinction is meaningful, although it should remain attached to the claim. Sage’s novelty lies in its combination of multi-step contract training, downloadable weights, permissive licensing, and published traces.

The release also reflects Ivo’s competitive position. The company focuses primarily on contracts for in-house teams, while larger rivals often address broader law-firm and enterprise legal work.

Ivo announced a major financing round earlier in 2026 to expand its contract intelligence business. Its funding announcement described a product that works inside Microsoft Word and helps lawyers review clauses, identify risks, and suggest redlines.

Publishing Sage can attract developers and researchers who might otherwise overlook a smaller specialized vendor. It can also encourage outside experimentation that reveals new uses, failure modes, and evaluation methods.

The decision carries competitive risk. Rivals can download the same adapter, study its training approach, and incorporate useful findings into their own systems.

Ivo appears willing to accept that risk because it believes the model itself will become interchangeable. Under that thesis, customer context and product execution create more durable differentiation than private weights.

Legal departments should view this as a testable proposition, not a settled industry outcome. Open models offer control and inspectability, while managed platforms can reduce technical and operational burdens.

The more interesting question is whether customers begin demanding comparable transparency from every vendor. Public evaluations could pressure providers to disclose stricter completion rates, failure traces, and task-level evidence.

If that happens, Sage’s largest impact may not come from direct adoption. It may come from changing what sophisticated buyers expect to see before trusting an AI contract system.

The Strict Scores Reveal the Legal Judgment Gap

Sage improves measurable contract performance, but its results also show why legal AI still requires qualified supervision.

The most important number in the release is not the 91.3 percent contract criterion rate. It is the 16.4 percent rate for completing every contract criterion successfully.

That gap captures a fundamental problem in legal automation. A system can satisfy most rubric items while still producing a deliverable that a lawyer cannot safely accept without review.

A missing criterion might involve formatting or a secondary instruction. It might also involve an escalation decision, a negotiation position, or language affecting commercial exposure.

Aggregate benchmark scores cannot express those differences by themselves. Legal teams need failure taxonomies that separate harmless omissions from errors with material consequences.

Sage’s intended-use statement acknowledges this boundary. Ivo says the model produces legal work for review by a qualified lawyer, rather than replacing professional judgment.

The warning becomes more significant when contract workflows span many actions. Each additional search, edit, and tool call creates another opportunity for error or drift.

Long context does not guarantee consistent attention. A model can technically process an agreement and supporting documents while failing to connect the right provision with the correct playbook rule.

Ivo previewed a second benchmark at its conference to examine these judgment failures. The Ivo-micro1 Contract Bench evaluates prioritization, restraint, deal adaptation, escalation, and playbook compliance.

According to Ivo’s preliminary findings, tested frontier models handled acceptance or no-change decisions correctly 82 percent of the time. Their reported success fell to 46 percent when they needed to counter or reject language.

The company also reported that models added a required new negotiating point only 23 percent of the time. They satisfied 23 percent of expert-defined escalation criteria on average.

These figures come from Ivo and its research partner, not an independent publication. The company said a public result set would follow, so outsiders cannot yet fully audit every preliminary claim.

The pattern still identifies an important distinction. Contract review is not merely detecting clauses and producing redlines. It involves knowing when silence, resistance, or human escalation is appropriate.

Ivo reported another behavioral difference in editing style. Attorneys made 64 percent of their changes inline, with an average length of 92 characters.

The tested models reportedly made only 14 percent to 37 percent of changes inline. Their edits averaged between 207 and 369 characters, suggesting a stronger tendency to rewrite whole passages.

Broad rewrites create practical problems even when the language appears reasonable. They increase review time, disturb negotiated wording, and make it harder to understand exactly what changed.

Deal-specific context created another weakness. Ivo said models met 51 percent of criteria involving standard playbook positions, but only 38 percent when the deal changed the correct response.

That decline supports Jung’s argument about context. A playbook supplies general rules, yet the system must still determine when facts justify an exception.

Organizations considering Sage therefore need more than model access. They need curated documents, permission controls, task-specific instructions, evaluation sets, escalation rules, and accountable human reviewers.

They also need reliable knowledge retrieval. Contract agents can only apply internal policies when those policies remain current, discoverable, and linked to the right matter.

A searchable knowledge base illustrates the surrounding challenge. Specialized models still depend on well-governed source material rather than relying entirely on their trained parameters.

Deployment brings additional concerns. Sage’s DeepSeek lineage can trigger procurement, jurisdiction, and supply-chain questions even when organizations run the model on their own infrastructure.

Self-hosting keeps contract content within infrastructure controlled by the operator. However, local deployment does not automatically solve access control, logging, model governance, or licensing review.

The hardware requirement creates another barrier. Sage uses adapter weights, but the underlying base model remains large enough to require serious GPU capacity for practical serving.

“Free” therefore describes access to the weights, not the total cost of operation. Teams still need computing infrastructure, engineers, evaluation work, security review, and legal oversight.

Open access can make those costs easier to investigate because buyers are not limited to a vendor demonstration. It can also transfer more implementation responsibility onto the customer.

The right conclusion is neither that Sage is production-ready for unsupervised legal work nor that its low strict score makes it irrelevant. The release provides a measurable starting point for improving a difficult workflow.

Its transparency exposes limitations that closed vendors might present less clearly. That candor could strengthen trust if independent researchers reproduce the gains and identify predictable failure boundaries.

Three Signals Will Decide Whether Ivo’s Bet Works

Independent replication, real adoption, and competitive disclosure will determine whether Sage changes legal AI or remains a research release.

The first signal is independent technical validation. Researchers need to reproduce Sage’s reported gains using the published artifacts and comparable infrastructure.

Replication should test more than the average criterion rate. It should examine strict completion, judge consistency, failure severity, tool-use stability, and sensitivity to different prompts.

Researchers should also compare Sage with newer open and closed models under identical conditions. A benchmark result can age quickly when base models, inference systems, and agent frameworks change.

The strongest validation would come from legal experts reviewing blinded deliverables. Human assessment can reveal whether benchmark improvements translate into clearer, safer, and more usable work.

The second signal is meaningful adoption outside Ivo. Download counts alone cannot show whether a model has become useful.

Watch for legal technology vendors, universities, corporate legal teams, and research laboratories publishing adaptations or evaluation results. Those projects would demonstrate that Sage can support work beyond its original environment.

Fine-tuned variants would provide another indicator. Teams might adapt Sage for procurement contracts, employment agreements, privacy addenda, or jurisdiction-specific drafting.

Production use will require evidence about latency, infrastructure demands, error handling, and governance. Case studies should explain the complete workflow rather than crediting every improvement to the model.

Negative evidence will matter as much as success. If teams find that serving requirements or integration work outweigh the benefits, proprietary platforms will retain a strong advantage.

The third signal is how competing legal AI companies respond. They do not need to release their own weights to answer Ivo’s challenge.

A competitor could publish task-level evaluation traces, allow customer-controlled testing, or support deployment within customer-managed infrastructure. It could also disclose stricter completion metrics alongside average scores.

Silence would not prove that closed systems underperform. However, buyers may increasingly ask why a vendor cannot provide evidence comparable with an open model’s public record.

The new Ivo-micro1 benchmark deserves particular attention. Public results would let researchers inspect claims about excessive concessions, weak escalation, broad rewriting, and poor deal adaptation.

Those behaviors sit closer to professional judgment than generic legal knowledge. If the benchmark measures them reliably, it could become more consequential than Sage itself.

The release also creates a clear falsification test for Ivo’s broader thesis. Sage succeeds strategically if accessible models become adequate foundations while context and workflow determine real-world quality.

That judgment weakens if proprietary foundation models maintain a durable lead that specialized post-training cannot close. It also weakens if organizations cannot operate Sage economically or safely.

For legal teams, the immediate action is not replacing existing systems. It is developing evaluation cases that reflect actual agreements, policies, exceptions, and escalation rules.

Those cases let buyers compare open and closed systems using the same evidence. They also expose whether a vendor’s headline score matches the work that matters inside the organization.

Ivo Sage open-source legal AI gives technical teams a model they can inspect instead of accepting a product demonstration at face value. Its strict results also prevent easy triumphalism.

The release shows that focused post-training can improve long contract workflows. It simultaneously shows that dependable end-to-end legal judgment remains unsolved.

That combination makes Sage worth watching. Will independent teams reproduce the gains, publish useful adaptations, and force proprietary vendors toward greater transparency? The answers will determine whether open legal models become infrastructure or remain compelling experiments.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page