Merlin AI Model Benchmark: Nine Times the Cost, No Accuracy Win
Merlin Search Technologies tested seven AI models and found a sharp reversal: the costliest model consumed over nine times more per report without delivering higher citation accuracy.
The Merlin AI model benchmark gave every system the same opioid litigation documents, two investigative questions, and one reporting assignment. The cheapest and fastest model finished first on the vendor’s fidelity measure. The most expensive finished fourth.
That result challenges a common buying shortcut. A newer or more expensive frontier model is not automatically the best model for every professional task. However, Merlin did not disclose which anonymized label belonged to each product, and six models were statistically tied.
The study also measured a narrow form of quality. It checked whether statements traced back to the supplied documents. It did not determine which report offered the strongest legal reasoning, identified the most important evidence, or exercised the best judgment.
Those qualifications make the result more useful, not less. Merlin’s test shows why enterprises need task-specific evaluations, confidence intervals, and auditable citations before routing every assignment to a premium model.
What the Merlin AI Model Benchmark Actually Tested
Merlin compared seven models inside one controlled legal investigation workflow, then graded their citations and source fidelity.
The company tested models from OpenAI, Anthropic, and Google through Alchemy, its investigation and discovery platform. The lineup included GPT-6 Astra, GPT-5.6 Sol, GPT-5.6 Terra, Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5, and Gemini 3.1 Pro.
Each model received the same collection of publicly available opioid litigation documents. The materials covered Mallinckrodt and also referenced McKesson, Insys, and Valeant.
Merlin avoided client data. That choice reduced confidentiality concerns and made the collection safer to reuse, although the underlying report was not presented as a public, independently reproducible benchmark package.
The models answered two questions. One concerned issues across the company’s pain-management products. The other asked what internal notes and analyst commentary revealed about financial trouble and credit problems.
The second question contained a deliberate trap. Some credit commentary concerned other companies, while parts of the Mallinckrodt record described ordinary billing disputes. A careless system could merge those threads and imply a financial crisis unsupported by the evidence.
Every model then produced a full investigative report with citations. Across the seven reports, the systems generated 777 citations and 454 cited sentences, according to the published methodology.
Merlin anonymized the outputs as Models A through G before grading. The evaluators therefore did not receive provider names, although the company retained a recorded random seed so it could reconstruct the labeling.
The public scorecard kept that masking in place. Readers can compare the results, but they cannot independently determine whether Model D was an OpenAI, Anthropic, or Google product.
Model D led with a fidelity score of 3.3, where lower was better. It completed each report in an average of 37 seconds and recorded no broken or misdirected citations in the displayed results.
Model C received a fidelity score of 4.1 and took 110 seconds on average. It was the most expensive system in the run, consuming more than nine times the report cost of Model D.
Model E finished last with a fidelity score of 4.9. The remaining six models had overlapping confidence ranges, so Merlin treated them as statistically tied rather than presenting a falsely precise ranking.
That distinction matters. The headline reversal is real within the study, but it does not establish that Model D has a universal accuracy advantage. It establishes that paying much more did not produce a detectable fidelity improvement in this particular run.
The Cheapest Model Won a Narrow but Important Test
The benchmark undermines price as a proxy for accuracy, but it does not identify one model as the universal winner.
Model D had the cleanest citation record, the shortest average completion time, and the lowest operating cost in the comparison. Model C cost more than nine times as much and placed fourth on the displayed scorecard.
Neither of the two newest releases, GPT-6 Astra and Claude Fable 5.1, finished first. Since Merlin withheld the label key, the public cannot tell exactly where either model placed.
That secrecy limits purchasing value. A legal team cannot read the article and confidently select a specific provider. It can, however, reject the assumption that the newest premium option deserves every workload by default.
The results also exposed differences that one rank cannot capture. Model A covered the broadest range of distinct issues, recording a breadth score of 24.9. It produced no broken citations but had two citations that pointed to the wrong supporting passage.
Model G achieved nearly as much breadth, with a score of 23.6. However, nine of its 93 checked claims pointed readers to the wrong paragraph. Merlin said the underlying facts were present elsewhere, but a lawyer following those references would need to search for them.
Model D was more precise but narrower. Its breadth score reached 15.9, well below Models A and G. A careful report that misses a decisive theme can still be less useful than a broader report with a few repairable citation errors.
This is the central reversal. The cheapest model led the fidelity ranking, but its win did not erase the value of coverage, synthesis, or legal judgment.
The study found no fabricated facts in any of the seven reports. Every specific name, figure, date, and quotation traced back to the document collection under Merlin’s grading system.
That finding is encouraging, although it should remain bounded by the task. The models worked from a controlled evidence set and produced cited reports inside Alchemy. They were not asked to conduct open-web legal research or retrieve authorities from an unrestricted database.
A model that stays grounded within a curated collection faces a different problem from a chatbot answering questions from memory. Retrieval controls, citation structures, and the surrounding application can materially change failure rates.
Earlier academic work illustrates that context. Stanford researchers evaluating commercial legal research products defined a hallucination as a response that was incorrect or not properly grounded. Their legal AI study found that retrieval alone did not eliminate unsupported answers.
Merlin’s result therefore speaks to the entire system, not only the base models. Alchemy supplied the document summaries, citation markers, prompts, report structure, and checking process. The winning output emerged from that combined environment.
The correct conclusion is not that cheaper models are always more accurate. It is that model price alone provides little evidence about performance on a defined enterprise workflow.
How the Eight-Pass Evaluation Checked Every Citation
Merlin’s strongest contribution is the evaluation pipeline, which uses deterministic checks wherever arithmetic can replace model judgment.
The process began by hiding model identities and removing accidental mentions from reports. Later stages saw only the fixed anonymous labels.
The second pass audited Alchemy’s own source summaries. Each summary contained markers connecting it to paragraphs in the original documents, but some markers were defective.
The audit found 24 flawed summary sections. Merlin excused citations that failed solely because the platform had supplied a bad marker. This prevented the benchmark from penalizing a model for an upstream data error.
That step is easy to overlook. Evaluations often assign every visible failure to the language model, even when retrieval, parsing, or document preparation caused the problem.
The third pass resolved citations mechanically. Code checked whether each cited document and section existed and whether the cited paragraph fell inside the identified range.
A simple document reference without a section or paragraph was treated as prose, not as a failed citation. This rule prevented the checker from inflating its error count through overly aggressive parsing.
The fourth pass evaluated whether cited sentences were grounded in the supplied summaries. This was the first stage that used an AI judge.
Merlin said an earlier version treated every fact absent from the listed citations as invented. That approach punished synthesis when a model cited only part of the evidence behind a multi-source sentence.
The revised system used two stages. First, a grader identified information it could not find in the cited material. A second stage then searched the wider collection before assigning a verdict.
The possible outcomes were supported, uncited source, overreach, or fabricated. An uncited source meant the fact existed elsewhere in the collection. Overreach meant the conclusion exceeded the evidence without inventing a checkable fact.
Fabrication covered a specific name, figure, date, or quotation found nowhere in the collection. Merlin weighted that category three times more heavily because an invented detail can enter a filing or client memorandum.
The fifth pass measured breadth and repetition. The system converted sentences into semantic vectors, mathematical representations of meaning, and estimated how many independent ideas each report contained.
That design tried to avoid rewarding length for its own sake. A longer report received no automatic benefit if its additional sentences repeated the same points.
The sixth pass calculated scores and uncertainty ranges. Cost and speed appeared in the scorecard but did not affect the accuracy ranking.
Merlin used resampling across the two questions to estimate confidence intervals. When intervals overlapped, it called the models tied. That is why six of the seven systems shared the leading statistical group despite their ordered display.
The seventh pass graded the grader. Two additional AI judges from different providers reviewed every flagged claim and a sample of passed claims.
A majority vote resolved disagreements. Three of 33 initial flags were overturned, producing a reported false-alarm rate of nine percent. The reviewers also spot-checked 37 passed claims and did not flag any.
All three graders fully agreed on 58 of 70 reviewed claims. That leaves meaningful disagreement, but the disclosure gives readers a clearer view of measurement uncertainty.
The eighth pass generated a written memo. An AI system drafted prose, while computed metrics supplied every number and code produced the scorecard.
Merlin also required every criticism to name a better-performing comparison model. Its full evaluation report includes definitions, metrics, and finding-level appendices.
This combination of deterministic checks and bounded AI judgment follows a sound general principle. Use code for existence tests, ranges, counts, and reproducible calculations. Reserve probabilistic graders for semantic questions that arithmetic cannot answer.
It also resembles NIST’s recommendation that organizations document test sets, metrics, uncertainties, and deployment conditions. The AI RMF Core also calls for repeatable evaluation and continued monitoring after deployment.
For teams building a searchable knowledge base, the lesson extends beyond litigation. Citation checking should test both retrieval and generation because either layer can break the evidence chain.
Why the Result Pressures Enterprise AI Buyers
The benchmark puts pressure on buyers who standardize on one premium model before measuring their actual workloads.
Enterprise procurement often treats models like software editions. A higher-priced or newer option appears safer because it promises stronger reasoning, a larger context window, or better scores on general benchmarks.
Merlin’s test shows the weakness in that logic. A general capability advantage does not guarantee better performance inside a document-grounded investigation workflow.
The underlying assignment demanded several distinct abilities. Models had to read summaries, separate evidence about different companies, avoid overstating routine disputes, synthesize product issues, and attach useful citations.
One model handled the citations cleanly but covered fewer issues. Other models found more distinct threads while producing more citation friction. The right choice depends on which error creates the greater professional risk.
A first-pass review of a large collection may favor speed and broad recall. A final report prepared for counsel may place more weight on precise citations, complete synthesis, and disciplined conclusions.
That difference supports model routing, which assigns each task to a model selected for its measured requirements. It does not support automatically choosing the cheapest system.
A defensible routing policy would begin with risk. Teams should identify whether a task requires factual extraction, issue spotting, long-form synthesis, legal judgment, or final client communication.
They should then test representative examples under the same prompts, retrieval settings, and document conditions used in production. A public leaderboard cannot reproduce those conditions.
Model choice also affects review costs. A fast system that produces narrower reports can shift work back to lawyers who must find omitted themes. A broad system with weak citation precision can create a different burden because reviewers must trace claims manually.
The benchmark’s operational cost metric excluded that downstream labor. It measured model usage for producing reports, not the total cost of reviewing, correcting, and approving them.
For legal buyers, this distinction is especially important. The American Bar Association’s AI ethics opinion says lawyers must consider duties involving competence, confidentiality, communication, supervision, candor, and reasonable fees.
A model’s fluent output does not transfer responsibility away from counsel. Lawyers still need review procedures proportionate to the assignment and the risks created by incorrect or unsupported statements.
The same principle applies outside law. Financial analysts, compliance teams, researchers, and product managers all rely on outputs that compress large evidence collections into decisions.
A model can cite real documents and still miss the decisive issue. It can also reach a sensible conclusion while pointing to the wrong passage. Those are separate failure modes and deserve separate metrics.
Buyers should therefore resist a single composite “quality” score. They need a scorecard that reflects the work: citation validity, claim grounding, coverage, contradiction handling, latency, and human correction time.
Merlin also argues for user choice across model providers. That position serves the company’s commercial design because Alchemy offers multiple model families.
Still, the evidence supports the narrower point. When six models are statistically tied on fidelity but differ in speed, breadth, and operating cost, locking every task to one provider leaves measurable tradeoffs unused.
What the Benchmark Does Not Prove
This was a vendor-run, single-workflow evaluation with hidden model identities and a deliberately limited definition of accuracy.
Merlin developed Alchemy, ran the benchmark, designed the scoring process, and published the interpretation. EDRM republished the article with a note stating that the views belonged to the authors.
That does not invalidate the work. It does mean readers should treat the findings as a documented vendor evaluation rather than an independent comparative trial.
The hidden answer key creates another constraint. Anonymization protected the graders from brand bias, but continued secrecy prevents public model-level scrutiny.
Readers cannot compare the reported positions against other benchmarks. They also cannot test whether a model’s known behavior explains a particular pattern in breadth, speed, or citation errors.
The benchmark used two questions drawn from one litigation collection. Two questions can expose meaningful failure modes, but they cannot represent every discovery assignment or legal domain.
The reported confidence intervals try to acknowledge this uncertainty. However, resampling two questions cannot create the diversity of a larger test set covering contracts, regulatory matters, privilege review, chronology building, or conflicting witness accounts.
The models apparently completed one evaluated run under one configuration. Repeated trials could reveal output variance, particularly when generation settings permit different phrasing or evidence selection.
The public materials also do not provide enough information to calculate the total system effect of prompts, summary quality, retrieval limits, and model configuration independently. A different application harness could change the order.
Most importantly, fidelity was not synonymous with legal accuracy. The test measured whether statements traced to sources and whether citations pointed to the right material.
It did not grade the quality of legal reasoning. It did not determine whether a model identified the most important evidence, reconciled conflicts persuasively, or anticipated an opponent’s argument.
Merlin acknowledged this limitation directly. A model could rank well by listing carefully supported facts while avoiding the difficult conclusions that make an investigative report useful.
The absence of detected fabrication also requires restraint. It means the evaluation pipeline found no specific invented fact in these seven reports. It does not establish a zero-hallucination rate for the models, Alchemy, or legal AI generally.
A system’s failure rate changes with the task. Open-ended research, incomplete collections, ambiguous questions, and adversarial documents create different pressures from a controlled synthesis exercise.
Automated grading introduces its own uncertainty. Merlin’s additional judges overturned three initial flags, demonstrating why a single model should not serve as unquestioned evaluator.
The graders agreed unanimously on 58 of 70 reviewed claims. That is substantial agreement, but it also shows that semantic classification remained contestable.
Independent human legal review would strengthen the evidence. Domain experts could judge whether the evaluation categories matched professional expectations and whether a supposedly supported claim preserved the source’s meaning.
NIST’s generative AI profile emphasizes pre-deployment testing while recognizing that risks differ across model, application, and ecosystem levels. Merlin’s benchmark evaluates all three layers together but cannot isolate them fully.
The result should therefore change buying behavior without ending the investigation. It is evidence against price-based selection, not proof that one anonymous economical model should replace frontier systems.
Three Signals That Will Test the Cost Versus Accuracy Reversal
The next test is whether Merlin’s reversal survives broader tasks, independent review, and repeated model updates.
The first signal is a larger judgment benchmark. Merlin says it is building a different evaluation for conflicting evidence, causal reasoning, issue importance, and opposing arguments.
That test matters because premium models often justify their higher operating cost through reasoning rather than citation mechanics. If economical models remain competitive there, the case for default frontier routing will weaken substantially.
If premium systems open a clear lead, the current result will look more like task specialization. Teams could then reserve frontier models for complex synthesis while using economical systems for evidence-grounded reporting.
The second signal is independent replication. A credible follow-up should disclose prompts, settings, model versions, repeated runs, scoring code, and a reusable document package where licensing permits.
It should also include human legal reviewers who do not work for the platform vendor. Agreement between deterministic checks, automated judges, and domain experts would strengthen the method.
Replication might preserve anonymity during grading while publishing the label key afterward. That design would retain protection against brand bias and give buyers actionable model-level evidence.
The third signal is performance drift after model and platform updates. Providers revise model snapshots, inference systems, and pricing structures frequently. Retrieval pipelines and summarization prompts also change.
Merlin says new arrivals receive the same test. Publishing a stable time series would show whether rankings persist or merely capture one favorable snapshot.
Teams should watch more than the winner. Changes in broken citations, misdirected references, breadth, grader disagreement, latency, and correction time would reveal whether the workflow is improving.
The Merlin AI model benchmark leaves buyers with a practical challenge. Stop asking which model is best in the abstract. Define the work, identify the costly failures, and run a controlled evaluation against your own evidence.
For legal teams, that means checking whether citations exist, whether passages support claims, and whether the report captures the issues that matter. It also means keeping lawyers responsible for final judgment.
For other knowledge-intensive teams, the question is similar: does the system help people reach a defensible decision, or does it merely produce fluent text?
The ninefold cost gap makes a memorable headline. The durable lesson is more demanding. Measure the task, audit the evidence chain, include uncertainty, and retest whenever the models or workflow change.



