Editor’s Note: A vendor benchmark landed in June with a finding built to unsettle procurement assumptions: the highest-cost model among nine tested for legal document classification finished second to last, while models costing a fraction of its price clustered near the top. DecoverAI’s working paper argues that review architecture, not model capability, sets the accuracy ceiling, an interpretation its single-pipeline design cannot prove, and two other provider studies offer related though limited support.

For cybersecurity, data privacy, regulatory compliance and eDiscovery professionals, the stakes are concrete. Inadvertent production of privileged material can implicate Federal Rule of Evidence 502, and amended Rule 26(f)(3)(D), effective Dec. 1, 2025, requires early views and proposals on privilege claim timing and method, which means model choice and validation evidence may surface in meet-and-confer sessions. The paper also leaves questions open, from its 100-document single-run sample and unpublished methodological details to the governance implications of routing privileged material to models from Chinese developers.

Watch for privilege-specific benchmarks, court scrutiny of AI-assisted privilege logs, and controlled validation runs to become standard requests in the year ahead.


Content Assessment: On a 100-document responsiveness test, model price did not predict accuracy

Information - 91%
Insight - 89%
Relevance - 90%
Objectivity - 89%
Authority - 90%

90%

Excellent

A short percentage-based assessment of the qualitative benefit expressed as a percentage of positive reception of the recent article from ComplexDiscovery OÜ titled, "On a 100-document responsiveness test, model price did not predict accuracy."


Industry News – eDiscovery Beat

On a 100-document responsiveness test, model price did not predict accuracy

ComplexDiscovery OÜ Staff

The highest-cost model tested in a new nine-model legal AI benchmark finished next to last. The cheapest one, priced 68 times lower, landed 0.075 behind the leader on the study’s F1 score.

Those results anchor a working paper published in June 2026 by DecoverAI, a San Mateo, California-based eDiscovery vendor, and they carry an uncomfortable message for legal teams that equate model price with review quality. On this test, the paper concludes, buyers of the top-priced option were “paying a premium for brand rather than performance.”

What the benchmark measured

Decover Research Working Paper 2026-02 tested nine large language models from Alibaba, DeepSeek, MiniMax, Moonshot AI and Anthropic on a fixed 100-document responsiveness classification task with a gold-labeled answer key of 27 responsive and 73 nonresponsive documents. The vendor ran its own two-step production pipeline for every model, changing only the model underneath, and metered cost from actual token usage at June 2026 prices.

Eight of the nine models landed inside an 11-point F1 band, from 0.76 to 0.87, across a per-document cost spread from $0.0012 to $0.0813. Qwen 3.6 Plus posted the best F1 score, 0.868, at $0.0205 per document. Anthropic’s claude-opus-4-8, the highest-cost model tested at $0.0813 per document, scored 0.760, second-lowest in the study and below models costing 30 times less. DeepSeek V4 Pro delivered an F1 of 0.815, with precision and recall balanced at 0.815 each, for $0.0027 per document, which the paper designates its value operating point.

At 100,000 documents, the paper calculates, the spread between that value option ($270) and the highest-cost model ($8,130) is $7,860, which it equates to roughly 52 hours of contract attorney quality control at $150 per hour.



Read the fine print before repeating the numbers

The caveats deserve as much attention as the rankings. The benchmark is a single run over 100 documents, which the paper itself labels “indicative, not audited.” Small samples produce wide confidence intervals, and a handful of borderline documents could reorder the middle of the table.

The paper also does not publish the dataset, the labeling protocol, the classification instructions, model settings, the inference route, token counts, raw predictions or repeated-run results. Without those materials, neither the observed rankings nor the per-document costs can be independently reproduced.

The scope note contains a larger qualification: this is a responsiveness benchmark, not a privilege benchmark, despite the paper’s title. Privilege calls are harder and more judgment-intensive, and Decover acknowledges model rankings may differ on a privilege-specific gold set, which it says it is building as a follow-on study.

The design limits run deeper than sample size. Because the test held Decover’s pipeline constant and compared no alternative architectures, it shows only that most models clustered within this pipeline on this 100-document set. It does not establish that architecture caused the clustering, that the observed F1 range transfers to other collections, or that a model upgrade could not improve performance elsewhere. And Decover sells the pipeline credited with producing the plateau, which gives its preferred interpretation, that architecture rather than model choice sets the accuracy ceiling, a commercial interest. Buyers should treat the paper as vendor-published research making a testable claim, not as an audited independent evaluation.

The precision finding is the practical headline

For privilege review, the paper’s most useful contribution may be its warning about precision. One model, Anthropic’s claude-sonnet-4-6, achieved the highest recall in the study at 96.3 percent but flagged 58 documents as responsive when only 27 were, a precision score of 0.448. Fewer than half of its positive calls were correct.

That figure is responsiveness precision measured on this one set. It does not establish what the model’s precision would be on privilege classification, and the paper’s model-specific privilege recommendations remain hypotheses awaiting a privilege-labeled test. The conceptual point survives the caveat, though, because the error asymmetry is specific to privilege work. A missed privileged document risks inadvertent waiver, though Federal Rule of Evidence 502(b) protects against waiver where the disclosure was inadvertent, the holder took reasonable steps to prevent it, and the holder promptly took reasonable steps to rectify the error. An over-inclusive model creates the opposite problem: it can produce a privilege log padded with improper withholding claims, inviting entry-by-entry challenges and, in some circumstances, sanctions requests. The paper argues a reasonable recall rate paired with high precision produces a more defensible log than maximum recall at low precision, and recommends avoiding any model with precision below 0.75 for privilege classification without attorney review of every positive call.

Two other provider studies offer related evidence, with limits

Neither supporting study is disinterested; both come from companies with products in the fight, and each supports a narrower proposition than Decover’s. Percipient, a legal services provider, gave 10 model variants a 78-document review task that included responsiveness, privilege coding and production of a privilege log, graded by experienced attorneys. Nine of the 10 scored within eight points of one another on that composite task, though gaps widened to as much as 37 points on insurance coverage analysis, a reasoning-intensive assignment. “The industry’s conversation has largely focused on model size, speed, and cost,” Chad Main, Percipient’s founder, said in the study’s announcement. “What this research demonstrates is that the next phase of competition will be won by systems that reason better, not simply systems that generate text faster.”

LegalOn Technologies makes an adjacent argument from its 2026 contract review benchmark of 11 models, a vendor-run study in which the publisher’s own system ranked first. “The model matters, but the harness matters too,” Daniel Lewis, LegalOn’s chief executive, wrote in a sponsored thought-leadership article for Artificial Lawyer. “For contract review, that can be the difference between a system that is impressive in a demo and a system that is reliable enough for daily legal work.” LegalOn’s claim is that a purpose-built harness can beat raw foundation-model use, which is adjacent to, not proof of, Decover’s convergence thesis.

The open question sits in the privilege detail. Percipient’s composite score included privilege coding and still converged, but the study published no separate precision and recall figures for the privilege calls, so whether model performance converges on privilege specifically remains unresolved by these studies. Privilege determination, with its fact-specific judgments about legal advice, dual-purpose communications and third-party waiver, plausibly sits closer to the reasoning-heavy end where Percipient found divergence. That is this article’s analysis, not a finding of either study.

Model provenance is the question the paper skips

The governance question a chief information security officer will ask first goes unaddressed in the paper. Six of the nine model entries come from Chinese developers: Qwen 3.6 Plus, MiniMax M3, the two DeepSeek models and the two Kimi models from Moonshot AI. Several are distributed as open weights that legal teams can self-host or run on infrastructure in a jurisdiction of their choosing, but the paper does not say where its test inference ran or what deployment it assumes for client matters.

For privileged attorney-client communications, that silence is the gap to close before any cost comparison matters. Legal teams should confirm where inference occurs, what the model provider or host retains, whether outside counsel guidelines and protective orders permit the routing, and whether client consent is needed. A model that saves $7,860 per 100,000 documents buys nothing if the deployment breaches a confidentiality obligation.

New privilege log rules raise the stakes

The timing gives the debate procedural teeth. Amendments effective Dec. 1, 2025 revised Federal Rule of Civil Procedure 26(f)(3)(D) to require that the discovery plan state the parties’ views and proposals on issues about privilege claims, including the timing and method for asserting them, and revised Rule 16(b)(3)(B)(iv) to permit the scheduling order to adopt those provisions. Privilege logging moves from a late-stage task to an early negotiation. When AI-assisted classification forms part of a proposed privilege process, model selection, validation evidence and quality control protocols may become subjects of those discussions, and a party that cannot explain why it selected a given model, or produce validation results on its own document population, will be poorly positioned if the question arises.

That points to the paper’s most defensible advice: validate a proposed system against a representative, gold-labeled sample from the matter’s own document population under a documented protocol rather than relying on a vendor benchmark. Demand precision figures alongside recall, because a recall-only quote conceals the false-positive rate that inflates privilege logs. Run the cost arithmetic at the matter’s real volume, where the difference between $0.003 and $0.08 per document is trivial at 1,000 documents and decisive at 500,000. And treat model savings as budget for attorney quality control on escalated documents rather than as margin.

The paper packages its recommendations into four proposed buyer profiles: Qwen 3.6 Plus for maximum accuracy, DeepSeek V4 Pro for value, claude-haiku-4-5 for precision-first work and DeepSeek V4 Flash for first-pass review at extreme volume. These are Decover’s hypotheses, not validated privilege-review guidance, and the precision-first designation contains an internal inconsistency: the paper calls claude-haiku-4-5’s 0.875 precision the highest in the study, but its own results table shows Qwen 3.6 Plus higher on precision (0.885) as well as recall, F1 and accuracy, leaving a $0.001 per-document cost edge as the Haiku model’s only listed advantage. The paper’s architecture-floor test, the idea that a platform demonstrating F1 between 0.81 and 0.87 on a controlled sample has made model selection a pure cost decision, should be read the same way: a score below that band could reflect collection difficulty, class prevalence, labeling disagreement, criteria design or model configuration as easily as architecture.

Model pricing will keep shifting, and the specific rankings in any single-run benchmark will age quickly. The durable question for legal teams is different: when opposing counsel asks how you validated the AI that built your privilege log, what will your answer be?



News sources



Assisted by GAI and LLM Technologies

Additional reading

Source: ComplexDiscovery OÜ

ComplexDiscovery’s mission is to enable clarity for complex decisions by providing independent, data‑driven reporting, research, and commentary that make digital risk, legal technology, and regulatory change more understandable for practitioners, policymakers, and business leaders.

 

Have a Request?

If you have information or offering requests that you would like to ask us about, please let us know, and we will make our response to you a priority.

ComplexDiscovery OÜ is an independent digital publication and research organization based in Tallinn, Estonia. ComplexDiscovery covers cybersecurity, data privacy, regulatory compliance, and eDiscovery, with reporting that connects legal and business technology developments—including high-growth startup trends—to international business, policy, and global security dynamics. Focusing on technology and risk issues shaped by cross-border regulation and geopolitical complexity, ComplexDiscovery delivers editorial coverage, original analysis, and curated briefings for a global audience of legal, compliance, security, and technology professionals. Learn more at ComplexDiscovery.com.

 

Generative Artificial Intelligence and Large Language Model Use

ComplexDiscovery OÜ recognizes the value of GAI and LLM tools in streamlining content creation processes and enhancing the overall quality of its research, writing, and editing efforts. To this end, ComplexDiscovery OÜ regularly employs GAI tools, including ChatGPT, Claude, Gemini, Grammarly, Midjourney, and Perplexity, to assist, augment, and accelerate the development and publication of both new and revised content in posts and pages published (initiated in late 2022).