10 min read

Your Legal AI Vendor Is Selling You Cheap Wins

Your Legal AI Vendor Is Selling You Cheap Wins

THE TECHNOLOGY BLIND SPOT

On a Monday afternoon in April 2026, a twenty-three-year-old amateur named Liam Price prompted ChatGPT and produced a proof of an Erdős conjecture that had resisted professional mathematicians for sixty years. Eighty minutes. One run. Within a day, Fields Medalist Terence Tao read the proof and called it a meaningful contribution to number theory. The website that catalogues unsolved Erdős problems updated its status: Problem #1196, proved.

That is the version of the story the AI vendors want you to read.

Two months earlier, in February 2026, Tao told The Atlantic something more useful. Most of the AI Erdős solutions arriving every week were not breakthroughs. They were “cheap wins,” in Tao’s phrase. Sophisticated literature retrieval dressed in the language of discovery. The Price proof was Tao’s exception, not his rule. The rule was the rest of the catalogue. And the rule told a quieter story about what AI actually does well. It enables what Tao called “population studies.” It cannot do what mathematicians actually do. A mathematician knows two plus two is four, not five.

That distinction maps directly onto the legal AI vendor pitch deck Catherine reviewed last week. AI cannot do what attorneys actually do. Attorneys know “may” and “shall” are not synonyms. One word creates discretion. The other creates obligation. Population-study work is real value: contract triage, document review, issue spotting across thousands of files. Bespoke judgment work is something else: the novel-issue brief, the settlement strategy, the doctrinal argument. Not what these tools do well, and the tools cannot tell Catherine which side they sit on. The vendors are selling one tool for both categories. Tao gave the profession the vocabulary to call the bluff.

The Methodological Split

In the Atlantic interview, Tao reached for an analogy from medicine. Eighteenth-century physicians studying a rare disease examined one patient at a time and recorded symptoms in meticulous notes. Twenty-first-century clinical trials administer a drug to a thousand people and produce statistics that no individual case study could yield. Mathematics, Tao said, “is still very much at the case-study level. A paper will take one or two problems and study them to death in a very handcrafted, intensive way.” AI tools enable what Tao called population studies.

The legal profession sits at the same juncture. A litigator preparing for trial works through case law one opinion at a time, reading for facts, holdings, and implicit reasoning that no headnote captures. A transactional attorney negotiating a complex agreement works through clauses one provision at a time, weighing language against jurisdictional context and client intent. That is the case-study mode. The handcrafted, intensive style. It is what attorneys have done for centuries because it is what the work requires.

Large portions of legal practice are not case-study work at all. They are population-study work in everything but name. A document review across forty thousand files for privilege, responsiveness, and key terms is a population study. A first-pass review of two hundred vendor agreements for unusual indemnification clauses is a population study. An issue-spotting exercise across a year of regulatory filings for early signals of enforcement priorities is a population study. Tao’s framework fits cleanly here. The work is voluminous, pattern-based, and benefits from systematic application of known criteria across many similar items. AI does this well.

What AI does not do well is the case-study work that sits on the other side of the methodological line. The novel-issue research memo. The settlement-value calculation in a matter with no comparable precedent. The strategic recommendation about whether to file in state or federal court when the client’s industry is being reshaped by an unsettled regulatory regime. Those are case-study problems. They require what Tao said current AI lacks: creativity and accurate self-confidence ratings. The model produces output. The output is fluent. The output also cannot tell Catherine when it is operating outside its competence.

The Vendor Conflation

The vendor pitch deck does not draw this line. It cannot. The financial model depends on the conflation.

Open any major legal AI platform’s marketing materials. The same architecture appears every time. A capabilities matrix lists contract review next to brief writing next to research synthesis. A case study from a firm that used the tool for due diligence sits next to a testimonial from a partner who used the tool for strategic memo drafting. A pricing tier grants access to the population-study workflows and the bespoke-judgment workflows without distinguishing between them in the user interface. The tool feels like one tool because it is one tool. The marketing presents it as one tool because the marketing reflects the underlying architecture.

Tao identified the failure mode in the Atlantic interview directly. Current AI systems, he said, lack “accurate self-confidence ratings.” The model that produces a precedent-pattern match for a contract clause does not signal that it has done so with eighty-percent confidence. The model that fabricates a quotation from a Fourth Circuit opinion does not signal that it has produced a hallucination. Both outputs arrive in the same prose register, with the same surface fluency, attributed to the same authority. The interface gives Catherine no mechanism to know which category she is reading.

That is not a feature gap that an interface refresh will close. It is a structural property of how the tools work. A February 2025 paper by Hagai Simhi and colleagues, known in the field as the CHOKE study, documented the phenomenon: high-certainty hallucination produced by models that possess the correct underlying knowledge. The phenomenon resists prompt-quality variation as a fix. Architecture, not tuning. [See Your AI Knows the Right Answer, The Technology Blind Spot (2026).]

The Cost Is Named

The cost of that gap sits in the sanctions docket. The cost is named.

Steven Howe argued an immigration appeal before the Sixth Circuit in 2026. The brief he submitted contained quotations attributed to United States v. Washington and United States v. Lawrence. The quotations did not exist. The cases existed. The quotations had been generated by Westlaw’s CoCounsel, a legal-research AI marketed by Thomson Reuters as producing “100% hallucination-free linked legal citations.” On April 3, 2026, the Sixth Circuit removed Howe from the appeal. United States v. Farris, No. 25-5623, 2026 WL 915082 (6th Cir. Apr. 3, 2026) (per curiam). [See Your AI Research Tool Fabricated the Quotation, The Technology Blind Spot (2026).]

Matthew Reeves was a practice group co-leader at Butler Snow. The firm had a written AI policy. The policy required written approval before using ChatGPT for legal research. Reeves used ChatGPT without approval and without verifying the output. Judge Anna Manasco found his conduct “tantamount to bad faith.” Johnson v. Dunn, 792 F. Supp. 3d 1241, 1268-69 (N.D. Ala. 2025).

Both attorneys made the same mistake at the architectural level. Both treated population-study output as case-study output. Howe asked CoCounsel to produce something the tool was structurally incapable of producing without error: novel quotations from primary sources. The tool produced fluent text. Howe submitted it. The Sixth Circuit removed him from the appeal. Reeves asked ChatGPT to produce case citations for a brief. The tool produced citations that did not exist. Judge Manasco found the conduct sanctionable.

In neither case did the tool tell the attorney it was operating outside its competence. The interfaces signaled nothing. The output arrived with the same surface fluency that document-triage output arrives with at the same firms. That is the conflation Tao’s framework exposes. The attorneys learned from their training, from the marketing, and from the interface to treat the tool as if it produced one type of output across all task categories. Tao’s mathematics community had the same problem until the cheap-wins framing arrived to make the distinction visible.

The Steelman

The strongest version of the vendor counterargument runs like this. AI capability advances each year. Hallucination rates dropped sixty-four percent year over year in 2025. Retrieval-augmented generation cuts errors by seventy-one percent on standard benchmarks. Agentic AI architectures reduce single-turn errors through multi-step verification loops. The Howe and Reeves cases were both cases where attorneys violated firm policy and bar guidance. The architecture is improving. The training will catch up. The verification duty exists. Catherine should not refuse to use a generation of tools that solves real problems because an earlier generation produced sanctions.

That argument is not wrong on the technical facts. It is wrong on the framing.

Tao’s distinction does not depend on hallucination rates. It depends on what the model does. A population-study tool that matches contract clauses against a precedent corpus with ninety-eight-percent accuracy belongs in a different category from a bespoke-judgment tool that produces an answer to a question no precedent corpus contains. The first tool does what AI does well: high-volume pattern matching across known criteria. The second tool needs to do what Tao identified as the structural gap: case-study reasoning that requires creativity and the ability to know when one is wrong. Improving hallucination rates on the first task does not transfer to the second.

The legal AI market has not internalized this. The market sells one product for both categories at the same price tier. The Robin AI collapse made the structural problem visible at the company level: a London-based startup that launched contract review, drafting, negotiation, and managed services simultaneously, raised sixty-nine million dollars, claimed eighty-percent contract review time reduction, and ended up on an insolvency marketplace by October 2025 with ten million in revenue against fourteen million in losses. [See Every Failed AI Project Breaks the Same Rule, The Technology Blind Spot (2026).] Robin AI did not fail because the product was bad. It failed because no single tool could deliver across the four categories the company sold simultaneously.

Where the Math Analogy Breaks

One asymmetry between Tao’s domain and legal practice deserves naming up front. Tao developed his framework inside a discipline where ground truth is verifiable in principle. A proof either works or it does not. Lean, the formal proof assistant, mechanically checks every step. Legal work has no compile-time check. The check is appellate review, years after the work is filed. Catherine cannot verify her associate’s privilege determination through a formal proof system. She has to trust the work or do it herself.

That limit makes the vendor problem worse, not better. In mathematics, AI’s missing self-confidence rating slows the work down. Humans verify the proof, refine the argument, and the field absorbs the contribution. In law, AI’s missing self-confidence rating produces a sanctions order, a malpractice claim, or a privilege waiver. By the time appellate review arrives, Howe is off the case and the court has disqualified Reeves. Tao described the failure mode before Howe filed his brief. The legal AI market sells the failure mode as a feature.

Three Actions This Week

Three actions belong on Catherine’s calendar.

Email the firm’s primary AI vendor with one specific request. Not “what is your accuracy rate?” That question gets a marketing answer. The specific request: produce the underlying study, the methodology, and the error-rate distribution by query type. Distinguish between population-study task accuracy (contract clause matching, document classification, precedent retrieval) and bespoke-judgment task accuracy (novel-issue research, primary-source quotation, original argument synthesis). If the vendor cannot produce the breakdown, Catherine has Tao’s answer. The tool sits on the cheap-wins side of the line, regardless of how the marketing presents it.

Audit the firm’s verification protocol against the methodological split. Most firm protocols treat AI output as one category and verify accordingly. Replace that structure with two protocols. Population-study output gets sampled review and quality-control statistics. Bespoke-judgment output gets attorney verification against primary sources before any external transmittal. Build the distinction into the workflow, not the training deck.

Cancel any subscription that prices the two categories identically. The firm is paying premium-tier rates for a tool that delivers commodity-tier value on most matters and unverified-tier risk on the rest. The pricing model assumes both categories are worth the same. Tao’s framework says they are not. Catherine signs the renewal.

The Close

Liam Price’s eighty-minute proof did not rewrite mathematics. It demonstrated, briefly, what AI can do when the question matches the architecture and the verification matches the stakes. Tao read the proof. Lichtman extended it. The mathematics community absorbed the contribution. None of that happened by accident. Tao had been distinguishing cheap wins from real contributions for two years before Price typed his prompt. The vocabulary made the verification possible.

The legal AI market has not done that work. No one in the market has called the bluff. No one in the supply chain has the vocabulary or the incentive to make Catherine’s distinction for her. Not the vendor. Not the bar association. Not the procurement committee. Not the malpractice carrier. Catherine has to call the bluff herself.

Her vendor is selling cheap wins at breakthrough prices. She paid for breakthroughs. The reply to her Thursday email will tell her what she actually bought.

About the Author

JD Morris is Co-Founder and COO of LexAxiom, an Agentic AI platform for the business of law. Over a 25-year career, he has built and scaled enterprise technology products across Dell, EMC, VMware, and Cisco, including the first exabyte eDiscovery platform. He holds dual MBAs from Columbia Business School (Finance) and UC Berkeley Haas (Marketing), a Master of Legal Studies in Cybersecurity Law from Texas A&M, and a Master of Engineering from George Washington University. He writes The Technology Blind Spot on the intersection of emerging technology and law. Connect with him on LinkedIn at www.linkedin.com/in/jdavidmorris, on X at @JDMorris_LTech, or on Bluesky at @JDMorris-ltech.bsky.social.

References

1. Matteo Wong, The Edge of Mathematics, Atlantic (Feb. 2026), https://www.theatlantic.com/technology/2026/02/ai-math-terrance-tao/686107/.

2. Lauren Leffer, Amateur Armed with ChatGPT ‘Vibe Maths’ a 60-Year-Old Problem, Sci. Am. (Apr. 28, 2026), https://www.scientificamerican.com/article/amateur-armed-with-chatgpt-vibe-maths-a-60-year-old-problem/.

3. T. F. Bloom, Erdős Problem #1196, https://www.erdosproblems.com/1196 (last accessed May 7, 2026).

4. United States v. Farris, No. 25-5623, 2026 WL 915082 (6th Cir. Apr. 3, 2026) (per curiam).

5. Johnson v. Dunn, 792 F. Supp. 3d 1241 (N.D. Ala. 2025) (Manasco, J.).

6. Hagai Simhi et al., Trust Me, I’m Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer, arXiv:2502.12964 (Feb. 2025), https://arxiv.org/abs/2502.12964.

7. Varun Magesh et al., Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Stan. RegLab (May 2024), https://reglab.stanford.edu/publications/hallucination-free.

8. Model Rules of Pro. Conduct r. 1.1, cmt. 8 (Am. Bar Ass’n 2024).

9. ABA Comm. on Ethics & Pro. Resp., Formal Op. 512 (2024).



Originally published on LinkedIn Newsletter — The Technology Blind Spot

Leave a Reply

Discover more from The Technology Blind Spot

Subscribe now to keep reading and get access to the full archive.

Continue reading