10 min read

When Attorneys Stop Checking AI’s Work: The Anthropomorphism Liability

On July 17, 2025, Judge David Leibowitz of the Southern District of Florida opened his sanctions order with a quotation from the late Justice Antonin Scalia on the importance of candor in judicial proceedings. The quotation was an AI hallucination. Judge Leibowitz had generated it with a ChatGPT prompt the week before. He knew it was fabricated. He used it as an epigraph anyway, to make a point about the technology he was about to sanction an attorney for trusting. Attorney James Martin Paul had cited hallucinated cases across eight different matters. He continued filing AI-generated fabrications even after opposing counsel flagged them, even after a court order put him on notice, even after a show cause hearing. Judge Leibowitz dismissed all four federal cases and ordered Paul to pay $85,567.75 in attorneys’ fees. It was the largest AI hallucination sanction in American legal history. Six days later and 700 miles north, Judge Anna M. Manasco of the Northern District of Alabama issued a different kind of opinion. In Johnson v. Dunn (N.D. Ala. July 23, 2025), the attorneys who filed AI-generated hallucinated citations were not rogue actors. They were partners at Butler Snow LLP, a large, well-regarded firm that had done everything the consultants recommend. Butler Snow had circulated an email warning about generative AI dangers. It had prohibited AI use without practice group leader approval. It had adopted verification policies when it deployed Westlaw’s CoCounsel platform. Matthew B. Reeves, the partner who used ChatGPT to generate the fabricated citation, held the title of assistant practice group leader. He was the gatekeeper the policy relied on. Judge Manasco disqualified the attorneys from the case, directed the clerk to notify bar regulators in every jurisdiction where they held a license, and suggested that existing rules may not adequately address the harms caused by this category of conduct. I analyzed the governance failure in the prior installment of this series, “Every Failed AI Project Breaks the Same Rule,” through the lens of systems design. That analysis explained why Butler Snow’s policy failed structurally. This analysis addresses the cognitive mechanism the policy could not reach: anthropomorphism, the attribution of human judgment to a machine that has none. **What Is Happening and Why It Matters Now** Anthropomorphism is a cognitive default produced by neural architecture that evolved to detect agency quickly. A tool that speaks in complete sentences, acknowledges uncertainty, apologizes for errors, and adapts its tone to context activates the same trust pathways that govern human relationships. The attribution is automatic. It operates below deliberate judgment. Large language models are, by design, optimized to produce outputs that are fluent, contextually coherent, and conversationally natural. They are not optimized for accuracy. Those two objectives are not the same thing, and conflating them is the operational error at the center of this liability exposure. I watched a version of this pattern play out at EMC in 2011. When the RSA breach compromised 40 million SecurID tokens, the first instinct across the company was to trust the systems that reported everything was fine, because the systems spoke in the language of normalcy. The dashboards looked normal. The logs looked normal. It took weeks for the organization to accept that the indicators of compromise had been hiding in plain sight behind screens designed to project confidence. AI tools present a faster, more personal version of the same trap: the conversational surface projects competence, and the user’s brain fills in the rest. If a first-year associate handed you a memorandum citing three cases, and two of them did not exist, you would question the supervising attorney who signed the filing without reading it. When the memorandum comes from an AI tool, the verification instinct that would catch the junior associate’s fabricated citation does not engage. The screen shows fluent, authoritative text. The brain registers a colleague who has done the work. Matthew Reeves was not a careless attorney. He was an experienced partner reading prose that looked exactly like the research he had reviewed thousands of times before. A 2025 preprint study found that anthropomorphic attributions to AI increased 34 percent in a single year, with humans increasingly viewing AI systems as warm and competent. A study by Angela Duckworth and Lyle Ungar published in the Harvard Business Review in January 2026, surveying nearly 2,500 Generation Z adults, found this cohort uses AI tools even in situations where they have been explicitly told not to. Generation Z is entering law firms as first-year associates, paralegals, and legal operations staff, carrying a baseline level of AI trust categorically higher than any prior cohort. But anthropomorphism does not stop at generational lines. It operates across every experience level, because the tool activates trust regardless of the user’s age, training, or skepticism. Reeves had decades of experience. Butler Snow had written policies. Johnson v. Dunn proved that neither one interrupts the moment when the screen looks right. Why AI Sounds Like It Knows What It Is Talking About An LLM does not retrieve information from a database. It generates text by predicting the statistically most probable next sequence of words given the input and its training data. When an LLM cites a case, it is not looking up that case in Westlaw. It is generating a string of characters that resembles a case citation based on patterns learned from legal documents. The citation may be syntactically perfect, jurisdictionally plausible, and entirely fabricated. Stanford University’s RegLab and Institute for Human-Centered AI tested the purpose-built legal AI research tools sold by LexisNexis and Thomson Reuters. Both providers had publicly claimed their tools avoided hallucinations through proprietary retrieval architecture. The study found that even these tools hallucinate between 17 and 33 percent of the time on legal queries. General-purpose LLMs perform substantially worse, producing incorrect responses between 69 and 88 percent of the time. Put those numbers in a room. A senior associate who stayed up all night researching sits at one end of the conference table. An AI assistant that generated its output in four seconds sits at the other. Both present their conclusions with identical surface confidence. Neither hesitates. Neither qualifies without prompting. The attorney reading the AI’s output has no auditory, visual, or behavioral cue to signal that the confident prose on the screen was generated by a system that fabricates between one in six and one in three of its legal conclusions. Stanford’s team identified a second failure mode more dangerous than outright fabrication: the AI cites a real case that does not actually support the proposition for which it is cited. Johnson v. Dunn documented both types across five hallucinated citations. One cited a real case with a fabricated holding. Another cited a real reporter volume containing an entirely different case. A hallucinated citation can be caught with a Westlaw search. A misgrounded citation requires reading the opinion. That is the verification step anthropomorphism suppresses. **What the American Bar Association Has Already Said** The professional responsibility framework governing this conduct is not ambiguous. Competence. Model Rule 1.1 requires competent representation, including the legal knowledge, skill, thoroughness, and preparation reasonably necessary for the representation. Comment 8 was amended to include a duty to keep abreast of the benefits and risks associated with relevant technology. Forty-two jurisdictions have adopted this requirement. ABA Formal Opinion 512, issued July 29, 2024, concluded that attorneys who use AI tools must understand their limitations, verify AI-generated content for accuracy, and supervise AI use with the same rigor applied to work from junior attorneys. Confidentiality. Model Rule 1.6 requires reasonable efforts to prevent unauthorized disclosure of client information. Formal Opinion 477R established that cloud-based services implicate this obligation. As I documented in the Heppner privilege analysis, Judge Rakoff’s February 2026 ruling in the Southern District of New York confirmed that consumer AI platforms expressly disclaim the confidentiality protections privilege requires. Firms that have issued AI acceptable use policies without verification mechanisms are not managing this risk. They are documenting it. Candor. Model Rule 3.3 prohibits attorneys from making false statements of fact or law to a tribunal. Filing AI-generated citations without verification is a Rule 3.3 violation when the attorney knew or should have known the tool generates fabrications at documented rates. Judge Manasco described the Johnson v. Dunn conduct as reflecting “extreme dereliction of professional responsibility.” The court in Noland v. Land of the Free described it as a violation of “a basic duty counsel owed to his client and the court.” Supervision. Model Rule 5.1 requires partners and supervising attorneys to ensure the firm has measures giving reasonable assurance that all attorneys conform to the Rules. Judge Manasco declined to sanction Butler Snow itself, recognizing that its institutional policies predated the misconduct. But the three individual attorneys who signed the filings received public reprimand, disqualification, and bar referrals regardless of the firm’s compliance posture. Visible steps and effective steps are not the same thing. **Addressing the Counterargument** The standard pushback runs as follows: attorneys are trained professionals with a duty of competence and an ethical obligation to verify their work product. They already have every incentive to catch errors. Anthropomorphism is a consumer phenomenon, not a professional one. The market will self-correct because malpractice exposure creates accountability. Each of those propositions is defensible in isolation. Collectively, they describe the risk management framework that produced Johnson v. Dunn. Butler Snow’s safeguards did not interrupt the moment when an experienced partner accepted AI output because it presented with the confidence and fluency of trusted expertise. Professional training does not override an automatic reflex that operates below the level of deliberate judgment. As I documented in the AI competence analysis, “Word Can’t Even Spell-Check After 40 Years,” the same cognitive abilities that make attorneys exceptional at legal analysis make them overconfident in domains where their expertise does not transfer. A partner who would never advise a client on patent prosecution without understanding the underlying technology will paste case facts into ChatGPT and submit the output to a federal court without understanding how the model generates text. Dunning-Kruger operates precisely in that gap. Remove the AI element entirely. If a law firm hired an outside research service staffed by recent law school graduates who fabricated one in three of their citations, and the firm submitted the research to courts without independent verification, no one would call the resulting sanctions surprising. They would call them inevitable. The technology changes the speed of the failure. It does not change the obligation. Mata v. Avianca in 2023 established the legal standard. Johnson v. Dunn, two years later, demonstrated that the standard is still being violated by sophisticated practitioners at firms that believed they had addressed the problem. Damien Charlotin’s tracking database has identified 979 judicial decisions worldwide addressing AI hallucinations, with 90 percent issued in 2025 alone. The self-correction argument requires a timeline that the documented evidence does not support. Malpractice liability misreads the incentive structure. Liability is a lagging indicator. The error occurs, the filing is made, the proceeding concludes, the claim is filed, and the sanction follows, often years after the original conduct. A profession that relies on malpractice liability as its primary AI governance mechanism has chosen to accept the harm as the price of admission. **The Pattern Beneath the Cases** Every documented sanctions case shares a structural signature: trust transferred from a category where it belongs, a verified human colleague, to a category where it does not, a probabilistic text generator. The transfer happens because the interface is designed to invite it. James Martin Paul in the Southern District of Florida demonstrates the pattern in its extreme form. He continued citing hallucinated authorities after opposing counsel flagged them, after the court issued a show cause order, and in his own response to that show cause order. Judge Leibowitz found the conduct constituted bad faith. Paul’s case represents the floor of the sanctions scale: willful indifference after explicit notice. Matthew Reeves at Butler Snow demonstrates the pattern in its more dangerous form. He was not indifferent. He was experienced, policy-compliant, and supervising. He read the ChatGPT output the same way he had read thousands of associate memoranda, and the prose did not trigger the verification reflex that a junior associate’s work would have. Judge Manasco called it “more than mere recklessness and tantamount to bad faith.” When the fabricated citations came to light, Reeves’s first instinct was to try to skip the show cause hearing. His second was to explain how little personal review he had given the filing. In Noland v. Land of the Free, L.P. (Cal. App. 2d Dist., Sept. 12, 2025), the California Court of Appeal found that nearly all quotations in the attorney’s opening brief were fabricated. The court sanctioned the attorney $10,000 and referred the matter to the State Bar. What made the opinion notable was the question it raised about the other side. The court declined to award attorneys’ fees because opposing counsel had failed to detect or report the fabricated citations. The reasonable attorney standard is being expanded in real time: the obligation to verify now extends to your opponent’s work product. Digital sociologist Julie Albright has described AI tools as offering frictionless engagement: available, non-judgmental, and consistently affirming. In a professional context, that framing describes the identical dynamic operating at Butler Snow and in Paul’s Florida litigation: the tool presents output without hesitation, without qualification, and without any signal that verification should engage. The design suppresses the instinct that would catch the error. **Where This Shows Up in Practice** The liability exposure is not limited to attorneys. It extends to any professional whose work product is mediated by a generative AI tool. The problem is not the tool. It is the human tendency to anthropomorphize the tool, to attribute to it a judgment it does not possess, and to trust its output without verification. The solution is not to ban AI. It is to train professionals to verify AI-generated content with the same rigor they apply to human-generated content.

Originally published on LinkedIn Newsletter: The Technology Blind Spot

Leave a Reply

Discover more from The Technology Blind Spot

Subscribe now to keep reading and get access to the full archive.

Continue reading