THE TECHNOLOGY BLIND SPOT
In 2022, Jason Allen won first prize in the digital arts division of the Colorado State Fair. The award was three hundred dollars and a blue ribbon. The work was “Theatre D’opera Spatial,” a baroque opera house filled with Victorian figures in space helmets. Allen produced it by typing 624 prompts into Midjourney, then cleaning the result in Photoshop and upscaling it with Gigapixel AI. The fair’s judges later said the use of AI would not have changed the outcome. Four years later, that three-hundred-dollar ribbon sits at the center of a federal lawsuit that the Copyright Office is winning.
Allen filed to register the work. The Office refused. He asked the Office to reconsider, twice, and it refused both times. In September 2024 he sued in federal court in Colorado, where the case is still pending on his motion for summary judgment. On March 2, 2026, the Supreme Court declined to hear the parallel case of Stephen Thaler, leaving the human authorship requirement intact for work generated by a machine.
Most analysis of these cases asks a forward question: how many prompts, and how much human editing, it takes to become an “author” under the Copyright Act. That is the wrong question. The right one runs in the opposite direction, and answering it changes how a lawyer should think about copyright, AI products, and the value of anything a generative model produces.
The right question is mechanical. What is the model actually doing when it produces “Theatre D’opera Spatial”? Once the answer is honest, the legal outcome stops looking like a policy choice. It starts looking like arithmetic.
What the model is doing
Set the marketing language aside. Diffusion models like Midjourney learn a probability distribution over the space of possible images by examining millions of training examples. To generate an image, the model starts from random noise and removes it step by step, steered by the prompt, toward the dense regions of that learned distribution. Transformer language models do the same for text. At each step the model samples the next token from a probability distribution, weighted by the patterns it absorbed in training.
This is sampling. It is statistical inference. In the most literal mathematical sense, it is regression to the mean.
Engineers built the model to land in the high-likelihood region of its training distribution, conditioned on the prompt. That is what it means for output to “look right.” A face that fails to regress to the mean grows a sixth finger and a third eye. A sentence that fails to regress to the mean reads as gibberish. Coherence and aesthetic appeal are not features bolted onto the sampling process. They are the signature of successful regression to the mean.
When Allen typed his 624 prompts, he was not directing an artist. He was running queries against a statistical engine that returned the average of what sat near his query. The model chose the hair, the bone structure, the helmet design, the lighting, the composition, the brushwork, the palette. Sampling from a distribution built out of human creative work produced every expressive element in the final image. Allen’s role was to constrain the query until the output landed somewhere he liked.
That distinction is the whole case. The better the model gets at producing coherent, on-prompt work, the more of the expressive arrangement traces back to the learned distribution and the less of it traces back to the person at the keyboard.
The doctrine that decides it
American copyright law turns on two requirements that most AI commentary blurs together. The first is human authorship. The second is originality. The Allen and Thaler denials rest on the first, not the second, and the difference matters more than it looks.
Originality sets a famously low bar. In Feist Publications, Inc. v. Rural Telephone Service Co., 499 U.S. 340 (1991), the Supreme Court held that copyright requires only a “modicum of creativity,” not labor. Rural had compiled a phone directory through exhausting data collection, and the Court still ruled it uncopyrightable, because an alphabetical list of subscribers contains no creative choice no matter how much sweat produced it. Feist killed the “sweat of the brow” theory. A human can produce an unprotectable work by exercising too little creativity.
Authorship is a separate prong. It asks who made the expressive choices, and it requires that the answer be a person. This is where the machine cases live. The Copyright Office did not deny Allen because his image lacked creativity. It denied him because the expressive choices were executed by Midjourney, and a person who issues prompts has not, in the Office’s words, controlled the elements that a machine determined.
Regression to the mean explains why that line falls where it does. The expressive content of a pure prompt-to-output image is the recombined center of the training data. The prompt narrows the search. It does not supply the brushwork. A critic will answer that selection and arrangement can themselves be authorship, and the critic is right. The Office agrees: where a human arranges, edits, or contributes perceptible expression, that contribution is protectable, on terms the Office compares to a derivative work. The vulnerability in the regression argument is real and worth stating plainly. Originality is cheap, and a novel arrangement of average parts can clear the Feist bar. The denials do not turn on the output being statistically ordinary. They turn on the absence of a human who made the arrangement. The regression point is not that the output is unoriginal. It is that the arrangement Allen wants to claim was performed by the model, sampled from everyone else’s work, and a prompt is a request, not a brushstroke.
The Office’s awkward framing
The Copyright and Artificial Intelligence report the Office published on January 29, 2025, states the position cleanly. Outputs are copyrightable only where a human determines the expressive elements. Mere prompts do not qualify, because the user does not control what the machine produces. Register of Copyrights Shira Perlmutter put the principle directly: extending protection to material whose expressive elements are determined by a machine would undermine the constitutional goals of copyright rather than serve them.
That framing is workable and evasive at the same time. A model does not “determine” anything in the sense the word implies. It performs matrix multiplication. It samples from a distribution. There is no deliberation in a forward pass and no intent in a denoising step. Saying the machine determined the expression treats statistical sampling as if it were decision-making. The more honest statement is that no entity made expressive choices in the copyright sense. The choices that produced the work were the aggregate choices of every human whose output landed in the training set, recombined by sampling. The expression lives in the training data, not in the prompt and not in the weights.
The Office does not say this, because saying it points somewhere uncomfortable. If the expression in any specific output traces to identifiable training works, the people who made those works have a claim, and the vendors that built businesses on mass aggregation of that work have a problem. Treating the model as a black box that “determines” expression protects the industry from output-copyright claims and training-data claims at the same time. It is a fiction with a job to do.
Why the contradiction has not surfaced yet
That fiction is unstable, and the obvious place to test it is the courts. The interesting fact is that the courts have refused to test it.
In June 2025, two federal judges in the Northern District of California ruled within two days of each other. In Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA (N.D. Cal. June 23, 2025), Judge William Alsup held that training a model on lawfully acquired books was “exceedingly transformative” fair use, while pirating the source copies was not. Two days later, in Kadrey v. Meta Platforms, Inc., No. 23-cv-03417-VC, 2025 WL 1752484 (N.D. Cal. June 25, 2025), Judge Vince Chhabria reached the same fair-use result on training, on the narrower ground that the authors failed to prove market harm. Both opinions did something deliberate. They confined themselves to inputs. Alsup noted in plain terms that the case would look different if the outputs were infringing, and the plaintiffs had not alleged that. Anthropic later settled the piracy claims, in the largest copyright settlement on record, without anyone reaching the output question.
So the seam runs exactly where no court has cut. The same output cannot be authored by the machine for one doctrinal purpose and derived from training data for another. A vendor cannot tell the Copyright Office that the model determined the expression, then tell a district court that the model merely learned uncopyrightable patterns. The classification problem that runs through every emerging corner of AI law is here too: the same artifact cannot be a tool when that helps and an author when that helps.
The output cases are where the contradiction gets forced, and they are pending now. Disney and Universal sued Midjourney in June 2025 in the Central District of California, alleging that the service generates and distributes images of their characters on demand, a “virtual vending machine” for infringing output. The New York Times case against OpenAI and the artists’ suit against Stability AI raise the same theory from the output side. If any one of them holds that a specific output substantially derives from identifiable training works, the “the machine determined the expression” framework does not survive contact. Watch that seam over the next eighteen months.
For a managing partner, this is not abstract. If your firm’s deliverables or your clients’ products incorporate pure AI output, the most common contract clause in the building may be assigning rights that do not exist. Pull one engagement letter or one client vendor agreement on Thursday and find the intellectual property clause. If it warrants that AI-assisted deliverables are protectable, or assigns “all intellectual property in deliverables” without distinguishing the AI-generated layer, the warranty covers ground the Copyright Office and two federal courts have now mapped as empty.
Recommended by LinkedIn
Run the same logic into the brief itself. An attorney who prompts a model to draft a motion, skims it, signs it, and files it has authored none of the expressive content in the sense the Copyright Office means. The structure, the phrasing, and the cadence all regressed to the mean of every brief in the training set. That collides with the billable hour. ABA Formal Opinion 512 is blunt: a lawyer who bills hourly must bill for actual time spent and not more, so the work the model collapsed from six hours to ninety seconds cannot be billed as six hours of drafting. What stays billable is the part that does not regress to the mean, the reading, the verification, and the judgment about whether the citations exist and the argument holds. The drafting belonged to no one and is worth nothing on the invoice. Judgment is the only thing the client is paying for. The signature is not a claim of authorship. It is an assumption of responsibility. Rule 11 makes the lawyer who signs answerable for every word a court later finds fabricated, which is how the attorneys in Johnson v. Dunn drew sanctions for citations a model invented and no one checked. The same human the Copyright Office will not call an author is the one the court sanctions by name. A tool for the copyright question, an author for none of it, a defendant for all of it. [See When Attorneys Stop Checking AI’s Work, The Technology Blind Spot (2025).]
What this means for the products your clients build
The same mechanics that sink the copyright claim reshape the market your clients operate in, and the firms buying legal AI should read it the same way.
If a product sells pure AI output as proprietary work product, three things hold at once. The output is probably not copyrightable, so neither the vendor nor the customer owns it, and a competitor who copies it faces no infringement claim. The output regresses to the mean, which means it converges across vendors, because frontier models trained on overlapping corpora toward similar objectives produce similar results for similar prompts. And a product that is a thin layer over a frontier model has a moat that shrinks on a schedule set by frontier vendor research, not by its own roadmap. [See Your Legal AI Vendor Is Selling You Cheap Wins, The Technology Blind Spot (2026).]
Defensibility has to come from layers that do not regress to the mean, and there are only four candidates. The first is proprietary data that does not exist in the training distribution, such as a litigation-analytics firm sitting on a private corpus of sealed settlement outcomes. The second is fine-tuning that shifts the output distribution toward a domain, such as a medical-coding vendor that trains on its own adjudicated claims. The third is workflow integration that binds model output to external systems and human judgment the model cannot reproduce alone, such as a contract-review tool that routes each clause into a firm’s own playbook and a named reviewer’s sign-off. The fourth is encoded domain methodology, the evaluation harnesses and verification layers that took a specialist years to build and that the model does not contain. A product whose answer to “what is your moat” is “a better prompt” is describing the thing the mean reaches first.
The unresolved question
Two scenarios will tell whether this holds.
Start with the output litigation. If Disney’s case against Midjourney produces a finding that named outputs derive from identifiable training works, the vendor’s authorship story and its training-data defense collide in a way no settlement can paper over. The contradiction stops being theoretical and acquires a docket number.
The second is capability. The next generation of frontier models trains on more specialized content and reasons better over technical material. Some of the methodology a vendor thought was defensible gets absorbed into the base layer, and the mean shifts toward territory that used to be safe. The honest plan assumes the base capability keeps eating defensible ground, and requires the data flywheel and the workflow integration to compound faster than the frontier vendors generalize. That is the actual race.
Where this leaves Allen
Allen will likely lose in Colorado on the same statutory ground that defeated Thaler. The court will reach the result through human authorship and never engage the mechanism underneath it. The case that confronts the mechanism honestly is still coming, and it will ask the question the doctrine keeps deferring. If the machine made no expressive choices in the copyright sense, and the prompter made none either, and the only parties with an expressive claim are the diffuse community whose work fed the training set and cannot be individually identified, then the pure output belongs to no one. Public domain by default.
That is the finding behind the three-hundred-dollar ribbon. The value of generative AI does not live in the output, because the output is no one’s to own. It lives in the layers built around an undefendable center. Regression to the mean is not only a copyright problem. It is the central economic fact of the technology, and the contract on Catherine’s desk is already written as if it were not.
About the Author
JD Morris is Co-Founder and COO of LexAxiom, an Agentic AI platform for the business of law. Over a 25-year career, he has built and scaled enterprise technology products across Dell, EMC, VMware, and Cisco, including the first exabyte eDiscovery platform. He holds dual MBAs from Columbia Business School (Finance) and UC Berkeley Haas (Marketing), a Master of Legal Studies in Cybersecurity Law from Texas A&M, and a Master of Engineering from George Washington University. He writes The Technology Blind Spot on the intersection of emerging technology and law. Connect with him on LinkedIn at www.linkedin.com/in/jdavidmorris, on X at @JDMorris_LTech, or on Bluesky at @JDMorris-ltech.bsky.social.
References
1. Allen v. Perlmutter, No. 1:24-cv-02665-SKC-KAS (D. Colo. filed Sept. 26, 2024).
2. Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA (N.D. Cal. June 23, 2025).
3. Disney Enters., Inc. v. Midjourney, Inc. (C.D. Cal. filed June 11, 2025).
4. Feist Publ’ns, Inc. v. Rural Tel. Serv. Co., 499 U.S. 340 (1991).
5. Johnson v. Dunn, 792 F. Supp. 3d 1241 (N.D. Ala. 2025).
6. Kadrey v. Meta Platforms, Inc., No. 23-cv-03417-VC, 2025 WL 1752484 (N.D. Cal. June 25, 2025).
7. Thaler v. Perlmutter, No. 23-5233 (D.C. Cir. Mar. 18, 2025), cert. denied, No. 25-449 (U.S. Mar. 2, 2026).
8. 17 U.S.C. § 102 (2018).
9. Fed. R. Civ. P. 11(b).
10. Model Rules of Pro. Conduct r. 1.5 (Am. Bar Ass’n 2024).
11. ABA Comm. on Ethics & Pro. Resp., Formal Op. 512 (2024).
12. U.S. Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability (Jan. 2025), https://www.copyright.gov/ai/.
Originally published on LinkedIn Newsletter — The Technology Blind Spot
