Editor’s Note: Two Boston University law professors have given policymakers something the AI slop argument has lacked: a test with edges. Jessica Silbey and Woodrow Hartzog provisionally define slop as machine output produced with little exertion that shifts the burden onto recipients and erodes the domain it lands in. The test turns on effort, imposition and domain degradation rather than on quality, and that is what makes it usable. It separates a clumsy first draft, which is fine, from a polished report nobody will stand behind, which is not.

The timing sharpens the point. Transparency duties under Article 50 of the EU AI Act and California’s AI Transparency Act both became operative Aug. 2, with three more compliance dates through 2028 and a penalty formula that inverts for smaller firms. Meanwhile, the courts, where the counting has actually been done, are what the paper’s policy catalog never reaches. A public database of decisions involving hallucinated material stood at 2,022 when checked Sept. 7.

Watch two developments next. Whether detection tooling hardens into an enforcement layer, carrying its false-positive problem. And whether governance programs start treating unattributable AI output as a retention and defensibility question rather than an HR one.


Content Assessment: Law professors propose a three-part test for what counts as AI slop

Information - 94%
Insight - 95%
Relevance - 95%
Objectivity - 94%
Authority - 94%

94%

Excellent

A short percentage-based assessment of the qualitative benefit expressed as a percentage of positive reception of the recent article from ComplexDiscovery OÜ titled, "Law professors propose a three-part test for what counts as AI slop."


News Analysis – Artificial Intelligence Beat

Law professors propose a three-part test for what counts as AI slop

ComplexDiscovery OÜ Staff

Two Boston University law professors have proposed a policy definition for “AI slop,” a phrase that has spread through courts, journals and workplaces without one. Where that definition draws its line will shape which machine output policy can reach, and which is just work done with a tool.

Jessica Silbey and Woodrow Hartzog make the case in “AI Slop,” an unpublished draft posted to SSRN that carries a Sept. 6, 2026, timestamp in its running header and is labeled a draft on its interior pages, meaning the language quoted here may change before publication. Silbey is a professor of law and Frank R. Kenison Distinguished Scholar in Law at Boston University School of Law; Hartzog is the Andrew R. Randall Professor of Law there. Their provisional definition: AI slop is “any output of a generative probabilistic automated system produced with little exertion that asymmetrically burdens recipients and tends to degrade cultural domains.” That is deliberately narrower than anything a machine had a hand in, and the narrowing is the point.



Three parts, and a spectrum rather than a line

The test has three components, and the authors treat them as dials rather than switches. Negligible exertion asks how much thought, revision and judgment went in. Asymmetrical imposition asks who absorbs the shortfall, and the answer is the recipient, who reads twice, fills gaps and corrects errors the sender declined to catch. Domain degradation asks whether the output, at scale and over time, corrodes the practices of the field it lands in.

Output is “more or less ‘sloppy,'” the authors write, depending on how much of each component is present. That spectrum is the part practitioners should sit with, because compliance frameworks want a binary and this one refuses to supply one.

The framing has a lineage the paper traces. Developer Simon Willison drew the spam analogy in May 2024, writing that “not all AI-generated content is slop” but that “if it’s mindlessly generated and thrust upon someone who didn’t ask for it, slop is the perfect term for it.” Merriam-Webster made slop its 2025 word of the year, defining it as “digital content of low quality that is produced usually in quantity by means of artificial intelligence.” Silbey and Hartzog keep Willison’s emphasis on imposition and drop the dictionary’s emphasis on quality.

Why the test does not turn on quality

That drop matters. A quality-based definition catches every clumsy first draft and misses the polished report nobody stands behind. The authors’ test catches the second and lets the first go.

It also lets useful automation through. The paper names AI-generated earthquake alerts and low-stakes live translation as cases where the first two components are met, effort is negligible and recipients absorb the correction cost, and the third is not, because the domain does not erode and may improve. Automation that supplies information people cannot produce on their own is not slop on this reading, even when it is sometimes wrong.

For anyone drafting an acceptable-use policy, that distinction does the most work. The question is not whether a tool was involved, but whether the sender kept the judgment that makes output worth another person’s time or pushed that work downstream. Regulators, so far, have been asking a different one.

What is already on the books

Governments have mostly reached for disclosure, and the paper catalogs the result. The timing is sharper than the draft lets on: two of the instruments it describes became operative five weeks ago, on the same day.

Transparency obligations under Article 50 of the European Union’s AI Act began applying Aug. 2, 2026. The Digital Omnibus on AI, Regulation (EU) 2026/1744, pushed back other parts of the AI Act but left that date alone. It entered into force July 27 and touched Article 50 only at paragraph 7, on codes of practice, so the operative duties below are the ones now in force.

Providers of systems intended to interact directly with people must design and develop them so the person is informed they are interacting with an AI system, unless that is obvious to a reasonably well-informed observer. Providers of systems generating synthetic audio, image, video or text must mark the outputs in a machine-readable format and make them detectable as artificially generated or manipulated. That second duty falls away where the system performs an assistive function for standard editing, or does not substantially alter the input data the deployer provided or its semantics. That carve-out may matter inside review workflows. Paragraph 4 then carries two duties for deployers. A deployer generating a deep fake must disclose that the content was artificially generated or manipulated, and where the work is evidently artistic, creative, satirical or fictional the obligation narrows to disclosing that such content exists. A deployer publishing AI-generated text to inform the public on matters of public interest must disclose that too, unless the text has undergone human review and a person or organization holds editorial responsibility for it.

Every one of those duties carries a law-enforcement exception, and the one attached to direct-interaction systems is the narrowest. It is subject to safeguards for third-party rights, and it does not reach systems the public can use to report a crime. The omnibus wrote in one transitional break: systems already on the market before Aug. 2 have until Dec. 2, 2026, to meet the machine-readable marking requirement.

Penalties sit in Article 99, and the formula runs in two directions. A breach of Article 50 draws up to 15 million euros or, where the offender is an undertaking, up to 3 percent of total worldwide annual turnover for the preceding financial year, whichever is higher. For small and medium-sized enterprises, including start-ups, the same two figures apply with the comparison reversed: whichever is lower. The omnibus added a paragraph 6a extending that reversal to small mid-cap enterprises.

California’s AI Transparency Act, Senate Bill 942, became operative the same day. Assembly Bill 853, signed Oct. 13, 2025, as Chapter 674, moved it off its original Jan. 1, 2026, date. Section 22757.6 of the Business and Professions Code sets the operative date in one sentence: the chapter becomes operative Aug. 2, 2026. The act reaches generative systems with over 1 million monthly visitors or users that are publicly accessible in California.

For image, video and audio content, and combinations of those, it requires a free detection tool, an optional visible disclosure, and a latent disclosure that, so far as is technically feasible and reasonable, conveys the provider’s name, the system name and version number, the time and date the content was created or altered, and a unique identifier. That disclosure also has to be detectable by the provider’s own tool, and permanent or extraordinarily difficult to remove to the extent that is technically feasible. Text sits outside that scope, a direct difference from Article 50, where synthetic text is expressly covered. Under Section 22757.3(c), a covered provider that finds a licensee has modified its system so it no longer includes the required latent disclosure must revoke that license within 96 hours. Civil penalties run $5,000 per violation, each day counting as a discrete violation, enforceable by the attorney general, a city attorney or county counsel.

Two later dates follow from AB 853. Large online platforms and generative-system hosting services take on provenance obligations Jan. 1, 2027, and capture-device manufacturers take on latent-disclosure obligations Jan. 1, 2028, a category the statute defines to reach still and video cameras, phones with built-in cameras or microphones, and voice recorders. Compliance calendars that do not carry Aug. 2, Dec. 2 and the two January dates are behind.

Platform policy is messier and, in the authors’ assessment, weaker. Spotify said Aug. 11 it would begin applying an “AI Persona” badge in mid-September to artist identities that may not represent a real person, a label about who is presenting the music rather than how it was made, and it will keep those artists out of editorial and algorithmic recommendations by default. The paper reports that Bandcamp bars music generated wholly or in substantial part by AI, and that Medium keeps AI-generated work from behind its paywall while leaving it readable free, which the authors call half-hearted. Silbey and Hartzog see industry and government converging on four moves: “limited bans, de-privileging, disclosure requirements, and provenance tracing.” They argue disclosure-led approaches underestimate how thoroughly sheer volume shapes what a reader can find.

Detection tools invite a policing problem

Research publishing has already felt that volume, and it has the most developed rules of the paper’s three domains, plus the numbers to argue for them. Editors at Organization Science reported in April that the journal took 42 percent more submissions in January 2024 through November 2025 than in January 2022 through November 2023, a baseline window that itself straddles ChatGPT’s November 2022 release. They put the pandemic-era rise at 20 percent, half that size, and while they allow that thinner backlogs or rising productivity could account for some of it, they read their own figures as pointing to AI use. Editor-in-Chief Lamar Pierce and three colleagues scored submission abstracts with the Pangram detector. Abstracts in the 70 percent and above band were desk-rejected 69.6 percent of the time, against 43.7 percent in the 0 to 15 percent band. Editors reject on quality and fit rather than on a detection score, they took care to say.

Journals and law reviews have answered with disclosure rules and prohibitions. The paper reports that the California Law Review requires editor-approved disclosure and bars pieces written entirely or substantially with AI, and that the Virginia Law Review makes authors complete a form covering AI use to support factual assertions, to advance legal claims, and to generate text.

Silbey and Hartzog commend those rules, then warn about the enforcement layer going up beside them. By their account detection has improved lately while still struggling with second drafts, polishing and short passages. Others put it harder. MIT Sloan Teaching & Learning Technologies publishes instructor guidance under the title “AI Detectors Don’t Work. Here’s What to Do Instead.” That guidance says the tools carry high error rates that can lead instructors to falsely accuse students, and it points to OpenAI’s withdrawal of its own detector for poor accuracy. At scale, even a small false-positive rate produces writers defending themselves against an accusation the evidence cannot support. Bolting a policing layer onto institutions that ran on presumed sincerity, the authors argue, compounds the damage rather than repairing it.

That warning should land with anyone weighing detection tooling for internal investigations or review quality control. A tool that is right 98 percent of the time across a million documents misclassifies 20,000 of them, and the people behind those documents are the ones who have to answer for it.

One security team stopped absorbing the cost

Asymmetrical imposition is easiest to see where somebody finally refused to carry it. The curl project ended its bug bounty Jan. 31, 2026, and founder Daniel Stenberg explained why on his own blog five days earlier: the share of submissions confirmed as real vulnerabilities fell below 5 percent in 2025, down from above 15 percent previously, which he attributed to a surge in AI-generated reports that began in late 2024 and accelerated through 2025. The program had confirmed 87 vulnerabilities and paid over $100,000 across its life. Reporting moved to GitHub’s private vulnerability channel and a security mailbox, with no reward attached. Stenberg’s post does not separate submissions that were AI-generated from those that were merely poor, so the AI share is his assessment rather than a measured figure.

Read against the paper’s test, that sequence is the whole argument in miniature. Effort collapsed on the sending side, the triage cost landed on a volunteer security team, and the domain response was to shut a door that had been open for years. Security leaders running vulnerability disclosure or bug bounty programs may want to watch their own confirmation rates, since that ratio can move well before a queue becomes unmanageable.

The domain the paper leaves out

Courts get the same treatment as security teams, which is to say almost none. The paper’s abstract puts legal tribunals alongside peer-reviewed journals as institutions buckling under low-reliability submissions, and its opening paragraph repeats the pairing. Journals then get a full policy section. Courts get nothing. A search of all 70 pages, footnotes included, finds the word tribunal only in that sentence and its echo in the abstract, the word court only inside a single footnote citation, and the word judicial only in the title of a cited article. Sanction, docket, litigation and Rule 11 do not appear anywhere.

The gap is worth naming, because the legal domain is where the counting has actually been done. Damien Charlotin’s AI Hallucination Cases database listed 2,022 decisions when checked Sept. 7, 2026, carrying a last-updated date of Sept. 5, with 1,379 in the United States, 217 in Canada, 110 in Australia, 57 in Israel and 41 in Brazil, and a long tail of national entries running down to single cases. The page states no jurisdiction total. The tracker excludes bare allegations, counting matters where a court or tribunal found or implied reliance on hallucinated material, though its own note concedes that entries where AI use was alleged but not confirmed involve a judgment call. On those figures, close to 68 percent of the recorded worldwide total sits in U.S. courts. The counts come from that tracker alone and were not confirmed against a second register.

The judiciary’s own guidance runs the other direction. The Administrative Office of the U.S. Courts broadcast interim AI guidance judiciary-wide July 31, 2025, and its director, Judge Robert J. Conrad Jr., described it to Senate Judiciary Chairman Chuck Grassley in an October 2025 letter. The guidance recommends, rather than requires, that users review and independently verify AI-generated output, and it reminds judges and staff that they remain accountable for work done with AI. A judiciary spokesperson declined to give FedScoop a copy when it reported on the guidance.

That covers the people on the bench. For what arrives from counsel, Conrad’s letter points to no judiciary-wide AI rule at all, but to instruments already on the books: Rule 11 of the Federal Rules of Civil Procedure, which lets judges sanction attorneys and self-represented parties over offending filings, and Canon 3B(6) of the Code of Conduct for United States Judges, on attorney misconduct. That is the quiet irony of the paper’s gap. Rule 11 is one of the terms appearing nowhere across its 70 pages.

Scholars have not settled the definition

None of which makes the definition wrong. Silbey and Hartzog are entering a live argument rather than closing one. In June 2026, Sachita Nishal of Northwestern University, Marijn Sax of the University of Amsterdam and Kimon Kieslich of the University of Hohenheim posted to arXiv a critique of a different slop framework, arguing that definitions built on individual preference satisfaction foreclose the social and political questions that matter most. Their target is a preference-based account rather than the effort-and-imposition account offered here, but it establishes that the conceptual work is contested.

The workplace evidence is unsettled too, and the $9 million figure needs care. BetterUp Labs and the Stanford Social Media Lab surveyed 1,150 full-time U.S. desk workers in September 2025 and reported that 40 percent had received workslop in the previous month, that resolving an incident took about two hours, and that the cost ran $186 per employee per month. The $9 million annual figure attached to a 10,000-person company follows only if that monthly cost is applied to the 40 percent who reported receiving workslop, about 4,000 people. Applied across a full 10,000-person headcount, the same inputs produce roughly $22.3 million. BetterUp’s page does not explain the calculation, so anyone putting the $9 million in a board deck should say which population it describes.

Underneath all of it sits a governance point information professionals will recognize. Workslop is not only a morale problem. The plausible-looking report nobody will vouch for still enters the record, still gets retained, still gets collected, and still has to be defended by somebody who did not write it. Organizations treating AI acceptable-use as an HR memo rather than a records policy are storing the problem instead of solving it.

Silbey and Hartzog close where policy arguments usually open, arguing the window for treating AI slop as nascent is closing and that lawmakers should act while it is still open. Their bet is that naming the thing precisely is what makes it governable.

So here is the question worth putting to your own organization: if someone asked tomorrow how much of your AI-assisted work product a named human is prepared to stand behind, could anyone answer?



News sources



Assisted by GAI and LLM Technologies

Additional reading

Source: ComplexDiscovery OÜ

ComplexDiscovery’s mission is to enable clarity for complex decisions by providing independent, data‑driven reporting, research, and commentary that make digital risk, legal technology, and regulatory change more understandable for practitioners, policymakers, and business leaders.

 

Have a Request?

If you have information or offering requests that you would like to ask us about, please let us know, and we will make our response to you a priority.

ComplexDiscovery OÜ is an independent digital publication and research organization based in Tallinn, Estonia. ComplexDiscovery covers cybersecurity, data privacy, regulatory compliance, and eDiscovery, with reporting that connects legal and business technology developments—including high-growth startup trends—to international business, policy, and global security dynamics. Focusing on technology and risk issues shaped by cross-border regulation and geopolitical complexity, ComplexDiscovery delivers editorial coverage, original analysis, and curated briefings for a global audience of legal, compliance, security, and technology professionals. Learn more at ComplexDiscovery.com.

 

Generative Artificial Intelligence and Large Language Model Use

ComplexDiscovery OÜ recognizes the value of GAI and LLM tools in streamlining content creation processes and enhancing the overall quality of its research, writing, and editing efforts. To this end, ComplexDiscovery OÜ regularly employs GAI tools, including ChatGPT, Claude, Gemini, Grammarly, Midjourney, and Perplexity, to assist, augment, and accelerate the development and publication of both new and revised content in posts and pages published (initiated in late 2022).