Editor’s Note: Regulators moved the deadline. Litigation did not. When the EU AI Act reached general application on Aug. 2, 2026, the record-keeping duties most practitioners associate with it, automatic event logging under Article 12 and the six-month retention floor under Article 26(6), did not take effect. The AI omnibus in force July 27, 2026 pushed those Chapter III obligations to Dec. 2, 2027 and Aug. 2, 2028. The questions they were written to answer arrived anyway, in 2026 rulings on AI prompts, privilege and generative review.

This tutorial gives ComplexDiscovery OÜ’s three audiences a way to work the problem before it is theirs. Discovery teams face preservation and production decisions about agent traces that sit on two sets of retention clocks and rarely appear on a custodian chart. Information governance teams set the instrumentation and retention policies that decide, years in advance, which questions a record can answer. Security teams own the telemetry everyone else will treat as evidence.

Three things are worth watching this autumn: the Advisory Committee’s return to proposed Rule 707, the outcome and publication of ISO/IEC 24970, and the first order that squarely tests authentication of an agent log.


Content Assessment: When the agent becomes a witness: an Oxford-style tutorial on AI evidence, accountability and discovery

Information - 93%
Insight - 93%
Relevance - 94%
Objectivity - 92%
Authority - 91%

93%

Excellent

A short percentage-based assessment of the qualitative benefit expressed as a percentage of positive reception of the recent article from ComplexDiscovery OÜ titled, "When the agent becomes a witness: an Oxford-style tutorial on AI evidence, accountability and discovery."


News Analysis – Artificial Intelligence Beat

When the agent becomes a witness: an Oxford-style tutorial on AI evidence, accountability and discovery

ComplexDiscovery Staff

An AI agent that searches, classifies, recommends and acts leaves a trail behind it. Whether that trail can establish what the agent did, months later, in front of an examiner paid to doubt it, is a different question, and it is the one ComplexDiscovery OÜ has built its newest tutorial around.

The Oxford tutorial, a form built on preparation, argument and close questioning, has served the publication before as a way to slow a fast-moving field down. This installment points the method at a narrower target: the evidentiary standing of machine conduct.

What it is

The result, “The ComplexDiscovery Tutorial: When the Agent Becomes a Witness,” takes the Oxford tutorial (a meeting in which a student prepares written work and defends a contestable claim under sustained challenge) and points it at the records that agentic systems leave behind. The document is built around 21 contestable propositions across three academic terms, with a fully worked example for each term, a model essay paired with the interrogation that follows, so readers can watch the method at work rather than take its value on faith.

One question runs through all 21. Organizations increasingly rely on AI agents to search, classify, recommend and act. When an agent’s conduct becomes relevant to litigation, an investigation or an audit, can its logs establish what actually happened?

Why now, and a correction worth making

The timing is not incidental, though the reason is widely misstated. The EU AI Act reached a milestone on Aug. 2, 2026, when the Chapter IV transparency duties in Article 50 and the Chapter IX right to explanation in Article 86 became applicable, and the enforcement framework took hold. That framework is split three ways: the AI Office enforces against providers of general-purpose AI models, against AI systems developed by the same provider as the underlying model or by a provider within the same business group, and against AI systems integrated into the very large online platforms and very large online search engines designated under the Digital Services Act; national competent authorities take other AI systems; and the European Data Protection Supervisor supervises AI used by EU institutions.

The provisions practitioners most often associate with the record did not arrive with them. Article 12 (automatic recording of events over the lifetime of a high-risk system), Article 19 (a provider’s duty to keep those logs) and Article 26(6) (a deployer’s duty to keep them “for a period appropriate to the intended purpose … of at least six months”) all sit in Chapter III. Regulation (EU) 2026/1744, the AI omnibus that entered into force July 27, 2026, moved those Chapter III obligations to Dec. 2, 2027 for the high-risk systems listed in Annex III and to Aug. 2, 2028 for those embedded in regulated products.

That gap is the tutorial’s opening. A deferred obligation is not a deferred question. Litigation, investigations and audits run on their own calendar, and the first party to ask what an agent did on a particular afternoon will not wait until December 2027 to ask it. Meanwhile the standards work has moved: ISO/IEC 24970, a standard devoted to AI system logging, entered its final draft ballot on Aug. 28, 2026, two days before this publication.

The form

The tutorial, as commonly associated with Oxford teaching, is less a lecture than a disciplined conversation. The university describes a tutorial as normally involving two or three students and their tutor, though some are held one to one. The document adopts the one-to-one form as its working model: a weekly argument in which a student prepares written work, defends it in discussion and absorbs sustained challenge from a tutor under no obligation to be generous. For this document the essay is kept short, roughly 1,500 to 2,000 words, and is expected to take a position rather than survey a field. Oxford itself sets no standard length; expectations vary by tutor and by discipline.

The essay is not the point. What survives the hour after it is the point. A proposition that arrives intact has usually been narrowed, and the narrowing is the work.

A method, then three terms

Oxford divides its academic year into three terms, and the tutorial borrows the names: Michaelmas in the autumn, Hilary in the winter, Trinity in the spring. The rest of the scheme is the document’s own adaptation rather than Oxford practice. Collections, at Oxford, are the examinations a college sets its own students, normally at the start of a term to test the one before. Here each term runs seven weekly propositions, where an Oxford Full Term runs eight, and closes with collections rather than opening on them: a final exercise that asks a reader to argue earlier positions against themselves. The propositions are written to be defensible in either direction, which is what makes them usable.

Michaelmas: the record

The first term asks what an agent actually leaves behind, and whether it amounts to evidence or only to telemetry.

Week one opens with the claim that an AI agent’s log is not a neutral record but the first draft of its operator’s defense, training the reader to see instrumentation as authorship, against readings that pair Federal Rule of Evidence 803(6)(E) and Lorraine v. Markel American Insurance Co. with Lorraine Daston and Peter Galison on mechanical objectivity and Cornelia Vismann on the file as an instrument of authority rather than a report of one.

Week two proposes that observability is an engineering practice rather than an evidentiary one, and that the difference is retention, set against the AI Act’s Articles 12 and 19, the six-month floor in Article 26(6), the draft ISO/IEC 24970 logging standard and the monitoring subcategories of the NIST AI Risk Management Framework.

Week three argues that a prompt is a document and an embedding is not, pressing the reader on where the boundary of discoverable electronically stored information falls, against Rule 34(b)(2)(E), Williams v. Sprint/United Management Co. on metadata, and the 2025 and 2026 orders in Concord Music Group v. Anthropic and Conservation Law Foundation v. Shell Oil Co. treating prompts as inputs a court can reach, the second of them stayed and under district court review.

Week four holds that model version is metadata, and that unversioned output is unauthenticated output, tested against Rule 901(b)(9), the authentication discussion in Lorraine, and the validation framework Paul Grimm, Maura Grossman and Gordon Cormack set out in “Artificial Intelligence as Evidence.”

Week five claims that retrieval is the part of the record that disappears first, pairing Rule 37(e) and Zubulake v. UBS Warburg LLC with the preservation history in the consolidated OpenAI copyright litigation, where a broad going-forward preservation order gave way to a narrower, negotiated arrangement.

Week six takes up the comfort that a machine’s output is not hearsay, and argues the comfort is smaller than it sounds, against Rule 801(a)’s requirement of a person, United States v. Lizarraga-Tirado, United States v. Washington and the still-unsettled proposed Rule 707.

Week seven closes the term on the claim that the log which would exonerate you is written by the system you are defending, against Helen Nissenbaum on the barriers to accountability in computerized systems, Danielle Citron on technological due process and Joshua Kroll and colleagues on accountability built in before the fact rather than inspected after it.

The term’s collections then ask the reader to defend, in a single sitting, both that logging is an evidentiary practice and that it is not.

Hilary: the responsibility

The second term moves from the artifact to the actor: delegated authority, human approval, privilege, vendor opacity and what happens when several agents contribute to one decision.

Week one proposes that delegation to an agent transfers the work but not the duty, against Upjohn Co. v. United States, ABA Formal Opinion 512 on lawyer competence and generative AI, and Moffatt v. Air Canada, in which a Canadian tribunal declined to treat a chatbot as a separate legal entity from the company that ran it.

Week two argues that privilege cannot be delegated to a system whose conduct cannot be reconstructed, set against United States v. Kovel and the 2026 split between United States v. Heppner and Warner v. Gilbarco, Inc., decided a week apart and reaching opposite results.

Week three claims that human-in-the-loop is an allocation of blame before it is a control, against Lisanne Bainbridge on the ironies of automation, Madeleine Clare Elish on moral crumple zones and Article 14 of the AI Act on what oversight must actually enable.

Week four holds that an approval which cannot be refused is not an approval, pairing Diane Vaughan on the normalization of deviance with Onora O’Neill on intelligent accountability and the override and interruption requirements in Article 14(4).

Week five argues that vendor opacity is a choice the buyer makes at procurement rather than a condition the buyer inherits, against Deirdre Mulligan and Kenneth Bamberger on procurement as policy, Rebecca Wexler on trade secrets in criminal proceedings, State v. Loomis and State v. Pickett.

Week six proposes that when several agents contribute to one decision, no one is responsible unless someone was already, against Nissenbaum’s problem of many hands, Charles Perrow on tightly coupled systems and the agent-visibility mechanisms Alan Chan and colleagues catalogue.

Week seven closes the term on the claim that an explanation the deployer cannot give is one the provider owes, against Articles 13 and 86 of the AI Act, Andrew Selbst and Solon Barocas on the appeal of explainable machines, and Lilian Edwards and Michael Veale on the limits of a right to an explanation.

The term’s collections ask the reader to argue that human oversight is the primary safeguard, and then that it is primarily a liability arrangement.

Trinity: the contest

The third term assumes the dispute has arrived: preservation, collection, authentication, proportionality and testimony.

Week one proposes that the duty to preserve attaches to a system rather than to a custodian, against Rule 37(e), the Zubulake line and the split-record problem raised when an appellate court treats the user rather than the provider as the actor, leaving agent traces on two sets of retention clocks.

Week two argues that agent logs held by a vendor are within your control until you have to prove they are not, against Rule 34, the Sedona Conference commentary on possession, custody and control, and the Stored Communications Act constraints on reaching an operator’s records directly.

Week three claims that authentication of machine output is a proportionality question dressed as a foundation question, against Rules 901(b)(9) and 902(13), Lorraine, and the Grimm, Grossman and Cormack validation framework.

Week four holds that reconstruction is not replay, and that a system which can only replay cannot be examined, pairing Michael Polanyi on tacit knowledge with Selbst and Barocas on inscrutability and the draft ISO/IEC 24970 logging model.

Week five proposes that proportionality protects the party with the worse records, against the six factors in Rule 26(b)(1), Sedona Principle 6, Livingston v. City of Chicago and In re Valsartan, Losartan, and Irbesartan Products Liability Litigation.

Week six argues that generative review is technology-assisted review, and that the answer is either reassuring or evasive, against Da Silva Moore v. Publicis Groupe, Rio Tinto PLC v. Vale S.A. and the June 2026 order in Schulte v. LinkedIn Corp., which analyzed a generative review workflow as technology-assisted review and refused to order further validation because no specific production deficiency had been shown.

Week seven closes the tutorial on the claim that no organization should deploy an AI agent it could not explain under oath, against proposed Rule 707, Loomis and Pickett on proprietary systems, O’Neill on what accountability requires and the Sedona Conference decision tree for evaluating AI-generated evidence.

The term’s collections ask the reader to argue that the existing rules are adequate to agentic AI, and then that they are not.

The form in motion

Three worked examples follow the curriculum, one for each term. Each pairs a model essay with the tutor’s interrogation that follows it, and each is an illustrative demonstration of the method, a written simulation of the exchange, rather than a transcript of a live tutorial between named people. The cases, rules, standards and scholars the essays cite are real and were checked against primary or authoritative sources. The student and the tutor are composite, illustrative voices.

The first, from Michaelmas, defends the claim that an agent’s log is the first draft of its operator’s defense, and then loses most of it. The tutor grants the observation about design and takes away the imputation of motive, and what remains is a narrower and more usable proposition: the log is a designed artifact, and discovery has to reach the design.

The second, from Hilary, takes on privilege and a genuine 2026 split. The essay argues that reconstructability is a precondition of privilege; the interrogation pulls it apart until the student concedes that privilege is not a property of systems at all, and that what an unreconstructable agent destroys is not the privilege but the ability to assert it, which arrives at the same place at the moment it matters.

The third, from Trinity, is the longest and the hardest. It defends the proposition that no organization should deploy an AI agent it could not explain under oath, and spends most of its length deciding what “explain” has to mean before the sentence can be either true or useful. The interrogation presses on technical feasibility, proprietary systems, proportionality and the difference between a witness who is confident and a witness who is right.

Running it against your own work

The tutorial is designed to be run rather than read. The document sets out three self-tutorial modes: solo, in which the reader writes the essay and then writes the interrogation against it; paired, in which a colleague takes the tutor’s chair with instructions to be unhelpful; and machine-assisted, in which an AI system is prompted to argue the other side, which has the useful property of never getting tired and the less useful property of never getting embarrassed.

The examiner’s standard is simple to state. A claim survives if the person making it can say what would have to be true for it to be false, and can point to where that evidence would live. The failure modes are equally simple: an assertion that cannot be falsified, a citation that carries a proposition it does not support, a concession made so early that the argument never gets tested, and the habit of answering a question about mechanism with a statement about intent.

The practical payoff sits one step away from the reading. Take a sentence your organization already says about its AI agents, that its activity is logged, that a human reviews the output, that the vendor maintains an audit trail, and put it in front of a hostile examiner. Ask what the examiner would ask next, and then ask who in the building could answer it.

The tutorial ends where any of these should end, on a question a reader can take to work on Monday. Could your organization reconstruct yesterday’s agent activity if litigation began tomorrow, and who would you send to say so?


The tutorials in full

The three worked examples below are provided by ComplexDiscovery OÜ as illustrative demonstrations of the method. Each pairs a model essay with the tutor’s interrogation that follows it, and each is a written simulation of that exchange, the kind the curriculum suggests a reader can run with a colleague or an AI interlocutor, rather than a transcript of a live tutorial between two named people. The student and the tutor are composite, illustrative voices.

Every case, rule, standard and scholarly work the essays cite exists and was checked against a primary or authoritative source, and the citations were corrected where verification disagreed with the draft. The essays also characterize those authorities and argue from them, and a characterization is an argument rather than a verified fact: where a cited decision is stayed, unresolved or record-bound, the essays say so. The propositions are written to be argued in either direction and are reproduced here as the tutorial states them. They are positions the document puts up to be tested, not positions ComplexDiscovery OÜ holds. In all three examples the proposition is weakened or narrowed under questioning, which is the outcome the form is designed to produce.


Tutorial one, Michaelmas: “An AI agent’s log is not a neutral record but the first draft of its operator’s defense”

Explainer

This example trains a single distinction: between a claim about design and a claim about motive. The essay’s strongest observation is that instrumentation encodes choices made before any dispute; its weakest is that those choices are defensive. Watch the two travel together until the tutor separates them, and watch what happens when the tutor points out that Rule 803(6)(E) already contemplates records made by interested parties. The proposition does not survive in the form it was written. What replaces it is narrower, and considerably more useful in a document request.

The essay

A model essay of the kind the tutorial asks for, followed by the tutor’s interrogation. It takes a position rather than surveying the field.

There is a photograph in every discussion of logging, and it is doing damage. The photograph is the idea that a log captures what happened the way a camera captures a room: imperfectly, perhaps, at limited resolution, but without an argument. On that picture the only questions worth asking are technical. Was the camera on? Was the resolution adequate? Did anyone tamper with the film?

I want to argue that the photograph is the wrong picture and that the right one is authorship. An AI agent’s log is a text composed by an interested party before the dispute in which it will be read. It is the first draft of its operator’s defense.

Begin with what a log is not. It is not the agent’s experience. An agent that retrieves 400 passages, ranks them, discards 380, calls three tools, retries one after a timeout and returns an answer has done something with a rich internal structure, almost none of which is recorded. What is recorded is what somebody decided to record. That decision has a name in engineering, instrumentation, and it is made months or years earlier by people optimizing for debugging, cost and latency. It determines in advance which questions the record will be able to answer.

Consider what instrumentation decides. Granularity: whether a tool call is logged as a call, or as a call with its arguments and its return. Retention: whether the record survives 30 days or 400. Sampling: whether every request is logged or one request in 100, an entirely defensible engineering decision that is ruinous for an evidentiary purpose. Redaction: whether the prompt is stored or a hash of the prompt. Identity: whether the log records which model version answered, or only that the service answered. Each choice is legitimate on its own terms, and each forecloses a category of question.

Now the sharper claim. These choices are not made behind a veil of ignorance about litigation. The people who set retention periods are frequently the same people who have been briefed on discovery costs. An organization that shortens a retention window is choosing, among other things, to have fewer records to produce. An organization that logs the approver’s identity but not the elapsed time between presentation and approval has chosen to be able to prove that a human approved, and not to be able to prove how long they looked. The log will show a name and a timestamp. It will be silent on whether the approval was considered, and that silence is a design output.

Lorraine Daston and Peter Galison give this a name. What they call mechanical objectivity is the nineteenth-century epistemic virtue of letting the instrument speak so the observer need not: the photographic plate rather than the trained draftsman’s hand. Their point is not that mechanical objectivity is false but that it is a moral posture, adopted at a particular moment for particular reasons, and that it works by relocating judgment rather than removing it. The draftsman’s judgment about what to draw becomes the photographer’s judgment about where to point the camera. The judgment does not disappear. It moves upstream, where it is harder to see, and the output now arrives with the authority of a machine.

Cornelia Vismann makes the harder version of the point about files. Files, in her account, do not record legal authority so much as constitute it. An administrative act becomes real when it is filed, and the filing system’s categories determine which acts are available to be performed at all. Read against agent logs this is not a metaphor. A system with no field for uncertainty cannot record uncertainty, and an organization reading its own logs six months later will not find uncertainty there, and will sincerely report that the agent was not uncertain.

The same structure shows up in the vocabulary. Logs describe agent behavior in the terms the designers used to think about it: tool_call, retry, guardrail_triggered, confidence. Those terms carry theories. A field named guardrail_triggered asserts that there is a guardrail and that it operates as a trigger, and a reader who adopts the field name has adopted the claim inside it. Cross-examination of a machine begins with cross-examination of its schema.

The evidentiary framework is at least half aware of all this. Rule 803(6) admits records of a regularly conducted activity, and 803(6)(E) allows exclusion where the opponent shows that “the source of information or the method or circumstances of preparation indicate a lack of trustworthiness.” I read that clause as an acknowledgment that a record made in the ordinary course is made by someone with ordinary interests, though I concede I am reading purpose into it rather than quoting a committee note. Lorraine v. Markel American Insurance Co. makes the parallel move on authentication, treating Rule 901(b)(9) as one route for computer-generated output, among the several nonexclusive methods it surveys, and demanding proof about the process rather than about the artifact.

Here is where the proposition earns its keep. If the log is authored, the discoverable object is not the log. It is the design. The right requests are not for the records alone but for the instrumentation specification, the retention policy and its change history, the sampling configuration, the schema and its versions, and the tickets in which somebody decided to stop recording something. A party that produces 40 gigabytes of agent telemetry and no logging design document has produced the text and withheld the grammar.

A particular failure follows from the photograph picture and deserves a name: the argument from absence. A log that does not record an event is routinely offered as proof the event did not occur. The agent did not open that document, because no access appears in the log. That inference is sound only where the instrumentation was capable of recording the event and configured to do so, and the party offering the absence is the party who chose both. In a sampled system an absence proves nothing whatever. An examiner who accepts an absence without first establishing coverage has accepted the operator’s design as the boundary of reality.

The strongest objection is that I have proved too much. Every business record is created by an interested party. Payroll ledgers are authored. If authorship defeats neutrality then nothing in the file cabinet is neutral, and the proposition collapses into a truism about institutions. I accept a version of this. What distinguishes the agent log is not that it is authored but the distance between how much conduct it appears to describe and how little of that conduct it retains, combined with an unusual readiness among readers to treat it as complete. Nobody believes a payroll ledger is the whole of an employment relationship. The same instinct does not seem to protect the agent log, which gets read as though it were the whole of an agent’s conduct.

The second objection is that “first draft of the defense” imputes motive, and that most instrumentation decisions are made by engineers who have never once thought about litigation. This is true, and I think it is survivable. The claim need not be that the choices are strategic. Diane Vaughan’s account of the Challenger launch decision is instructive precisely because it contains no villains. Choices are made under production pressure, they accumulate, they become normal, and the organization ends with a record shaped by its own convenience without anyone having chosen that shape. Motive is not required for the effect.

None of this argues against logging, and none of it assumes bad faith. It is an argument about where scrutiny belongs. An operator who wants its logs believed has a route available: version the schema, document the retention decision and its date, record coverage rather than only events, and be able to say who changed the configuration and when. That is roughly what Article 12 of the EU AI Act asks of high-risk systems, what the draft ISO/IEC 24970 logging standard is attempting to specify, and what Joshua Kroll and his co-authors mean when they argue that accountability has to be built in before the fact rather than inspected after it. An operator who has done this can hand over a record that answers questions. An operator who has not will hand over a record that answers exactly the questions it was built to answer, which is a different thing, and the difference is the whole of my proposition.

So: treat the log as the operator’s account rather than as the event. Read it the way you would read a party’s contemporaneous memorandum, presumptively admissible, genuinely useful, and never mistaken for the thing it describes. Then go looking for the design.

The interrogation

The tutor’s questioning that follows the essay.

Tutor. You have written a good essay about instrumentation and titled it an essay about defense. Which of those is your proposition?

Student. Both. Instrumentation is the mechanism and the defense is the consequence.

Tutor. Then let us test the consequence, because I think the mechanism is doing all the work. Rule 803(6)(E) already permits an opponent to attack a business record by reference to the method or circumstances of its preparation. If your proposition is correct, what does it add to a rule that has been on the books for decades?

Student. It changes what an opponent has to look for. The rule assumes you can inspect the circumstances of preparation. With a log, the circumstances of preparation are a configuration file and a set of tickets that nobody has asked for, because the log arrives looking like a finished record rather than a prepared one.

Tutor. That is an argument about discovery practice, not about the nature of the record. You could make it without the word defense.

Student. I could.

Tutor. Let us see whether you should. Take your Challenger paragraph. You concede that the engineers had no litigation in mind, and you rescue the proposition by saying motive is not required for the effect. But the word defense is a motive word. It means a document made in contemplation of an accusation. If you are entitled to it without motive, what work is it doing that designed would not do?

Student. It carries a warning. Designed sounds neutral. A reader hears designed and thinks about engineering. Defense makes them ask who benefits.

Tutor. So the word is rhetorical rather than analytic.

Student. Yes. Though I would say the rhetoric is earned by the asymmetry. The operator knows what the log does not contain. The requesting party does not, and cannot find out from the log itself.

Tutor. That is a better sentence than the one in your title. Keep it. Now the harder problem. Suppose an operator has done everything on your list: versioned schema, documented retention with dates, coverage recorded alongside events, a change log for the configuration. Is that operator’s log the first draft of its defense?

Student. It is still authored.

Tutor. That is not what I asked. Is the proposition true of it?

Student. In the sense I mean, yes, but the sense has become so weak that it stops being a warning. If the design is disclosed, the reader can see the boundaries, and a record whose boundaries are visible is not much like a defense.

Tutor. Then your proposition is not about logs. It is about undisclosed design. A log with its design disclosed is evidence. A log without it is advocacy that has not been labeled as advocacy.

Student. That is where I end up, yes.

Tutor. Good, because it also repairs your argument from absence, which I thought was the strongest passage in the essay and which you left underused. State the rule you actually want.

Student. An absence in a log has evidentiary weight only where coverage is established independently of the log. Otherwise the party relying on the absence is relying on its own configuration decision, and has offered no evidence at all.

Tutor. And who bears the burden of establishing coverage?

Student. The party asserting the absence. They chose the instrumentation and they hold the specification.

Tutor. Will a court agree?

Student. I do not know of a decision that says so. The closest analogy is authentication under Rule 901(b)(9), where Lorraine asks the proponent to show the process produces an accurate result. Establishing coverage is the same kind of showing, made about a negative rather than a positive.

Tutor. Say “I do not know of a decision” more often. It is the most persuasive thing you have said this hour. One more. You want the design as the discoverable object. Name the three requests you would actually serve, and tell me which one the other side will fight hardest.

Student. The logging specification and schema with version history. The retention policy with its change history and dates. The sampling and redaction configuration for the period at issue. They will fight the third hardest, because it is the one that establishes coverage, and because it is the one most likely to show that a decision was taken after the dispute was foreseeable.

Tutor. Then write your proposition again, and do not use the word defense.

Student. An AI agent’s log is a designed artifact whose boundaries are set by its operator, and it is evidence of conduct only to the extent that the design is disclosed.

Tutor. That is duller and true, which is the trade you will make most weeks. Next week you will argue that observability is an engineering practice rather than an evidentiary one, and that the difference is retention. You have already conceded half of it, so come prepared to defend the other half against yourself.


Tutorial two, Hilary: “Privilege cannot be delegated to a system whose conduct cannot be reconstructed”

Explainer

This example trains a habit that is harder than it looks: reading two decisions that appear to conflict and deciding whether they actually do. The essay builds on Kovel and on a genuine 2026 divergence between United States v. Heppner and Warner v. Gilbarco, Inc., decided a week apart on opposite sides. Watch the tutor attack the verb rather than the noun, and watch the proposition fail twice: once because privilege is not delegated to anyone, and again because a perfectly reconstructable system can still be unprivileged. What is left is a claim about burden rather than about doctrine, and it is worth more in practice than the sentence it replaced.

The essay

A model essay of the kind the tutorial asks for, followed by the tutor’s interrogation. It takes a position rather than surveying the field.

The obvious home for the claim that a machine can sit inside the attorney-client privilege is United States v. Kovel. Judge Friendly held that privilege reaches a non-lawyer whose participation is necessary, or at least highly useful, to the lawyer’s rendition of legal advice, and he reached it through an analogy to translation. If a client speaking a foreign language may bring an interpreter without breaking the privilege, an accountant who renders financial language into terms a lawyer can advise on stands in the same place. Accounting concepts, as Friendly put it, are a foreign language to some lawyers in almost all cases and to almost all lawyers in some cases.

The analogy is inviting, and I want to argue that it does not extend to an agent whose conduct cannot be reconstructed. Not because a machine is disqualified by being a machine, but because everything Kovel requires is a fact about a describable role, and an unreconstructable agent has no describable role. Privilege cannot be delegated to a system whose conduct cannot be reconstructed.

Start with what Kovel actually asks. The agent must be engaged at the lawyer’s direction. The communication must be for the purpose of obtaining or delivering legal advice. It must be made in confidence. Each of those is a fact, and each is a fact about a particular exchange rather than about a category of participant. The accountant in Kovel was privileged in that engagement, on those instructions, for that purpose. Nothing about being an accountant carried the privilege.

Two decisions from February 2026 show what happens when those facts cannot be established. In United States v. Heppner, a defendant who had received a grand jury subpoena took information learned from his lawyers, put it into a consumer AI assistant on his own initiative, generated roughly 31 documents outlining defense strategy, and later shared them with counsel. Judge Rakoff held the material neither privileged nor work product. There was no attorney-client relationship between a user and an AI service, which alone disposed of the privilege claim. The platform’s terms, which reserved the right to disclose to third parties including the government, defeated any reasonable expectation of confidentiality. And the documents had not been created at counsel’s direction or in a way that reflected counsel’s mental impressions. Sharing them afterward did not make them privileged in retrospect.

A week earlier, in Warner v. Gilbarco, Inc., a magistrate judge in the Eastern District of Michigan reached the opposite result on a pro se litigant’s queries to a generative tool, holding them protected work product. The reasoning was that such systems are tools rather than persons, and that work-product waiver requires disclosure to an adversary or in a manner likely to reach one, which putting a question to a software tool is not.

The two are routinely described as a split. I think that description is careless, and the carelessness is instructive. Heppner is mostly an attorney-client case about a client acting alone, and its work-product holding turns on the absence of counsel’s direction. Warner is a work-product case in which the litigant was the person preparing for litigation, so no direction was missing. The doctrines have different elements and, importantly, different waiver tests. Attorney-client privilege is waived by disclosure to a third party. Work product is waived by disclosure to an adversary or in a manner likely to reach one. The same act of typing into the same box produces different answers because the two doctrines are asking different questions about it.

What the cases share is more useful than what divides them. In each, the court could rule because someone could say what had been done: which tool, on whose instruction, under what terms, containing what. Morgan v. V2X, Inc. makes the point almost explicitly. On the facts before it the court recognized some level of work-product protection for a litigant’s AI-assisted materials, and then held that the identity of the tool was discoverable, on the ground that naming the tool reveals no mental impressions. It went further and conditioned the use of AI platforms with information designated confidential under the protective order on contract terms barring training on inputs, restricting third-party disclosure and permitting deletion, which as a practical matter puts most consumer-tier services out of reach for that designated material. Concord Music Group v. Anthropic, in December 2025, treated attorney-crafted investigative prompts as opinion work product while finding waiver as to a post-suit investigation the party had placed at issue. And in Conservation Law Foundation v. Shell Oil Co., a magistrate judge ordered production of an expert’s AI prompts on the reasoning that an expert’s methodology is fair ground for discovery and the way the expert used the tool is part of that methodology. That last one has not settled: the district judge subsequently stayed the order while he reviews it, and I have found no ruling upholding or vacating it. I cite it for the reasoning, which is the part I want, and not for an outcome it does not yet have.

Read together, these are not decisions about artificial intelligence. They are decisions about proof. In each one the claimant either could describe the conduct or could not, and the outcome followed. Which brings me to the operational form of my proposition. An organization that runs an agent it cannot reconstruct will be unable to satisfy Rule 26(b)(5)(A), which requires a party withholding material to describe its nature “in a manner that, without revealing information itself privileged or protected, will enable other parties to assess the claim.” Consider the log entry such an organization would have to write. Agent-generated analysis. Prompt not retained. Model version not recorded. Retrieval corpus unknown. Instruction source unknown. That entry does not assert a privilege; it announces that the party cannot describe what it is withholding, and it invites exactly the motion it was drafted to avoid.

There is a further problem specific to agents, which Kovel never had to face. Friendly’s accountant was employed by the law firm. He was in the room, on the engagement, subject to instruction. An agent that calls external services mid-task is employed by nobody in the room. Its confidentiality is governed by the terms of service of whichever provider handled the call, and those terms are frequently the whole of the confidentiality analysis, as Heppner shows. An organization that cannot reconstruct which services an agent contacted cannot establish confidentiality even where it is confident the exchange was confidential in fact.

The strongest objection is that I have chosen the wrong verb. Privilege is not delegated. It is not a permission that attaches to a participant and can be handed on. It is a characteristic of a communication, determined by the purpose for which it was made and the circumstances in which it was kept. Nobody delegates privilege to an accountant, and nobody delegates it to a model. I think this objection is correct, and I do not think it defeats the point I am making, though it does force me to state the point differently. The practical claim is not that an unreconstructable system lacks the capacity to hold a privilege. It is that an organization using one has forfeited its ability to carry a burden that has always been its own.

The second objection is more damaging. Reconstructability is not sufficient. Imagine an agent that records everything: every prompt, every retrieval, every model version, every tool call, all of it immutable and time-stamped. Now run it on a consumer platform whose terms permit disclosure to third parties. The conduct is perfectly reconstructable and the communication is not confidential, so there is no privilege to assert. Reconstructability and confidentiality are separate axes, and my proposition addresses only one of them. Nor is reconstructability strictly necessary: a lawyer who directed a specific task, in writing, on a service under contractual confidentiality obligations, may be able to establish privilege from the engagement record even where the system’s internal steps are opaque.

So the sentence I began with is false as written, twice over. What survives is narrower. Where an agent’s conduct cannot be reconstructed, an organization asserting privilege over that conduct will usually be unable to meet its burden of description, and a privilege it cannot describe is one it will not keep. That is a claim about proof rather than doctrine, and it has a practical corollary. The time to build the capacity to describe an agent’s conduct is when the agent is procured, not when the log is served. By the time you are drafting the entry, the facts you need either exist or they do not, and no amount of assertion will conjure them.

The interrogation

The tutor’s questioning that follows the essay.

Tutor. You spent your last two paragraphs demolishing your own proposition. That is either intellectual honesty or bad planning. Which was it?

Student. Honesty, I hope. I could not make the sentence survive contact with the objections.

Tutor. Then let us find out whether you gave up too early or not early enough. You concede that privilege is not delegated. Why did you write a proposition whose main verb you do not believe?

Student. Because the practice it describes is real. People do talk as though putting a task inside a system extends the protection to whatever the system does with it. The verb is wrong, but it names the mistake accurately.

Tutor. A proposition that names a mistake accurately by committing it is a poor proposition. Try the other direction. If privilege is a characteristic of a communication, what exactly is the communication when an agent runs for 40 steps?

Student. That is the difficult question. There is an instruction from counsel at the start and an output at the end, and in between there are steps that may or may not be communications with anyone.

Tutor. Are they?

Student. A retrieval from an internal index is not a communication with a third party. A call to an external model is. The trouble is that a party cannot say which happened unless the agent recorded it.

Tutor. Now you are being useful. Say what follows.

Student. That waiver analysis for an agentic workflow is not a single question about one act of disclosure. It is a question about each external call the agent made, and the answer for each depends on the terms governing that provider.

Tutor. Which is a proposition worth an essay, and it is not the one you wrote. Let us test your reading of Heppner and Warner. You say the split is careless. Would you be comfortable telling a client the two decisions do not conflict?

Student. I would say the holdings do not conflict. The dicta might. Warner says these systems are tools rather than persons, which reads as a general proposition, and Heppner reasons partly from the absence of a relationship with the service, which reads as though the service is more like a person than a tool.

Tutor. So there is a conflict in the reasoning even where the results are reconcilable.

Student. Yes. And the reasoning is what a later court will use.

Tutor. Good. Now the objection you called damaging. You admit reconstructability is neither necessary nor sufficient for privilege. What is left of your essay?

Student. The burden claim. Rule 26(b)(5)(A) requires a description adequate to let the other side assess the claim, and an organization that cannot reconstruct the conduct cannot write that description.

Tutor. Can it not? I have seen a great many privilege logs that describe things nobody reconstructed. Counsel writes “analysis prepared at the direction of counsel in anticipation of litigation” and the entry stands. Why can your client not do the same?

Student. They can write it. I am not sure they can defend it in camera.

Tutor. Say why.

Student. Because in camera review of an ordinary document shows the judge the document, and the document carries its own evidence of purpose and authorship. An agent output carries almost none. The purpose lives in the instruction, and the authorship lives in the configuration, and if neither was retained the judge is looking at text that could have been produced by anything.

Tutor. That is the best point in your hour and it is not in your essay. What does an organization do about it?

Student. Retain the instruction and the configuration as a matter of course, so that the privileged character of the output is provable from records other than the output.

Tutor. Can counsel simply attest to it instead? A declaration saying: I directed this task, on this date, for this purpose.

Student. Sometimes. It puts counsel’s credibility in place of the record, which courts accept until an adversary gives them a reason not to. And it does not answer the confidentiality question at all, because counsel’s intention has no bearing on which external services the agent contacted.

Tutor. So your rule is that attestation can establish direction and purpose but not confidentiality.

Student. Yes. Confidentiality is a fact about the world rather than about anyone’s intention, and only the record establishes it.

Tutor. Write the proposition again.

Student. An organization that cannot reconstruct which external services its agent contacted cannot establish confidentiality, and a privilege claim it cannot establish is one it will lose the first time it is tested.

Tutor. Narrower, provable and worth arguing. You may keep it. Next week the proposition is that human-in-the-loop is an allocation of blame before it is a control. You will find Bainbridge unsettling and Elish worse, and I want you to come in prepared to defend oversight rather than to concede it in the first five minutes, which is what everybody does.


Tutorial three, Trinity: “No organization should deploy an AI agent it could not explain under oath”

Explainer

The longest of the three, and the one the curriculum treats as the term’s hardest. The essay spends most of its length deciding what “explain” has to mean before the sentence can be either true or useful, and it arrives at a testimonial standard rather than a technical one. Watch four pressures in the interrogation: technical feasibility, proprietary systems, proportionality, and the difference between a witness who is fluent and a witness who is right. The tutor’s most damaging question is not about machines at all. It is about who decides, in advance, which agent will matter.

The essay

A model essay of the kind the tutorial asks for, followed by the tutor’s interrogation. It takes a position rather than surveying the field.

The proposition is easy to agree with and almost impossible to apply, which usually means the agreement is being purchased by vagueness. Everything turns on what explanation means, and there are three readings available, two of which destroy the sentence.

The first is mechanistic. To explain an agent is to say why the model produced this output rather than another, in terms of its internal computation. No published method I am aware of delivers this for a model at current scale, and the interpretability literature treats it as an open research problem rather than an available capability. If that is right, the proposition on this reading is not a standard but a prohibition, and a prohibition on essentially all current deployment. A rule that forbids everything gives no guidance about anything.

The second is documentary. To explain an agent is to produce its logs. This is achievable and it is not explanation. A log answers what and when. An examiner asks why, and on what authority, and what would have happened otherwise, and how you know. A party that responds to those questions by handing over telemetry has changed the subject, and an examiner worth the fee will say so.

The third reading is the one the proposition actually contains, and it is hiding in the last two words. Under oath is not decoration. It specifies a person, a setting with an adversary in it, and a sanction for being wrong. The question is not whether the system is explicable in the abstract. It is whether there exists a human being who can be put in a chair and examined about it. That is a testimonial standard, and testimonial standards are the ones the law is actually built to administer.

So: what must that person be able to say? Five things, and the fifth is where organizations fail.

What the agent was for. The purpose, in terms someone outside the engineering group would recognize, and the scope of authority it was given.

What it was permitted to do. The boundary, expressed as capability rather than intention, and the name of whoever set it. Permitted is a different question from expected, and the gap between the two is where most incidents live.

What it did on the occasion in question. The record of the specific episode, with enough coverage that an absence in it means something.

What would have happened had it been wrong. The failure mode, the control that was supposed to catch it, and whether that control has ever caught anything.

How they know each of the four. The provenance of the witness’s own answers.

Most witnesses can give the first four from belief. They have read the design document, they attended the review, they trust the team. Under examination, belief is worth very little, because the fifth question is the one an adversary will spend the afternoon on, and a witness who answers it with “that is my understanding” has told the room that the organization’s account of its own conduct rests on nobody in particular.

Onora O’Neill’s distinction is the useful one here. Her argument against accountability by indicator is that a proliferation of metrics and targets does not produce trust; it produces performances aimed at the metric, while the substantive judgment that would justify trust goes unexercised. What she calls intelligent accountability rests on people with the expertise and the time to assess evidence and to answer for it. A dashboard is an indicator. A witness who can say how they know is an account. The proposition, read testimonially, is a demand for the second.

It is worth noticing that the federal rulemaking process has been circling the same point without landing on it. Proposed Rule 707, published for comment in August 2025, would have applied Rule 702’s reliability requirements to machine-generated evidence offered without an expert witness. After a comment period that closed in February 2026, the Advisory Committee on Evidence Rules revised the draft in May 2026, declined to advance it and declined to republish it for the time being, carrying it to a meeting of technical experts later in the year. The revised working draft presumes an expert foundation while leaving room for other evidence of the output’s reliability. That is a rule about who testifies, not about how models work. None of it is law, and the committee has not decided whether the revised text goes back out for public comment. If it does, as substantial revisions ordinarily do, the earliest effective date is Dec. 1, 2029. If it does not, Dec. 1, 2028 remains possible. But the direction of travel is toward the witness rather than toward the artifact, which is where my proposition already is.

The hardest case is the proprietary one, and it is hard in a way that is often misdescribed. The usual framing is that vendors will not explain their systems and deployers are therefore stuck. The case law is less accommodating than that. In State v. Loomis, the Wisconsin Supreme Court allowed a sentencing court to consider a proprietary risk score, but only with written advisements about its limits and only where it was not determinative. In State v. Pickett, a New Jersey appellate court ordered production of probabilistic genotyping source code under a protective order at the pretrial stage, holding that hiding the source code was not the answer and that a protective order was. Rebecca Wexler’s argument runs further, that there should be no trade secret privilege in criminal proceedings at all, because protective orders already supply the commercial protection the privilege is invoked to secure.

The through line is that opacity is negotiable, and that it is negotiated at purchase. Deirdre Mulligan and Kenneth Bamberger have made the general case: when institutions adopt machine learning, the decisions that matter are made in procurement rather than in policy, and procurement has none of the procedural furniture that would make those decisions reviewable. An organization that cannot explain its agent has, in the ordinary case, bought that inability. It signed a contract with no audit right, no obligation to preserve logs for the duration of a limitation period, no commitment to make an engineer available, and no version-disclosure term. The proposition is therefore a procurement rule wearing a courtroom costume, and it is enforceable at exactly one moment, which is before signature.

The second difficulty is proportionality, and here the proposition as written is plainly too strong. Federal Rule of Civil Procedure 26(b)(1) makes discovery turn on six factors, among them the importance of the issues, the parties’ relative access to relevant information and whether the burden outweighs the likely benefit. A standard requiring courtroom-grade explicability for every agent an organization runs would apply to spam filtering, ticket routing, autocomplete and a hundred other things whose conduct nobody will ever contest. Applied literally it would either be ignored or would consume budgets that produce no protection.

The threshold should be consequence rather than capability. The question is not how autonomous the agent is or how sophisticated the model is. It is whether the agent’s own conduct could become a disputed fact: whether someone might one day need to prove, or disprove, what it did. An agent that drafts a first pass of a document a human then rewrites is unlikely to meet that threshold. An agent that decides which documents a reviewer sees, or which transactions are flagged, or which candidates advance, plainly does, because its conduct is the thing at issue rather than an input to something else.

The third difficulty is the one I find most uncomfortable, because it cuts against my own framing. A testimonial standard rewards good witnesses. A general counsel who explains fluently and confidently satisfies the proposition, and a nervous engineer who says “it depends what you mean by permitted” fails it, and the second person may be the one telling the truth. Explicability read testimonially can degrade into articulacy.

The answer is that the record does not disappear from the standard; its role changes. The log’s job is not to do the explaining. Its job is to make the explanation falsifiable, so that an examiner can test the witness against something other than the witness. This is where the term’s earlier propositions come back. If the log is a designed artifact whose boundaries are set by its operator, then a witness whose account cannot be checked against a disclosed design is a witness who cannot be caught being wrong, and a claim that cannot be caught being wrong is not evidence of anything.

Two objections remain. The first is that explanation is simply the wrong frame. Andrew Selbst and Solon Barocas argue that the intuitive appeal of explainable machines conceals the real question, which is not how a model works but whether its rationale is normatively justified; Lilian Edwards and Michael Veale argue that a right to an explanation is probably not the remedy anyone is looking for, and that impact assessment and design controls do more. I accept the force of both, and I think they are arguments against explanation as a consumer right rather than against explicability as an internal readiness standard. A right to an explanation is a promise made to an affected person. The proposition here is a discipline an organization imposes on itself before it needs one.

The second objection is that the proposition is unfalsifiable and therefore idle. If an organization deploys an inexplicable agent and nothing goes wrong, the proposition is not refuted; it was only ever a claim about exposure. That is fair, and it means the proposition should be stated as a rule about risk rather than as a prediction. It would be genuinely undermined if courts and regulators turned out to accept counsel’s attestation in place of records as a routine matter, in which case the capacity to explain would be a counsel of perfection with no consequence attached to its absence. Nothing in the 2026 decisions suggests that is where things are heading, but it is the shape the refutation would take, and it is worth watching for.

The refined proposition, then. No organization should place an AI agent in a role where its own conduct could reasonably be expected to become a disputed fact unless a named person could state, under examination, what the agent was permitted to do, what it did on the occasion in question, and how they know, with the last of those provable from records the organization already holds.

That is longer, and it costs something to comply with. The named person has to exist before the dispute, not be identified after it. The records have to have been designed to answer questions rather than to reduce storage. The contract has to have said so. None of this is a technical problem, and treating it as one is how organizations end up with excellent observability and no evidence.

The interrogation

The tutor’s questioning that follows the essay.

Tutor. Under oath. You lean the whole essay on those two words. Are they doing analytic work, or are they theater?

Student. Analytic. They convert a question about systems into a question about people, and they specify the conditions: an adversary, and a penalty for being wrong.

Tutor. Most agents will never be near a courtroom. Regulatory examinations are not under oath. Internal audits are not under oath. Have you written a standard for one percent of cases and called it a standard for all of them?

Student. I would say the oath is a stress test rather than a setting. The question is whether the account would survive that setting, not whether it will ever be given in it.

Tutor. Then say “could be examined” and drop the oath, which is doing rhetorical work you have just disclaimed.

Student. The oath adds the penalty. Examination without consequence is a meeting.

Tutor. That is a fair answer. Keep it, and be aware you will be accused of drama. Now your five requirements. The fifth, how they know. Is that a requirement about the witness or about the organization?

Student. About the organization. A witness can only say how they know if the knowing was arranged in advance.

Tutor. Arranged by whom? Name the function.

Student. In the ones I have seen, nobody. Engineering owns the telemetry, legal owns the holds, procurement owns the contract, and the question spans all three.

Tutor. Then your proposition is an organizational design claim that you have dressed as an evidentiary one. Does that bother you?

Student. It should probably be in the essay rather than in this conversation.

Tutor. It should. Let us go to Loomis and Pickett, because I think you have used criminal cases to make a commercial argument and I want to see whether you noticed.

Student. I noticed. The interests are not the same. A criminal defendant has confrontation and due process arguments that a commercial litigant does not.

Tutor. So why are they in your essay?

Student. Because they are the only places courts have been forced to decide how much opacity a proceeding will tolerate. They set a ceiling on the argument rather than a floor. If a court will order source code under a protective order where liberty is at stake, a vendor’s claim that its architecture is simply unavailable is a commercial position rather than a legal one.

Tutor. That is a more careful use of them than your essay makes. Rewrite the paragraph so it says that. And Pickett decided what, precisely?

Student. Pretrial access to source code under a protective order, on a showing of particularized need. The court expressly did not resolve the trial-stage confrontation question.

Tutor. Good. Now proportionality, which I think is where your proposition either becomes usable or falls over. You propose a threshold: could the agent’s conduct become a disputed fact. Who applies that test, and when?

Student. The deploying organization, at deployment.

Tutor. On what information? You are asking somebody to forecast litigation about a system that does not exist yet, in a business that will change, over a limitation period they have not thought about.

Student. It is a forecast, yes. But it is the same forecast the duty to preserve already requires. Reasonable anticipation of litigation is a forecast, and organizations are expected to make it.

Tutor. The preservation trigger is an event. Something happens and the duty attaches. You are asking for a judgment at procurement, years earlier, with no event.

Student. Then the honest version is that the test is coarse. Not “will this be disputed” but “is this agent’s conduct the kind of thing that gets disputed.” Decisions about people, about money, about what a reviewer sees. Those are the categories, and they are knowable in advance even when the specific dispute is not.

Tutor. Better. Now the question I actually wanted to ask. Who in the organization knows the agent was deployed?

Student. Often nobody with the standing to apply the test. A team adopts a tool, it works, it spreads.

Tutor. So your proposition has a precondition it never states: an inventory. You cannot apply a threshold to systems you do not know you are running.

Student. That is right, and it is a serious omission. The proposition presumes a register of agents and their roles, and I have no basis for saying how many organizations keep one.

Tutor. Which makes the Monday morning version of your essay much duller and much more useful than the essay. Now the objection you called uncomfortable. The fluent witness. How does an examiner distinguish an explanation from a performance?

Student. By asking a question the witness could be caught being wrong about. Not what the system does, which invites the prepared answer, but what it did on a specific occasion, and then checking that against the record.

Tutor. And if the record was designed by the same organization that prepared the witness?

Student. Then you ask about coverage before you ask about content, which is the Michaelmas point. If the witness cannot say what the log was configured to capture in the period at issue, their account of the episode cannot be tested and should not be credited.

Tutor. Now suppose the only person who can answer is at the vendor, and the vendor is not a party.

Student. Then the organization needed a cooperation term in the contract, and if it does not have one it is reduced to a subpoena against a third party who has no incentive to be helpful and may be outside the jurisdiction.

Tutor. So the proposition is enforceable at procurement and nowhere else.

Student. Enforceable cheaply at procurement. Enforceable expensively later, and sometimes not at all.

Tutor. Last question, and it is the one you should have asked yourself. Is your refined proposition falsifiable?

Student. Not by outcomes, because it is a claim about exposure and exposure is compatible with getting away with it. It would be falsified if the practice turned out not to matter: if courts routinely accepted counsel’s attestation about an agent’s conduct without any supporting record, so that the ability to explain carried no consequence. I do not see that in the 2026 decisions, which have gone the other way, but that is the form a refutation would take.

Tutor. Then you have a proposition, a threshold, a precondition you omitted, and a test for whether you are wrong. That is further than most people get by the end of term. For collections, you will argue that the existing rules are adequate to agentic AI. Then you will argue that they are not. I will be more interested in the second one, and I expect you to be worse at it, because you have spent eight weeks building the first.


An invitation

The tutorial ends without a verdict, which is the form working rather than the form failing. Nothing in the three worked examples resolves whether an agent log is evidence or telemetry, whether human oversight is a control or an allocation of blame, or how much of a machine an organization must be able to account for before it puts one to work.

The useful way to read it is to argue the side you do not believe. If you are confident that an agent’s log is a designed artifact whose boundaries are set by its operator, take the other chair and defend it as an ordinary business record. If you are confident that privilege survives the use of an agent whose conduct cannot be reconstructed, try writing the privilege log entry. If the proposition that no organization should deploy an AI agent it could not explain under oath strikes you as obviously right, work out what it would cost to comply with and who in your building would pay it.

The questions are free. The thinking is the work, and it is easier to do now than in the week someone asks for the logs.


Sources and reference materials

References are cited in full below. Where a freely accessible public version exists, the title links to it.

ComplexDiscovery OÜ

Cases and judicial decisions

  • Concord Music Group, Inc. v. Anthropic PBC, No. 24-cv-03811-EKL (SVK), 2025 WL 3677935 (N.D. Cal. Dec. 18, 2025).
  • Conservation Law Foundation, Inc. v. Shell Oil Co., No. 3:21-cv-00933 (D. Conn. May 18, 2026) (magistrate judge’s order compelling production of an expert’s AI prompts; subsequently stayed pending district court review of the order, with no ruling upholding or vacating it located as of publication).
  • Da Silva Moore v. Publicis Groupe, 287 F.R.D. 182 (S.D.N.Y. 2012).
  • In re OpenAI, Inc., Copyright Infringement Litigation, No. 1:25-md-03143 (SHS)(OTW) (S.D.N.Y.) (preservation order entered May 13, 2025; going-forward preservation obligation terminated as of Sept. 26, 2025 by stipulated modification entered Oct. 9, 2025; production of a de-identified log sample ordered Nov. 7, 2025, reconsideration denied Dec. 2, 2025, objections overruled Jan. 5, 2026).
  • In re Valsartan, Losartan, & Irbesartan Products Liability Litigation, 337 F.R.D. 610 (D.N.J. 2020).
  • Livingston v. City of Chicago, No. 16 CV 10156, 2020 WL 5253848 (N.D. Ill. Sept. 3, 2020).
  • Lorraine v. Markel American Insurance Co., 241 F.R.D. 534 (D. Md. 2007).
  • Moffatt v. Air Canada, 2024 BCCRT 149 (B.C. Civ. Resolution Trib. Feb. 14, 2024) (Canadian civil resolution tribunal; persuasive rather than precedential in United States proceedings).
  • Morgan v. V2X, Inc., No. 25-cv-01991, 2026 WL 864223 (D. Colo. Mar. 30, 2026).
  • Rio Tinto PLC v. Vale S.A., 306 F.R.D. 125 (S.D.N.Y. 2015).
  • Schulte v. LinkedIn Corp., No. 22-cv-00237-HSG (LB) (N.D. Cal. June 30, 2026).
  • State v. Loomis, 2016 WI 68, 371 Wis. 2d 235, 881 N.W.2d 749 (2016), cert. denied sub nom. Loomis v. Wisconsin, 137 S. Ct. 2290 (2017).
  • State v. Pickett, 466 N.J. Super. 270, 246 A.3d 279 (App. Div. 2021).
  • United States v. Heppner, No. 25-cr-503 (JSR), 2026 WL 436479 (S.D.N.Y. Feb. 17, 2026).
  • United States v. Kovel, 296 F.2d 918 (2d Cir. 1961).
  • United States v. Lizarraga-Tirado, 789 F.3d 1107 (9th Cir. 2015).
  • United States v. Washington, 498 F.3d 225 (4th Cir. 2007).
  • Upjohn Co. v. United States, 449 U.S. 383 (1981).
  • Warner v. Gilbarco, Inc., No. 2:24-cv-12333, 2026 WL 373043 (E.D. Mich. Feb. 10, 2026).
  • Williams v. Sprint/United Management Co., 230 F.R.D. 640 (D. Kan. 2005).
  • Zubulake v. UBS Warburg LLC, 220 F.R.D. 212 (S.D.N.Y. 2003) (Zubulake IV); Zubulake v. UBS Warburg LLC, 229 F.R.D. 422 (S.D.N.Y. 2004) (Zubulake V).

Statutes, regulations and rules

  • Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), arts. 12, 13, 14, 19, 26(6), 50, 86, 99.
  • Regulation (EU) 2026/1744 of the European Parliament and of the Council of 8 July 2026 amending Regulation (EU) 2024/1689 and others as regards the simplification of harmonised rules on artificial intelligence, published in the Official Journal July 24, 2026, in force July 27, 2026.
  • European Commission, “Regulatory framework for AI,” application timeline as revised by the AI omnibus.
  • Fed. R. Civ. P. 26(b)(1), 26(b)(5), 26(f); 34(b)(2)(E); 37(e).
  • Fed. R. Evid. 502; 801(a); 803(6); 901(b)(9); 902(13)-(14).
  • Proposed Fed. R. Evid. 707 (machine-generated evidence), published for public comment Aug. 15, 2025; comment period closed Feb. 16, 2026; revised and not advanced by the Advisory Committee on Evidence Rules, May 7, 2026.
  • American Bar Association, Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512, “Generative Artificial Intelligence Tools” (July 29, 2024).

Standards and frameworks

Scholarship and commentary



Assisted by GAI and LLM Technologies

Additional reading

Source: ComplexDiscovery OÜ

ComplexDiscovery’s mission is to enable clarity for complex decisions by providing independent, data‑driven reporting, research, and commentary that make digital risk, legal technology, and regulatory change more legible for practitioners, policymakers, and business leaders.

 

Have a Request?

If you have information or offering requests that you would like to ask us about, please let us know, and we will make our response to you a priority.

ComplexDiscovery OÜ is an independent digital publication and research organization based in Tallinn, Estonia. ComplexDiscovery covers cybersecurity, data privacy, regulatory compliance, and eDiscovery, with reporting that connects legal and business technology developments—including high-growth startup trends—to international business, policy, and global security dynamics. Focusing on technology and risk issues shaped by cross-border regulation and geopolitical complexity, ComplexDiscovery delivers editorial coverage, original analysis, and curated briefings for a global audience of legal, compliance, security, and technology professionals. Learn more at ComplexDiscovery.com.

 

Generative Artificial Intelligence and Large Language Model Use

ComplexDiscovery OÜ recognizes the value of GAI and LLM tools in streamlining content creation processes and enhancing the overall quality of its research, writing, and editing efforts. To this end, ComplexDiscovery OÜ regularly employs GAI tools, including ChatGPT, Claude, Gemini, Grammarly, Midjourney, and Perplexity, to assist, augment, and accelerate the development and publication of both new and revised content in posts and pages published (initiated in late 2022).