Experts on Chat
Some legal cases touch on difficult or complex matters, which justify hiring an expert to assist and testify on behalf of your approach.1 But then, how do you pick said expert ?
There are different criteria to take into account. Primarily, since relying on an expert represents to some extent an appeal to authority, you want to maximise that authority: credentials, reputation, accomplishments, fancy titles and accolades from learnèd societies all go a long way towards persuading a judge or a jury that your side is concrete-solid. Other considerations may include a talent for pedagogy, experience with a given court, and, why not, good looks too.
All this matters, but one point may be even more important: I am yet to see any case where lawyers would pick an expert that disagrees with their theory of the case.
Certainly, there is an equilibrium to be found here, whereas your expert’s credibility comes from a certain independence, and hacks are not listened to. But from a theoretical standpoint (if not from actual experience and selection effects), both client and expert ultimately act as a team, and an expert that does not eventually conclude broadly in favour of the party that appointed them should probably resign (and, in all likelihood, would not have been hired in the first place).
In other words, you expect your expert to, at the very least, broadly agree with you, and your expert knows that you expect that from them and acts accordingly. Meanwhile, everyone on the other side - be it the opposing party, or the jury and judge - will listen to your expert knowing that as well, accounting for it in how deeply they listen and are ready to follow the expert’s conclusions. Nothing wrong here, and much turns on the nuances and subtleties of how this plays out.
Which is why I was a tad unfazed when reading this nice story from 404 Media about an expert who had asked a chatbot to help him come to the conclusion that would help his client. As they put it:
An expert witness testifying in a lawsuit about liability for a Houston explosion that killed three people and destroyed roughly 200 homes used ChatGPT to write significant portions of his “expert report.” The man, who was hired by the industrial product conglomerate 3M, exposed his AI prompts publicly. They showed that he asked ChatGPT to help him “create an exceptional expert witness report defending the standard of care at 3M,” and that the report should “show how 3M is 0% at fault for the explosion at Watson Grinding.”
404 Media also obtained the transcript from the subsequent trial, where opposing counsel played along with the feigned shock of discovering that a counsel might try to reach a pre-determined answer.

Opposing counsel obtained the logs (available here) after noticing that some of the documents being disclosed looked as though they had been created by ChatGPT. They make for an interesting read,2 and I am certainly in awe of the expert’s trust in GPT to handle more than 200 different files at once. Even more damning - and in my view the one detail that should have shattered the expert’s credibility - was the fact that the conversations were publicly available, somehow.
Anyhow, the AI logs where the expert asked ChatGPT point blank to deliver a full report in favour of the client certainly don’t look good. And there is a (legally important) distinction between doggedly pursuing a set conclusion, or actually trying to exercise your expertise in a given case. Yet, and while I am sure many experts have better processes to get there, this is also why you never want to know how the sausage is made.
And more generally, this is yet another instance of AI shattering some fictions we had - just like the fiction that witnesses come at you virgin of all influence from counsel, or the fiction that the legal system is ready to enforce and uphold all of our rights. For as long as anyone can remember, independent and impartial experts have somehow always found in favour of their client, and we have long agreed to ignore that coincidence. But it will be harder to do if we can get their chat history.
Water(mark) cooler talk
Last month, Matt Yglesias wrote about “the lost joy of monoculture”, or the fact that - save for collective events such as the World Cup - we are slowly losing cultural moments or elements that are sufficiently shared or experienced by people to serve as fodder for small talk and ice-breakers.
The point is well-made, but I’d venture one new element of monoculture that, for the past three years at least, has fuelled countless conversations: AI, of course, and AI writing in particular.
If you are more than an occasional user of LLMs, you probably know what I am talking about. The tics, the cadence, the weird preferences for certain terms, you have seen them again and again. Some are specific to some models (I could list several Claudisms, and hiss every time I see the term “load-bearing”), others are shared across the whole species (perhaps an unexpected outcome of distillation). Some of these are readily understood based on the training data, others remain mysterious.3 Hollis Robbins offered a guide to identify AI writing a year ago, and it still mostly holds up:
And so, in some circles at least it has become a significant part of life and shared experience with friends and collaborators to look at some text and opine that “this is obviously Claude”, or “they could at least have removed the em-dashes, etc.” Not always in disbelief or out of being annoyed: as with everything, there is a gradation of talent in (re-)using AI writing, with some people acting as meat proxies, and others just integrating it in their own talents as writers. Or at least this is to be hoped for.
Now, the big recent news is that Anthropic (following Google) is deploying some watermarking, so that one could always know, with some confidence and through a future API check, that some text is or is not likely to be from the most recent Claude models.
This aroused some passions, not all of them well-founded. Part of it is probably downstream of fears that the shortcuts people have only recently adopted will be laid bare, and that others will be able to peer into their working process - never a pleasant prospect.
Another concern arises from the notion that this means the output will be manipulated, and quality be degraded somehow. There is little to it when you understand the mechanics of the watermarking (Sebastian Raschka has an excellent explainer here), meaning that open-ended outputs are likely not at risk.
But there is a legitimate query about what this means for very precise language: already, Anthropic has acknowledged that it won’t apply to most computer code, where you can’t simply sample between different tokens. Well, this is often true of legal text as well, and even more so when dealing, e.g., with quotes or citations. Anthropic’s point that watermarking will attach less to “passages where there are fewer choices that can be made without decreasing the accuracy of the text” slightly soothes this concern, but raises the question whether the algorithm will easily know the difference. It would be a terrible result if watermarking increases the risk of hallucinations.
Anyhow, my point is simpler: to a large extent, if you practice it enough, you already know it’s AI. Watermarking, just like Pangram, will be helpful to confirm your hunch, but that hunch is worthwhile as a hunch: it shows you know the model well. Not all hunches need verifying.
Indeed, there is value in suspecting, but not checking, that something is AI: in preserving a Schrödinger authorship, and with it a certain maintenance of disbelief. Knowing forces you to confront it, to maybe think less of a collaborator or a student, while doubting lets you carry on. Knowing also, often, answers the wrong question, which should not be whether AI has been used, but whether it has been used without judgment. Sometimes, it’s better not to know.
Artificial and Capricious
An important thing to understand about legal hallucinations is that LLMs are more likely to generate them when (i) there is no source or authority to back you; and (ii) the model is nonetheless trying to help and please you. This explains why fabricated authorities are often uncovered by opposing counsel or judges who probe the one argument that, surprisingly, has you winning - it’s too good to be true.
But this also means that the danger is particularly acute every time you have to offer an explanation or an argument for anything; “Chat, say I am right, with sources” is a recipe for disillusionment.
Alas, a lot of our life is spent looking for such explanations, arguments, or justifications. Sometimes, the duty to give reasons is even a legal requirement - as under the US’ Administrative Procedure Act (“APA”).
Which leads me to a very interesting case that, somehow, flew under the radar: last week, the U.S. District Court for the District of Columbia enjoined a policy change from the Department of Health and Human Services (“HHS”), as being likely arbitrary and capricious. HHS had tried to shift its Teen Pregnancy Prevention programs by focusing on abstinence and a vague concept of “body literacy”, for which HHS itself had acknowledged there was a “near absence” of existing standards.
This did not, however, prevent HHS from backing its new approach with scientific literature that, when checked, proved non-existent. The court:
On the topic of body literacy, the notices (remarkably) reference public health studies that appear either not to exist or not to support the propositions for which they are cited—a hallmark of AI-generated citations. See id. at 6–7 & nn. 1–5; FY 2026 Tier 2 NOFO at 7–8 & nn. 1–5; see also Decl. of Kate Talmor ¶¶ 6–12 (“[I]t appears that five out of seven of the NOFOs’ cited articles could not be found as cited in the NOFO. Two out of the seven appear to be completely made up. Three of the seven did not publish in the cited journals but appear to have similar titles to articles published in completely different journals.”)
They had it coming. In 2025, a team behind the “Make America Healthy Again” report at the HHS had already been caught smuggling hallucinations. In the same vein, the FDA, itself part of the HHS,4 hurried to roll out its own AI assistant, Elsa, only to see its employees spooked by the high level of hallucinations produced by the tool.
The recent case involved yet another team from the HHS, which partly points to what went wrong here: beyond a simple story of some employees going rogue and misusing AI, there is a question of inadequate institutional AI governance and culture. Shortly put, when you combine rapid adoption, strong top-down enthusiasm, and insufficient verification together, such errors become rather less surprising. No wonder, then, that HHS’s Office of Inspector General disclosed in July 2026 that it would launch an audit of its AI strategy.
But the broader point, as always, is what goes beyond these sorry mishaps. The court caught the hallucinations because the policy here had been litigated; who knows what other documents from the same period contain fabricated sources. Just as a hallucinated quote, taken up by a court without checking, poisons the epistemic commons, so fake policy material gets laundered through the authority of the state.
Which makes the APA requirement that agencies “show their work” even more critical, as a form of epistemic quality control against synthetic reason-giving. Crucially, the hallucination was (part of) what led to the conclusion that the policy was arbitrary and capricious, just like hallucinations in the text allowed a Québec court to set aside an arbitral award a few months ago; they are a smoking gun that standards have been relaxed.
Hallucinations are thus both part of the problem and the remedy. One almost wishes they never cease to appear.5
Though not all cases may need it: see this nice example.
We might need a new word for the rather sacrilegious feeling of reading someone else’s chat logs, it seems profane.
See, notably, this inquiry in The Atlantic regarding the ubiquitous “it’s not X, it’s Y”.
Which at this point may just be renamed the Health & Hallucinations Service.
There is a parallel here with plagiarism: easy to catch, and useful for spotting academic fraud, but likely to disappear as plagiarists adopt AI to rewrite everything.


