Flooding the Zone
Most of us have rights: things we can, or are entitled to do in certain circumstances. The “circumstances” bit is important, since rights are typically limited in many ways, with - depending on the right - multiple exceptions and carve-outs, some explicitly spelled out, and others arising from a mix of common sense and jurisprudential accretion.
One such limit, and maybe the ultimate fallback, is that of “abuse”. Countless research theses have been written on how the notion of abuse of rights has been deployed in various jurisdictions, sometimes in contradicting or incompatible manners. But this is a feature, not a bug: the notion is elastic, its contours are blurry, precisely because it needs to be applied to various, ever-different situations. A new scenario will occur, offering the temptation to push these contours further, and reasonable minds may disagree.
Case in point: is this abuse ?
On July 28, the full Supreme Court of Chile reviewed an unprecedented event for the country’s judiciary: a lawyer had filed 38,477 documents in civil courts across the nation in less than three days, using an automated system to process mass resignations and requests to reopen archived cases. […] The collapse of the justice system was not due to the sophistication of the tool used, but rather because the judicial infrastructure was unprepared to handle such a massive workload.1
This was not merely anecdotal, since the court’s electronic system buckled under what amounted, in effect, to a massive denial of service attack.
Reportedly, the Santiago civil judges’ committee considered that this conduct should be qualified as “procedural abuse”, invoking a provision that covers procedural bad faith. Their letter to the Supreme Court asked that the lawyer’s access to the platform - already blocked at the IP level by the judiciary’s administrative arm - be kept restricted, raised the possibility of criminal prosecution or at least disciplinary measures, and asked that “the use of AI for the entry of documents be prohibited” until an institutional policy exists. (The full Court, in response, asked its IT department for a technical report first.)
There is, however, no question that the lawyer had a right to accomplish the administrative acts in question; in fact, it appears that in some cases it was even his duty to warn courts that he would resign from some of these cases, since he had indeed stopped working for the (large) parties at stake. The difficulty lies in the fact that he did it at scale, reportedly with the help of an “external agent”,2 overwhelming a system that, though rather recently digitalised, was not ready for this level of throughput.
But is that an abuse ? To a large extent, there is something to the idea that a competent and diligent lawyer (or “officer of the court”, to echo the notion of some Common Law countries) should account for the limits and capabilities of the digital systems they are using. You can understand why the judges would be pissed off that their system broke, not due to a deliberate attack, but for the dumbest possible reason: too many uploads at the same time. But this is far from traditional abuse where a a right is used for an improper end; here the acts themselves and their ends appear legitimate, only their aggregation turned that legitimate conduct into system-level misconduct.
And there is the thing that this incident revealed: the judiciary’s digital tools and systems were not, in fact, ready for digitalisation; they were still premised on the notion of people doing things one by one with their hands, even if these hands stay on a keyboard and a mouse, and the signatures become embedded in a .pdf. But actual digitilisation should open the door for broader uses of technology, including that of AI and other robots.
Which is why the judges’ knee-jerk suggestion to prohibit any use of AI to file documents, as noted elsewhere, would not only be short-sighted and unworkable: it’s also a missed opportunity to improve and get ready for the new tools. It falls into the classic error of trying to protect a system and a status quo without wondering what that system is ultimately about: the litigants, who would probably prefer that mundane procedural acts be automated and made cheap.
Now, that may be an abuse; years from now, hundreds or thousands of procedural acts made in parallel may become a new normal.
Everyone can read your chats
I noticed last week that there is something a bit profane in looking at someone else’s conversation with a chatbot. There is a fuller essay to be written on this, but this is an altogether new type of writing material: not a diary, not a doodle, not the transcript of a conversation with another human, but a different way to interact with an intelligence unlike our own. That queasiness, as far as I am concerned, would extend to the history of chat conversations, although that’s likely a perfect way to get to know someone.3
And yet, some people love to read other people’s chat conversations, and we discussed it already: prosecutors. As I put it, “LLM logs are, among other things, a vast and growing archive of mens rea”, but that’s useful for any kind of party that opposes you: short of your own thoughts, which cannot (for now) be read or recorded directly, some AI conversations might be the second best place to find incriminating evidence against you. (As I also described, one next step will be querying the agent or chatbot directly for their opinion on what your conversation meant; fun times ahead !)
Which makes this story from the Washington Post unsurprising.
A Washington Post review of public records and local news stories found that chatbot logs were cited in 12 court cases over the past two years. It’s hard to know how often chatbot material is drawn into investigations and legal proceedings more broadly, because police, and parties in civil cases, don’t have to present in court all the evidence they obtain.
The Post also describes how OpenAI (and, probably, other labs) increasingly reach out to law enforcement authorities to, well, snitch on you. At the same time, Sam Altman is quoted as suggesting that conversations should receive some kind of protection similar to privilege. (You know my views on that.) Otherwise, as one of the experts being interviewed opines, “Your entire world is going to now be available for police.”
At the same time, evidence is not always inculpatory; it can also be exculpatory. The Post has no example of someone using chat logs to prove they were innocent of something, but that’s not far beyond the imagination.4
And beyond that: before the Summer, I predicted that “the AI made me do it” would become an exculpatory defence, a way to explain how one ended in legal troubles, and why they should walk away scot free.
It did not take long for an example to pop up. In Autofit, a case decided in August by an Administrative Law Judge, an employer explained in written termination explanation filed with a state agency that they had fired an employee for discussing pay with co-workers - an activity very much protected under American law. At the hearing, the employer suggested that the testimony had been mistaken, and that ChatGPT had added that detail out of the blue. A sceptical judge quashed that defence, and refused to see AI as a fall guy.
But that might work one day ! Some legal defences already offer ground for it, or at least provide reasons to lessen one’s culpability. Reasonable reliance on counsel may shield against an accusation of bad faith, just like the same reliance on (bad) accounting or tax advice may weaken the inference that someone deliberately broke the law. There is no reason that won’t apply to discussions with AI.
“You asked ChatGPT and it told you how to do it”, prosecutors will increasingly say; “exactly, and that’s why I am innocent”, some defendants will try to argue.
AI & Bluebook
A well-known observation in the field of AI is that there is a decoupling between what humans consider hard and what computers and automated systems manage to achieve first. This is the broader insight from Moravec’s Paradox: computers have beat humans at chess first, then go, and now can compose better poems than 95% of the population (a conservative estimate). Yet, they also lack grounding, are gullible, struggle to walk (though robotics is doing leaps), and for a long time could not tell you how many “r”s there are in strawberry.
This lieu commun clashes with another view of the increasing role of AI in our lives, one that views it as a normal technology that will be adopted gradually, one step and one purpose at a time. This view is particularly common in, well, anyone looking around and wondering what AI means for them, especially in a business and corporate context. Use-case identification and validation is largely premised on this gradual approach, of trying out and testing small things before giving AI agents the keys, shaking their hands, and going to the beach while they take care of increasing profits.
To a large extent there is no contradiction here, but two non-commensurable scales: some drudgery does require higher cognitive functions, or at least the kind of intuition and judgment that, currently, cannot be distilled in a way or through proxies that an AI may understand. And so, contrary to expectations, we might have AIs capable of doing extremely expert work before they manage to align bullet points perfectly.5
That thought has recently been bolstered by a recent paper (SSRN) by Matthew Dahl and Eric Martínez, which looks at how well AI can do Bluebook editing:6
This article presents the first empirical examination of AI performance on perhaps the most ubiquitous and lamented form of legal drudgery: citation formatting under the Bluebook. We make four contributions. First, we develop a new benchmark of 2,058 Bluebook queries and show that, on average, frontier language models produce a fully compliant legal citation only 42.6% of the time in a zero-shot setting. Second, we conduct an experiment with five top law reviews and show that even a “reasoning” model falls far below the average score of the human candidates in these journals’ annual editor-selection competitions.
They also found that Retrieval-Augmented-Generation (to retrieve the Bluebook rules in particular) was not really effective in this context: even when the exact rule to be followed was in the prompt, accuracy did not reach par with humans.7 Finally, the authors experimented with “a neuro-symbolic system” that mixes LLM parsing and deterministic rules to achieve much better results, and which they use to advocate for a “formal turn” in Legal AI: less focus on AI resolving hard, judgment-laden tasks, and more on using the new technology to accomplish non-discretionary, one-to-one tasks.
And indeed, what’s particularly interesting here, is that, ultimately, Bluebooking is some sort of many-to-one mapping: you have countless ways to refer to a source, and it forces you to renounce all that variance for one canonical form. (In turn, that form has meaning: citing is relying on authorities, and that authority is (partly) signalled by how it is displayed on the page. Ignore it, and it’s not only your credibility as a drafter that’s lessened; it’s also that of the source.)
The nice thing about the authors’ methods is that their use of AI is limited to a middle step: encoding the “many”, unstructured data into an intermediate step that can then be used by non-AI, deterministic rules to do the job. Which, ultimately, further complicates the story about AI being fit for higher-cognition tasks or drudgery: they can also, and sometimes only, work at middle steps that allows us to close the automation gap and put the question of cognition or judgment aside.
It’s unclear whether an “AI” tool or some kind of automated process, but the latter may well have been coded by an AI assistant, which would come to the same thing.
You read it there first: allowing someone access to your chatbot account will become a significant plot point in many a romance story.
The switch scenario, of using AIs to prove that your case is legit is, well, all too often what lands you in the hallucination database.
Another way to look at it is that this reveals that some tasks’ prestige is decorrelated from the cognition and judgment that goes into it, which is why AI can manage those with ease.
For those who did not encounter it in the American legal education system, first, keep that cherished innocence and, second, you just need to know that Bluebook is one citation style meant to be the standard for all sources in legal writing. And of course, it falls in the XKCD scenario.
While the authors rightly use it as a way to challenge the notion of “law-following” AI, there is one rebuttal here: Bluebook citations are, to a large extent, rather arbitrary, making it harder - if not meaningless - to follow them to the letter.

