Contract MadLibs
A co-author and I hid negotiated terms in real agreements and asked humans and machines to fill in the blanks. We then wrote up what we found for the law reviews.
At the end of my post on cross-examining agentic commercial agents, I promised that I’d return this summer to argue that AI tools might help get contract law out of the pickle that AI has gotten it into.
This is that post. It summarizes a new draft paper, Generative Gap Filling, written with Yonathan Arbel, and just out today.
Think about AI’s use in law and what do you picture? For most of us, it’s not a comforting scene: job losses, hallucinated briefs, and the death of the knowledge economy. And just under the surface is the suspicion that AI outputs simply aren’t super reliable. They are super plausible, but you can’t quite ever know if they are right.
It’s thus not surprising that when we published Generative Interpretation, arguing that courts could use language models to help parse contract text, sharp critics threw down a epistemic gauntlet: “no experiment can determine whether a generative method yields correct results, because there is no accessible source of ground truth for legal meaning.” In other words, generative interpretation had alot in common with bullshit.
At first, this stumped us. Our first paper showed that you could use 2024 LLM models to match judicial interpretation in litigated cases. That provided an arguably better method than dictionaries, hunches and latin maxims. But courts are infallible only because they are final. That a panel of judges say that that the word “flood” includes man-made water overflows doesn’t mean that the parties actually intended that result. It just means they have to live with it. So, indexing on courts satisficed, it didn’t really answer the critics’ charge.
But, having mulled over the problem for a year, eventually we borrowed a trick from the machine learning literature.
Mad Libs as an Interpretative Scoring Sheet
The trick is called masking, and it’s how language models are trained in the first place: hide part of a text, ask the model to predict what’s missing, and check against the original.1 We thought this technique provided a novel way of asking if a respondent could accurately engage in a type of legal interpretation. At the same time, it might in the future be a method to benchmark AI models at this critically-important judgment.
So, we took real, executed agreements — a talent agency’s deal with a touring rapper, a contingency fee agreement, a requirements contract between a bottle maker and a tea company — and redacted one negotiated term from each. These weren’t mere boilerplate (we thought) but rather contingency clauses that told courts what to do when the show got cancelled, or the client settled behind the lawyer’s back.2
Then we wrote a realistic dispute that would trigger the missing clause and asked three kinds of readers to predict what the parties had written: ordinary people (about 465 of them, via Prolific), law students and practicing lawyers, and a panel of frontier language models. Everyone got the whole contract except the masked term. Either they recovered what the parties wrote or they didn’t. The questions were multiple choice, with four answers. So guessing would get you the right answer 25% of the time.
What happened
In our masking experiment, humans were decent gap fillers. Lay subjects got the hidden terms right about half the time, twice as often as chance would predict. And legal experience helped. Law students edged out lay readers. Practicing lawyers were better still, mostly.
Then came the machines.
Given nothing but the rest of the contract, the models predicted what the parties themselves had written nearly nine times in ten.
Here’s the paper’s key figure.
Lots to talk about here apart from the headline machine overlords situation.
For example, look at the middle panel of the figure, the Bottles scenario. It’s a requirements contract. The buyer forecast a truckload of extra 20-ounce bottles for a marketing campaign, the manufacturer made them, and the buyer then refused to take delivery. The masked clause let the manufacturer invoice for inventory held on a forecast. Lawyers scored 37.5% on that one, worse than the lay readers, who scored 61.6%.
Why? The lawyers who missed it clustered on the answer holding that payment obligations run only to signed purchase orders. Which was, to be fair, a very sensible arrangement! The comments to their guesses made clear they’d thought about it and they thought it was the industry default. It just isn’t what these parties agreed to. Two decades of experience taught the lawyers what supply contracts usually say, and that knowledge pulled them away from what this one actually said. Legal expertise is pattern matching, and pattern matching cuts both ways: it’s a superpower when the deal is standard and a trap when the parties have contracted around the standard.
Did the models actually read the contract?
The obvious worry about the machines’ 88% is that they weren’t “reading” at all, just regurgitating what contracts like this usually say. We ran two checks.
First, we redrafted the contracts to flip their direction — pro-buyer deals became pro-seller — and asked again. Accuracy dropped from 88% to 60%. Second, we told the models the contract had been accidentally left unattached and asked them to guess anyway. Accuracy dropped to 68%.
I read the figure as answering two questions.
That the models’ performance got worse after we perturbed the contracts tells you that they really were extracting information from the specific document in front of them. When you change the document, the models change their prediction, to one that’s incorrect! So, models (presumably like, but better than, people) are inferring from the deal’s text.
In the paper, we analogize this to a degraded radio signal that can be recovered at a distance with the right receiver. Contracts contain overlapping sources of the parties’ intent, because lawyers overengineer them just that way.
But the no-contract bar, sitting well above chance, tells you something different: a lot (maybe two-thirds, in our data) of gap filling is just knowing what deals of this type look like. The models carry strong priors about contracts, and reading tells them when this deal is different. Which, come to think of it, is a pretty good description of what we hope experienced lawyers do. The models were just better at the second step than our lawyers were.
We also scaled up, masking terms in 119 commercial contracts pulled from 2025 and 2026 SEC filings. Overall accuracy held at 87%. And the pattern by contract type is what the priors-plus-reading story predicts: templated promissory notes and credit agreements came in at 95% and up, while bespoke, heavily negotiated provisions like indemnification (61%) and registration rights (67%) were much harder.
Notably, when the models were wrong, they converged on the same errors. All six models missed the same five contracts, each of which reversed a market default.
So what? Choice of model clauses
The doctrinal payoff, we think, is that contracts don’t run out of meaning nearly as soon as the gap-filling literature assumes. For a century, scholars have treated the hypothetical bargain — the deal the parties would have struck had they thought about it — as invisible, and worried accordingly that judges “filling gaps” are really just decorating deals with their own preferences. Our results suggest the hypothetical bargain is often statistically legible. It can be read off the rest of the document, by anyone with the right receiver.
How should courts use that? We definitely don’t think the answer is to turn on the opinion-writing-generator and go golfing. The natural home for this evidence is the ordinary adversarial process: one side runs the contract through a model and discloses everything (model, version, prompt, settings), the other side runs its own query, and the judge weighs the outputs the way she already weighs imperfect evidence from outside the four corners. Nothing about that displaces Judge Hercules; it just makes the inferential work visible and contestable instead of intuitive and unreviewable.
And because parties can anticipate all this, they can contract over it. We propose what we call a choice of model clause: just as parties choose their governing law and their forum, they can name, in the agreement itself, the model whose reading of any silence will get weight. This isn’t as exotic as it sounds. Parties have incorporated external interpretive resources by reference forever — ISDA definitions in derivatives, AIA conventions in construction, even clauses specifying which dictionary controls. A model is a stranger technical standard than Black’s, but it’s in the same family. To the extent their is pickup for this idea, my guess is that it’ll be first in the arbitration space, but courts ought to be at least curious what choice of model clauses will do long term, because it’s kind of nifty.
…And at equilibrium you’d expect?
I keep on gesturing to AI equilibrium effects and today is no exception. If choice of model clauses become enforceable, the rational move is to run the contract through the named model before signing. Both sides look at the draft, query the model about its silences, and see, together and in advance, exactly how every gap would be filled.
What comes next? Let’s consider the current ecology of contractual silence. The literature has long-sorted silence into three bins.
Sometimes silence is inadvertence: nobody thought about the contingency.
Sometimes it’s deliberate economy: the parties saw the issue, agreed on it, and didn’t bother writing it down.
And sometimes it’s strategic: the parties saw the issue, disagreed, and left it unresolved rather than blow up the deal.
Courts mostly treat the three alike, because from the outside they look alike.
Pre-testing changes the mix. Inadvertent gaps largely disappear for sophisticated parties, because the model surfaces the contingencies before signing. Contracts stay incomplete — they always will, compute isn’t free — but the silences that remain are far more likely to be chosen. And a chosen silence, under a choice of model regime, comes in exactly two flavors.
One is endorsement. The parties ran the model, saw what it predicted for the gap, and left the silence in place because they could live with the prediction. That silence now works like an incorporation by reference. The autonomy case for enforcing the model’s answer is about as strong as it gets: nobody is imposing a hypothetical bargain on these parties, because they saw the bargain and nodded.
The other is strategy. The parties ran the model, at least one of them hated the prediction, and raising it at the table would have surfaced a fight too expensive to finish. They signed anyway, fingers crossed against the contingency. Now the model’s answer is precisely what one party declined to agree to, and enforcing it looks less like respecting autonomy and more like awarding the pot to whoever liked the machine’s guess.3
The same silence thus produces two equal and opposite implications. That in turn means courts in a choice-of-model world get a new, cool, hard, job. The live question becomes which kind of silence is this. We argue that parties and courts will be channeled into discovery about the drafting process, i.e., what was run, what was predicted, what was raised and dropped. The result is that choice of model clauses will push lawyers to work on the hardest, most interesting set of problems in contract interpretation.
Limitations and Self-Congratulation
We understand that there are some gaps in our gap filling.4 Our hidden terms were manufactured, not natural: we hid terms the parties wrote, and terms the parties bother to write may leak into the rest of the document in ways terms they never wrote do not.5 The task was multiple choice, and recognition is easier than drafting. The correlated-error finding should give pause to anyone who imagines a panel of models as independent verification. We expect to continue to improve the draft before it’s published.
The paper is, limitations and all, optimistic. But it concludes with a more careful working through of the themes exploration of my recent agentic post. We consider a future where humans write fewer deals and agents write more. The equilibrium story here is complicated, and doctrine doesn’t have great tools to know what to do with information extracted from purely agentic contracts. That’s putting aside the model collapse issues that result when models are trained on artificial texts. Overall, as the paper concludes, we lack a way to talk interpretation of contracts that aren’t built on actual, human, social understandings. AI gap filling of AI-only deals is not a straightforward extension of our work. We thus sugest that the legal community have real intellectual work to do.
But, as we say in the paper, the stock of human-drafted paper is titanic. We’ve got a few years left to figure it out. And the claim that contracts contain gaps that can’t be filled with parties’ intended answers has structured contract scholarship for at least three generations. We think masking & measuring is a real advance for the field. And it could be deployed to benchmark models’ use for a bunch of other kinds of interpretative tasks. So, good enough for a day’s work!
The paper is here. Comments very welcome.
For some reason, this particular masking method reminds me of Kvothe playing Seek the Stone against himself in the Name of the Wind. Which reminds me that just as soon as this cycle of puffing scholarship-substacks is over, I’m writing about promissory estoppel claims against fantasy authors.
We sourced the contracts from PACER, which is paywalled and miserable to scrape, and therefore almost certainly absent from any model’s training data.
I once learned of an example where the parties disagreed about an eventuality, but decided to begin to perform anyway, both sides having signed a version of the contract with their own term, but not transmitting that version to the other. It turned out they settled, which is a crying shame because if they hadn’t it would’ve been a made-for-the-casebooks special.
I personally hope you appreciate how many bad puns I deleted from this substack before shipping it.
We think selection actually cuts the other way, but I won’t pretend the inference is airtight.





