Why do two models that score the same on a benchmark behave completely differently on a lease?
A benchmark score averages correct answers over somebody else's questions. Calibration is about behaviour under uncertainty — whether the model quotes, holds a schema and declines when the file is silent. Test it on your own matters.
What a public score actually measures
A published benchmark is a fixed set of questions with known correct answers. The model gets a score for producing those answers. That is a useful signal about capability, and it is close to useless as a prediction of how the same model will behave on your lease. Three things get lost in the average. First, the questions are usually short, well-formed and self-contained, where a real lease is long, cross-referenced and full of defined terms. Second, the score is an average over many kinds of task, so two models can land on the same number with opposite strengths — one excellent at extraction and careless at drafting, the other the reverse. Third, and most importantly, a benchmark rewards the right answer and rarely rewards the right refusal. A model that guesses on an answerable question and a model that guesses where the answer does not exist can share a score, and only one of them is safe in a matter file. So when two models score the same, you have learned that they are roughly comparable on tasks that are not yours. The behaviour on your lease is a separate question, and it is answered by testing rather than by reading.
Calibration is a statement about confidence
Underneath the interface, a language model produces a probability distribution over possible next words and picks one. Calibration is the relationship between the confidence implied by its behaviour and the chance that it is right. A well-calibrated system is uncertain when the evidence is thin; a poorly calibrated one sounds exactly as assured when it is inventing as when it is quoting. Legal work punishes miscalibration more than most fields, because the output goes into a file note and the file note goes to a client. Instruction tuning makes this worse rather than better. Models are tuned against human judgements of helpfulness, and a fluent, complete answer is rated as more helpful than a refusal — so the training signal pushes toward the confident answer even when the retrieved material does not support one. That is why over-assertion is the failure you should expect by default, and why it has to be managed with instruction and structure rather than hope. It is also why the reviews in this index score evidence and abstention discipline separately from raw capability, and weight it heavily: the model most likely to help your fee-earners is not always the model most likely to admit it does not know.
The lease that separates two identical scores
Take a lease of a retail unit with a tenant break right, a landlord option to review the rent, a schedule of notice provisions and a side letter. Ask a model a straightforward question: what are the tenant's exit rights, and what has to happen before they can be exercised? One model returns a short table: clause reference, right, notice period, conditions. It quotes the operative words. It notes that the notice provision is in the schedule rather than the body, and it flags that the side letter varies one of the dates. A reviewer can check that against the lease in a few minutes. The other model returns a fluent paragraph, also correct in outline, in which the tenant break has absorbed the notice mechanics of the landlord option and the side letter has vanished. Nothing in either model's published score distinguishes these two outputs, and both would pass a general reading-comprehension test. What separates them is not knowledge. It is quoting behaviour, schema discipline and attention to where a document keeps its operative provisions — three habits you can measure directly, on your own documents, in an afternoon. That measurement is the whole of calibration practice. Everything else is inference.
What to measure instead of a score
Build a small set of behaviours and count them. Quote fidelity: did the quoted words appear in the source, character for character. Schema adherence: did the output validate against a fixed structure of dates, parties, obligations and termination rights. Abstention: on questions your documents genuinely do not answer, did the model say so. Register: in a drafting task, did it drift into another jurisdiction's conventions. Stability: does the same task produce the same shape of answer twice, or does it wander. None of these need a research team. Six tasks from two closed matters, an expected outcome written down before you run them, and a partner or senior associate who marks the output against the source will tell you more than any leaderboard. Record the model version, because behaviour moves between point releases and a name on a tin is not a version number. The value is comparative and local. You are not trying to establish which model is best in the world; you are establishing which one you would let near a matter, and for which task. Those are different questions, and only the second one belongs in a policy.
Writing it into supervision
Calibration behaviour belongs in two places. The first is your AI policy: which model version is approved, for which task types, under which configuration, and what the reviewer must check. The second is the supervisory record, where the wording matters more than people expect. A file note that says the material was reviewed with AI assistance tells a reader nothing and protects nobody. A note that says which tool was used, which document set it saw, and that dates, party names and every quoted clause were checked against the source tells a reader exactly what was and was not verified. Keep the raw output alongside the reviewed version, at least while the matter is live. It is the difference between being able to answer a question six months later and reconstructing it from memory. And be honest with yourself about what the evidence supports. Testing shows that a model behaved in a certain way on a defined task set on a particular date. It does not establish that the model is safe, and it does not transfer to a version you did not test. State the limits in the record the same way we state them here.
What to watch
- Assuming a strong published benchmark score transfers to your lease, your precedent or your disclosure set
- Reading a summary of a model as a statement about its extraction and quotation discipline
- Treating a model as careful because it sounded cautious in a vendor demonstration
- Letting a long, confident reasoning trace stand in for a verified answer
- Comparing two models on the same task with different prompts or sampling settings, then attributing the difference to the model
- Never re-testing after a point release, when behaviour has moved but the name on the tin has not
Pick six tasks you actually do — one extraction into a schema, one chronology, one clause-conflict question, one drafting-from-precedent letter, one summarisation, and one question your documents cannot answer. Write down the expected outcome before you run anything. Run the approved model version, mark the outputs against the source, and keep the raw output and the version number in one place. Repeat when the model version changes or a new model arrives. An afternoon of this, done once, gives you a defensible basis for the task list in your AI policy and a supervision note you can stand behind. Treat the result as task-specific evidence, not as a verdict on the model.