How do we know whether any of this works before we roll it out?
An evaluation harness is a fixed set of tasks from your own matters, with expected outcomes written down and pass-or-fail scoring, run identically against every candidate. It is the only evidence that survives being questioned.
What a harness is, and what it is not
A harness is a rule that you will judge this thing the same way every time. Concretely: a fixed set of tasks drawn from your own matters, the documents those tasks run against, an expected outcome written down before you run anything, a scoring rule that produces pass or fail rather than an impression, and a record of the results including which system produced them. It is not a benchmark, and it is not an attempt to rank models in the abstract. It is your benchmark, built from your work, and its purpose is comparative and local: does this candidate, at this version, with these settings, do these tasks well enough for the supervision model we can actually staff. It also does not require software. A shared spreadsheet, a folder of written task definitions and two people with an hour a week is a functioning harness, and it will outperform a sophisticated tool nobody maintains. The point of writing it down is not process for its own sake. It is that the moment somebody asks why the firm approved this tool, or why it stopped using it, the answer is a dated sheet of tasks and results rather than a recollection that the demonstration looked good.
Choosing tasks from your own matters
Draw tasks from closed matters, or from live matters where the client's engagement terms permit and the material is handled correctly. Redact where you need to; the task is to test behaviour, not to test the model's knowledge of a particular client. Cover the failure modes rather than the highlights. Six to ten tasks is a realistic starting set: extraction of dates, parties, obligations and termination rights into a fixed schema; a chronology from a bundle with the dates checked against source; a quote-fidelity task where the output must reproduce operative words exactly; a question your documents genuinely cannot answer; a clause-conflict question that needs two passages read together; a drafting-from-precedent letter reviewed by the fee-earner who owns the precedent; and one task deliberately designed to invite over-assertion, such as a question whose premise the documents do not support. Include realistic distractors. A bundle with an adjacent clause that does not answer the question, and a schedule that holds the operative provision while the body cross-refers to it, will separate candidates far better than a tidy two-page test document. The tasks your fee-earners most dislike doing are the right place to look, because that is where adoption will be directed and where a failure will hurt.
Scoring behaviour, not vibes
Every task needs a pass condition that two different reviewers would apply the same way. Count facts rather than impressions. Every date in the output matched the source: yes or no. Every quoted passage appears verbatim in the document: yes or no. The model said the documents do not answer the question: yes or no. The output validated against the schema: yes or no. The letter kept the firm's register and did not import another jurisdiction's conventions: usually a judgement, so use two reviewers and record where they disagreed. The reason for such pedantry is auditability. A partner's considered view that the output seemed good is not evidence and cannot be re-examined later, and it cannot detect a change. A count of nineteen correct dates out of twenty can be compared with the same count run against a different model, or the same model after a version update, and the comparison means something. It also surfaces the trade-offs that matter. One candidate may be better at extraction and worse at abstention; another the reverse. That is not a ranking problem to be resolved by an overall score. It is a decision about which tasks you will deploy for, and the harness gives you the material to make it.
Running it honestly
Pin everything. A result is meaningless without the model version, the quantisation or build identifier, the tool version, the prompt, the sampling settings and the retrieval configuration. Two runs of the same task set on differently configured systems are two different experiments. Run every candidate on the same tasks, and where the work allows, blind the reviewer: strip the model name from the output before marking, so the person scoring does not know which system they are reading. The effect of knowing is not small, and it is the easiest bias to remove. Keep the raw output exactly as generated, before any human tidying, with a timestamp and the configuration used. A scored spreadsheet is not enough six months later: when somebody asks what the model actually said, the spreadsheet holds a number and the raw file holds the evidence. Store the outputs with the same care as the inputs, since the document sets are often confidential even when redacted. Finally, run the harness deliberately rather than opportunistically. Score the unanswerable questions first and by hand, because they are the ones people skip, and they are the ones that fail silently in production.
What the results are for
The harness is not a one-off procurement chore. It is the evidence base for three things the firm needs anyway. The first is the AI policy: which model versions are approved, for which tasks, under which configuration, and what the reviewer must check. A policy that says what the harness showed, in the firm's own words, is a document a COLP can defend to a client, an insurer or a regulator. Second, the supervision model. Where the harness shows a task is reliable, the control can be a spot check. Where it shows a task is unreliable, the answer is usually not to abandon the tool but to require a specific check — every quoted clause verified against source, every date confirmed, every assertion of absence re-scoped to what was searched. That is a supervision instruction, which is a thing supervisors already know how to write. Third, the decision to stop. Where the harness shows failure that instruction, retrieval and configuration cannot fix, the honest outcome is that the task stays manual. Recording that is not a failure of the project; it is the project working, and it is far cheaper than discovering the same thing through a client complaint. Re-run when the model version, build or retrieval configuration changes, and keep every set of results with its date. Behaviour moves; the record is what lets you see that it has.
What to watch
- Testing on vendor demonstration documents where the task works by construction
- Scoring with a one-to-ten impression that no two reviewers would apply in the same way
- Results kept without the model version, build, prompt or retrieval settings, so nothing can be reproduced
- Keeping only the human-edited output, so the model's actual errors are lost before they can be counted
- Running the harness once at pilot and never again after a model or tooling update
- Choosing tasks that flatter the tool instead of the tasks your fee-earners most dislike doing
Build a starting harness this quarter: six tasks drawn from closed matters, each with the source documents, an expected outcome written before you run it, and a pass condition that is a count rather than an opinion. Include at least two unanswerable questions, one two-passage question and one task designed to invite over-assertion. Run the approved configuration and one alternative, blind the outputs before marking, and keep the raw outputs with the date, version and settings. Then re-run on any model, build or retrieval change, and cite the results when you write or revise the AI policy. Two people, an hour a week, and a shared folder is a sufficient infrastructure for a firm of any size in this range.
Related
- How we score legal AI — our method
- AI audit and governance review
- AI training for fee-earners and supervisors
- Submit a tool or correction