The vocabulary you need before you buy anything.
This section explains the machinery behind the index. The model reviews say what we think of a particular set of weights; these explainers say what the words mean — calibration, abstention, effective context, quantisation — and why they decide whether a deployment is safe for client work. Each one is written for the person who has to sign the policy and answer for it: a COLP, a managing partner, an IT lead fielding hard questions from the partnership. Our vocabulary, plainly, with the uncertainty left in.
What calibration means for legal work
Why do two models that score the same on a benchmark behave completely differently on a lease?
read the explainer →Abstention: why refusing to answer is a feature
How do we stop a model inventing an answer when the file does not contain one?
read the explainer →Temperature, top-p and why zero is not always the answer
Which settings should we standardise across the firm?
read the explainer →Advertised context versus effective context
It says a million tokens, so why did it miss the clause in the middle?
read the explainer →Retrieval, fine-tuning or a better prompt?
The model keeps getting our work wrong — what do we actually fix?
read the explainer →Reading an open-weight licence
What does 'open' actually let us do with client data?
read the explainer →Quantisation: what you lose when you shrink a model to fit
Can we run a serious model on hardware we already own?
read the explainer →Building an evaluation harness from your own matters
How do we know whether any of this works before we roll it out?
read the explainer →Each one answers a question a COLP actually asks.
They are written for the person who has to sign something, not the person building it. Where a concept is genuinely uncertain, we say so rather than inventing a rule.
Reading is not evaluation.
The only way to know whether a model abstains on your files, or quotes your precedents faithfully, is to run it on your own matters with your own fee-earners scoring the output.
Start with a readiness assessment →