The models you would have to run yourself — reviewed for legal work.
Every closed AI product on our tools index sends your matter documents to somebody else's infrastructure. These are the models you can put inside your own perimeter instead — and the questions that decide whether you should: what does the licence actually permit, how does the model behave when the evidence is thin, and what does it cost to run.
A model score is not a capability score. It is a question about whether a law firm can use it.
Licence & commercial freedom
May a law firm run this model for client work, commercially, at scale, without restriction — and what happens if the publisher changes its mind. Includes redistribution, attribution, acceptable-use riders and any patent or indemnity position.
Legal-task calibration
How well the model's default behaviour matches legal work: instruction following on drafting, extraction and summarisation; formatting discipline; usefulness of reasoning traces; and the tendency to over-assert.
Evidence & abstention discipline
Whether it quotes what it was given faithfully, cites the passage rather than the vibe, and says 'the documents do not answer this' instead of inventing an answer. The single most important behaviour for legal use.
Context & retrieval behaviour
Effective rather than advertised context: how it behaves with retrieved passages in the middle of a long prompt, degradation curves, and whether long-context claims survive a real matter file.
Deployment reality
Hardware and quantisation practicality, throughput per pound, tooling ecosystem, supply-chain and provenance trust, and how quickly the family moves.
Llama
The default answer to 'can we run something capable ourselves?' — broad ecosystem support, predictable tooling, and a licence that a law firm's compliance team will want to read closely before signing anything.
Qwen
The best licence-to-capability ratio in open weights: Apache-licensed for most sizes, strong instruction following, and a small-model range that makes a genuinely private pilot affordable for a mid-sized firm.
DeepSeek
Exceptional reasoning per pound under a permissive licence, with a provenance and hosting question that a law firm's client due-diligence conversation will eventually have to answer.
Mistral
The best licence-to-capability spread among mid-size European open weights, with small releases that run on one card and larger ones that need a server — and a licence position that has improved to match.
Gemma
A small-model family that fits a single workstation, and whose current generation is Apache 2.0 — the most permissive licence position of any large publisher in this index.
Phi
The cleanest licence in the index wrapped around the weakest legal behaviour we have assessed: excellent for narrow, verifiable tasks and short documents, unreliable the moment a matter file needs judgement.
Command
A retrieval-tuned family that cites what it was given — under a licence position that changes by release: the research weights are non-commercial, and the newest release is Apache 2.0.
GLM
A permissively licensed mixture-of-experts family with strong bilingual capability and capable tool use, whose verbose, over-confident reasoning and heavy hardware footprint keep it a supervised tool rather than a system of record.
Kimi
A long-context mixture-of-experts family that holds up across very large matter files and follows complex drafting briefs closely, with an attribution-flavoured licence and a hardware appetite that make it a deliberate purchase.
Granite
The least complicated entry in this index: Apache-licensed small dense models, safety and document tooling around them, and an enterprise support posture — with capability that is solid rather than spectacular.
Nemotron
Frontier-class reasoning and the strongest summarisation we have measured from open weights, with disclosed training datasets — and a publisher-specific licence that a compliance team will want read rather than skimmed.
OLMo
The reference case for genuine openness — weights, data, code and logs — with the best abstention behaviour we have measured and a capability ceiling a fee-earner will find on hard legal synthesis.
Apertus
A narrow, genuinely excellent translator of legal material with a documented provenance story, Apache-licensed weights behind an acceptable-use gate, and weak summarisation and reasoning beyond its specialty.
gpt-oss
Two permissively licensed open-weight releases sized so the larger fits on a single high-memory GPU — strong reasoning, configurable effort, no licence conditions to negotiate, and reasoning traces to keep away from clients.
Legal-domain open models
A survey entry, not a recommendation: legal tuning buys terminology and register, not judgement or citation discipline — and most releases in this category are academic artefacts with unclear maintenance.