open models · September 2026 · v1.0

The models you would have to run yourself — reviewed for legal work.

Every closed AI product on our tools index sends your matter documents to somebody else's infrastructure. These are the models you can put inside your own perimeter instead — and the questions that decide whether you should: what does the licence actually permit, how does the model behave when the evidence is thin, and what does it cost to run.

15models reviewed 5weighted criteria 0publisher payments On your hardwareby design
the five criteria

A model score is not a capability score. It is a question about whether a law firm can use it.

25%

Licence & commercial freedom

May a law firm run this model for client work, commercially, at scale, without restriction — and what happens if the publisher changes its mind. Includes redistribution, attribution, acceptable-use riders and any patent or indemnity position.

25%

Legal-task calibration

How well the model's default behaviour matches legal work: instruction following on drafting, extraction and summarisation; formatting discipline; usefulness of reasoning traces; and the tendency to over-assert.

20%

Evidence & abstention discipline

Whether it quotes what it was given faithfully, cites the passage rather than the vibe, and says 'the documents do not answer this' instead of inventing an answer. The single most important behaviour for legal use.

15%

Context & retrieval behaviour

Effective rather than advertised context: how it behaves with retrieved passages in the middle of a long prompt, degradation curves, and whether long-context claims survive a real matter file.

15%

Deployment reality

Hardware and quantisation practicality, throughput per pound, tooling ecosystem, supply-chain and provenance trust, and how quickly the family moves.

What the tiers mean in practice. A means we would run it on client material inside a firm's perimeter under a written policy, with evaluation evidence on file. B means pilot it on your own matters. C means keep it to non-confidential exploration. D means the licence or the provenance rules it out for client work however good the output looks.
15 of 15 models shown

Llama

Meta
78 /100 B
Frontier-class open weights

The default answer to 'can we run something capable ourselves?' — broad ecosystem support, predictable tooling, and a licence that a law firm's compliance team will want to read closely before signing anything.

Licence17
Legal-task calibration21
Evidence15
Context13
Deployment reality12
Sep 2026read the review →

Qwen

Alibaba Cloud
84 /100 B
Frontier-class open weights · Efficient and on-premise

The best licence-to-capability ratio in open weights: Apache-licensed for most sizes, strong instruction following, and a small-model range that makes a genuinely private pilot affordable for a mid-sized firm.

Licence23
Legal-task calibration20
Evidence14
Context14
Deployment reality13
Sep 2026read the review →

DeepSeek

DeepSeek
76 /100 B
Frontier-class open weights

Exceptional reasoning per pound under a permissive licence, with a provenance and hosting question that a law firm's client due-diligence conversation will eventually have to answer.

Licence22
Legal-task calibration19
Evidence13
Context12
Deployment reality10
Sep 2026read the review →

Mistral

Mistral AI
83 /100 B
Frontier-class open weights · Efficient and on-premise

The best licence-to-capability spread among mid-size European open weights, with small releases that run on one card and larger ones that need a server — and a licence position that has improved to match.

Licence21
Legal-task calibration20
Evidence15
Context14
Deployment reality13
Sep 2026read the review →

Gemma

Google DeepMind
81 /100 B
Efficient and on-premise

A small-model family that fits a single workstation, and whose current generation is Apache 2.0 — the most permissive licence position of any large publisher in this index.

Licence22
Legal-task calibration19
Evidence14
Context13
Deployment reality13
Sep 2026read the review →

Phi

Microsoft
71 /100 C
Efficient and on-premise

The cleanest licence in the index wrapped around the weakest legal behaviour we have assessed: excellent for narrow, verifiable tasks and short documents, unreliable the moment a matter file needs judgement.

Licence24
Legal-task calibration15
Evidence11
Context9
Deployment reality12
Sep 2026read the review →

Command

Cohere
71 /100 C
Frontier-class open weights

A retrieval-tuned family that cites what it was given — under a licence position that changes by release: the research weights are non-commercial, and the newest release is Apache 2.0.

Licence18
Legal-task calibration17
Evidence13
Context12
Deployment reality11
Sep 2026read the review →

GLM

Zhipu AI
79 /100 B
Frontier-class open weights

A permissively licensed mixture-of-experts family with strong bilingual capability and capable tool use, whose verbose, over-confident reasoning and heavy hardware footprint keep it a supervised tool rather than a system of record.

Licence22
Legal-task calibration19
Evidence14
Context13
Deployment reality11
Sep 2026read the review →

Kimi

Moonshot AI
80 /100 B
Frontier-class open weights

A long-context mixture-of-experts family that holds up across very large matter files and follows complex drafting briefs closely, with an attribution-flavoured licence and a hardware appetite that make it a deliberate purchase.

Licence20
Legal-task calibration19
Evidence15
Context15
Deployment reality11
Sep 2026read the review →

Granite

IBM
82 /100 B
Efficient and on-premise

The least complicated entry in this index: Apache-licensed small dense models, safety and document tooling around them, and an enterprise support posture — with capability that is solid rather than spectacular.

Licence23
Legal-task calibration18
Evidence14
Context12
Deployment reality15
Sep 2026read the review →

Nemotron

NVIDIA
77 /100 B
Frontier-class open weights

Frontier-class reasoning and the strongest summarisation we have measured from open weights, with disclosed training datasets — and a publisher-specific licence that a compliance team will want read rather than skimmed.

Licence15
Legal-task calibration21
Evidence15
Context13
Deployment reality13
Sep 2026read the review →

OLMo

Allen Institute for AI
75 /100 B
Fully transparent

The reference case for genuine openness — weights, data, code and logs — with the best abstention behaviour we have measured and a capability ceiling a fee-earner will find on hard legal synthesis.

Licence22
Legal-task calibration14
Evidence15
Context10
Deployment reality14
Sep 2026read the review →

Apertus

Swiss AI Initiative / ETH Zürich and EPFL
73 /100 B
Fully transparent

A narrow, genuinely excellent translator of legal material with a documented provenance story, Apache-licensed weights behind an acceptable-use gate, and weak summarisation and reasoning beyond its specialty.

Licence17
Legal-task calibration16
Evidence15
Context12
Deployment reality13
Sep 2026read the review →

gpt-oss

OpenAI
84 /100 B
Frontier-class open weights · Efficient and on-premise

Two permissively licensed open-weight releases sized so the larger fits on a single high-memory GPU — strong reasoning, configurable effort, no licence conditions to negotiate, and reasoning traces to keep away from clients.

Licence23
Legal-task calibration20
Evidence15
Context12
Deployment reality14
Sep 2026read the review →

Legal-domain open models

Various academic and domain groups
60 /100 C
Legal-domain and legal-tuned

A survey entry, not a recommendation: legal tuning buys terminology and register, not judgement or citation discipline — and most releases in this category are academic artefacts with unclear maintenance.

Licence16
Legal-task calibration13
Evidence12
Context8
Deployment reality11
Sep 2026read the review →