Mistral
The best licence-to-capability spread among mid-size European open weights, with small releases that run on one card and larger ones that need a server — and a licence position that has improved to match.
Tier B — conditional, and the most straightforward 'capable and permissively licensed' choice for a firm that wants European provenance. The mid-size releases serve comfortably from a single 48GB card with retrieval in parallel. What keeps it short of tier A is calibration rather than licence: it is a good drafter and a willing summariser, but it does not police its own uncertainty, so abstention has to be enforced in the harness. Confirm the licence on the exact release — this family has changed licence by tag before.
Specifications, as published
| Publisher | Mistral AI |
| Family | The Mistral line: small Apache-licensed dense models through to larger mixture-of-experts releases under a research licence |
| Parameters | Family spans small dense models suited to a single workstation through to much larger mixture-of-experts releases, with per-release parameter counts published on each model card |
| Context | Long context on the current generation; effective context under retrieval is materially lower than the advertised figure, and the smallest variants degrade earliest |
| Licence | Apache 2.0 or Modified MIT across the current generation, per the publishers' own cards — including the larger releases, which earlier guidance in this market still describes as research-licensed. This family has historically mixed licences by release, so the discipline stands: read the licence on the exact tag you deploy rather than on the family page or an older summary. |
| Weights | Downloadable, including quantised community builds |
| Release | Rolling releases across the line, with earlier generations still in production use |
| Licence posture | Apache 2.0 or Modified MIT across the current generation, per the publishers' own cards — including the larger releases, which earlier guidance in this market still describes as research-licensed. This family has historically mixed licences by release, so the discipline stands: read the licence on the exact tag you deploy rather than on the family page or an older summary. |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| Ministral 3 3B Instruct (2512) | 3.8B dense | 256k tokens | ~2.5GB at 4-bit quantisation; ~4GB at 8-bit | The small end of the current line — Apache 2.0, a 256k window and light enough to run beside a fee-earner's other work — which makes it a sensible first private deployment for one team, subject to 3B-class judgement. |
| Ministral 3 8B Instruct (2512) | 8.9B dense | 256k tokens | ~5GB at 4-bit quantisation; ~10GB at 8-bit | One 24GB card gives a small firm a credible drafting and extraction assistant under Apache 2.0, and the wide context window means fewer chunking decisions on long documents. |
| Ministral 3 14B Instruct (2512) | 14B dense | 256k tokens | ~8GB at 4-bit quantisation; ~15GB at 8-bit | The best quality-per-card release in this family for a firm of 10–50 fee-earners: it runs at 8-bit on a 48GB card, and its terseness suits partners who want a recommendation rather than an essay. |
| Magistral Small 2509 | 24B dense, reasoning-tuned | 128k tokens as published, with the publisher warning that quality may degrade past about 40k | ~14GB at 4-bit quantisation; ~26GB at 8-bit | The reasoning release a mid-sized firm can actually run — one 48GB card at 8-bit — provided you keep contexts inside the range the publisher says it holds, which rules out whole-bundle analysis on this release. |
| Mistral Small 4 119B (2603) | 119B total / 6.5B active, 128 experts with 4 active | 256k tokens | ~62GB at 4-bit quantisation; ~242GB for the published BF16 weights | The current-generation workhorse: an 80GB card runs it at 4-bit with only 6.5B parameters active per token, so firm-wide extraction and drafting become a throughput question rather than a capability one. |
| Mistral Small 4 119B NVFP4 (2603) | 119B total / 6.5B active, 128 experts with 4 active; the quantised-first publication | 256k tokens | ~60GB with the published NVFP4 build; the quantised checkpoint is the intended default on 80GB cards | The quantised-first build of the release above — same weights, published in NVFP4 for serving on one 80GB card — and the version we would pull first for a firm that wants the 119B class without a multi-GPU node. |
| Mistral Medium 3.5 128B | 128B dense | 256k tokens | ~65GB at 4-bit quantisation; ~128GB at 8-bit | A dense 128B flagship sized for an 80GB card or more and the release to choose if you want maximum capability on owned hardware, with the caveat that it is published under a Modified MIT licence carrying revenue-based exceptions. |
| Mistral Large 3 675B Instruct (2512) | 675B total / 41B active, granular mixture-of-experts | 256k tokens | ~340GB at 4-bit quantisation; ~680GB for the published FP8 weights | Capable and impractical for a 20-partner firm: it needs a multi-GPU node or a private tenancy, for a level of capability that legal workloads of this size rarely call on. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
Instruction obedience is Mistral's strongest legal-work trait, and the reason it appears in so many of the private deployments we have run. Hand it a drafting brief with a required structure — parties, recitals, operative clauses, a defined-terms block — and it produces all four, in order, in the register you asked for, at roughly the length you asked for. That sounds modest until you compare it with models that quietly drop the fourth element of a five-part instruction. Extraction behaves the same way: given a contract set and a fixed schema for dates, parties, obligations, notice periods and termination rights, the mid-size releases return clean rows that a paralegal validates rather than repairs, and the small Apache-licensed variants handle the same task at lower volume. Summarisation is where a fee-earner first notices the family's character: it is terse. A bundle summary comes back shorter than most models produce, with fewer hedges and less invented connective tissue, which senior lawyers tend to like and junior lawyers occasionally find too clipped to send on. The trade is that compression can read as certainty — a tight summary of an ambiguous document can look more decided than the document is. Formatting discipline is good but not absolute: it respects a schema, a heading hierarchy and a table when those are specified, and drifts back into prose when the instruction leaves room for interpretation. On reasoning, the dedicated reasoning releases behave differently enough from the general ones that they should be evaluated separately; the traces are useful for triage and for showing where an assumption entered, but they are not a substitute for reading the source. Over-assertion is the main behavioural risk. Ask a question the retrieved passages do not answer and Mistral will frequently answer it anyway, in the same confident register it uses when the passages do answer, which is precisely the failure that turns a tidy summary into a hallucinated file note. It responds unusually well to an explicit negative instruction — if the documents do not answer this, say so and stop — but the instruction has to be present on every call rather than assumed. Jurisdiction drift is specific to this publisher. Its training leans European, and in our tasks it writes English-law documents that read more naturally than most US-centric models manage; that is a real advantage. The drift runs the other way too: it imports civil-law habits of structure and phrasing into common-law drafting, particularly around definitions, agreements to negotiate and termination mechanics, and it tends to follow the user's framing rather than correct it. If your precedent set or your prompt has a civil-law tilt, the model leans into it instead of flagging the mismatch. Multilingual behaviour is a genuine strength and the reason several firms shortlist the family at all — French, German, Spanish and Italian material is handled with noticeably less register loss than in most open weights, and cross-language summarisation inside a single matter file is workable. What fee-earners report in the first week is consistent: quick to become useful, easy to steer with firm precedent, and prone to sounding more certain than the evidence supports. Budget supervision time for the second problem, not the first.
Quotation fidelity is good when the instruction is explicit about verbatim text and a page or paragraph reference, and it is the family's most reliable abstainer once a negative instruction is in the system prompt. Without that instruction it paraphrases fluently, and on adjacent-but-not-answering material it will bridge the gap rather than say the corpus is silent. Two habits improve results: ask for the quoted span before the comment on it, and require a separate list of questions the documents do not answer. Treat any quotation produced without a reference as unverified.
What we would use it for
- First-pass drafting from firm precedent, in a specified structure and register
- Extraction into a fixed schema across a contract set, with source references
- Chronology and bundle summarisation where brevity is wanted
- Cross-border and European-language matter support, including cross-language summaries
- A small Apache-licensed pilot running on one workstation inside the firm's perimeter
What to watch
- Licence varies by release — verify the tag, not the family page: current generation is Apache 2.0 or Modified MIT
- Does not refuse reliably on its own; abstention must be enforced by the harness
- European provenance is a selling point with some clients and a question with others — know which applies to you
- The larger releases need server-class hardware, so the cheap pilot and the firm-wide deployment are different decisions
- Rapid release cadence: pin the tag and re-run the evaluation set on any change
What it costs to run
We would serve Magistral Small 2509 at 8-bit on one 48GB card for a firm of 10–50 fee-earners: the 24B dense release keeps 8-bit quality, has unambiguous Apache 2.0 terms, and holds the register that fee-earners like, while Mistral Small 4 119B gives considerably higher throughput on the same card at 4-bit for extraction at volume. Latency is interactive for a single user and comfortable for a team on schema-bound work, provided you respect the context limits the publisher states rather than the headline window. We would not put the 675B flagship in a firm of this size at all: it needs a multi-GPU node and the operational overhead that comes with it.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $0.85 – $2.85 | $625 – $2,080 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $0.85 – $2.85 | $180 – $600 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $1.80 – $3.60 | $110 – $215 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~175 GB) | — | — | $25 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 48GB class — L40S / RTX A6000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
Mistral's 24B releases are the easiest case in this index for buying rather than renting: a 48GB-class workstation at an indicative $8,000–15,000 runs the recommended release at 8-bit, and the hardware stays useful across several generations of a family that ships often. Renting still wins for a first pilot, and it wins outright for the 675B flagship, where a multi-GPU server costs substantially more than the 80GB class and would sit idle for most of the day in a firm of this size. Both bands are indicative rather than quotes and ignore your own power, support and hardware-refresh terms.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Extraction and classification: temperature 0–0.1, top_p 0.8, fixed schema, no creative latitude. Drafting from precedent: 0.3–0.4, with the precedent supplied in the prompt rather than trusted to memory. Summarisation of a bundle: 0.1, with a source-reference-per-paragraph rule. Analysis and triage: 0.2 on the general releases; keep reasoning releases at 0–0.2, since raising temperature buys rambling rather than variety. Pin the exact release in your evaluation record — the line refreshes often, and a licence or behavioural change between releases is entirely possible.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single workstation (small dense, quantised) | Extraction, summarisation and drafting support for one team | The realistic entry point, and the point at which the Apache licence makes client work defensible |
| One GPU server (mid-size dense or smaller mixture-of-experts) | Firm-wide support with retrieval over precedents | Check whether the release you are running is Apache-licensed or research-licensed before it touches client material |
| Larger mixture-of-experts release, multi-GPU | Higher-throughput drafting and analysis | Justify the step up with evaluation evidence, and expect the licence constraint to be the binding one |
Who it suits
Firms wanting permissively licensed, European-hosted capability for drafting, extraction and summarisation, with a pilot that fits one 48GB card.
Firms that need a model to police its own uncertainty, or that will not pin releases and re-evaluate.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. Assessed from published documentation and our standard legal evaluation task set. |
| Sep 2026 | Licence re-scored from 15 to 21 after reading the current model cards: the releases in this generation are Apache 2.0 or Modified MIT, and the research-licensed framing in our first pass reflects older guidance rather than the current tags. Total 77 → 83, tier unchanged at B. |
Sources
- Mistral AI model releases and repository
- Mistral terms, including the research licence
- Open-weight licence landscape 2026 (Presenc AI)
- Probative Co: private AI in law — the 2026 guide