AI Tools Index / Open models / Mistral
Mistral AI · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

Mistral

The Mistral line: small Apache-licensed dense models through to larger mixture-of-experts releases under a research licence

The best licence-to-capability spread among mid-size European open weights, with small releases that run on one card and larger ones that need a server — and a licence position that has improved to match.

Our verdict

Tier B — conditional, and the most straightforward 'capable and permissively licensed' choice for a firm that wants European provenance. The mid-size releases serve comfortably from a single 48GB card with retrieval in parallel. What keeps it short of tier A is calibration rather than licence: it is a good drafter and a willing summariser, but it does not police its own uncertainty, so abstention has to be enforced in the harness. Confirm the licence on the exact release — this family has changed licence by tag before.

Specifications, as published

PublisherMistral AI
FamilyThe Mistral line: small Apache-licensed dense models through to larger mixture-of-experts releases under a research licence
ParametersFamily spans small dense models suited to a single workstation through to much larger mixture-of-experts releases, with per-release parameter counts published on each model card
ContextLong context on the current generation; effective context under retrieval is materially lower than the advertised figure, and the smallest variants degrade earliest
LicenceApache 2.0 or Modified MIT across the current generation, per the publishers' own cards — including the larger releases, which earlier guidance in this market still describes as research-licensed. This family has historically mixed licences by release, so the discipline stands: read the licence on the exact tag you deploy rather than on the family page or an older summary.
WeightsDownloadable, including quantised community builds
ReleaseRolling releases across the line, with earlier generations still in production use
Licence postureApache 2.0 or Modified MIT across the current generation, per the publishers' own cards — including the larger releases, which earlier guidance in this market still describes as research-licensed. This family has historically mixed licences by release, so the discipline stands: read the licence on the exact tag you deploy rather than on the family page or an older summary.

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
Ministral 3 3B Instruct (2512)3.8B dense256k tokens~2.5GB at 4-bit quantisation; ~4GB at 8-bitThe small end of the current line — Apache 2.0, a 256k window and light enough to run beside a fee-earner's other work — which makes it a sensible first private deployment for one team, subject to 3B-class judgement.
Ministral 3 8B Instruct (2512)8.9B dense256k tokens~5GB at 4-bit quantisation; ~10GB at 8-bitOne 24GB card gives a small firm a credible drafting and extraction assistant under Apache 2.0, and the wide context window means fewer chunking decisions on long documents.
Ministral 3 14B Instruct (2512)14B dense256k tokens~8GB at 4-bit quantisation; ~15GB at 8-bitThe best quality-per-card release in this family for a firm of 10–50 fee-earners: it runs at 8-bit on a 48GB card, and its terseness suits partners who want a recommendation rather than an essay.
Magistral Small 250924B dense, reasoning-tuned128k tokens as published, with the publisher warning that quality may degrade past about 40k~14GB at 4-bit quantisation; ~26GB at 8-bitThe reasoning release a mid-sized firm can actually run — one 48GB card at 8-bit — provided you keep contexts inside the range the publisher says it holds, which rules out whole-bundle analysis on this release.
Mistral Small 4 119B (2603)119B total / 6.5B active, 128 experts with 4 active256k tokens~62GB at 4-bit quantisation; ~242GB for the published BF16 weightsThe current-generation workhorse: an 80GB card runs it at 4-bit with only 6.5B parameters active per token, so firm-wide extraction and drafting become a throughput question rather than a capability one.
Mistral Small 4 119B NVFP4 (2603)119B total / 6.5B active, 128 experts with 4 active; the quantised-first publication256k tokens~60GB with the published NVFP4 build; the quantised checkpoint is the intended default on 80GB cardsThe quantised-first build of the release above — same weights, published in NVFP4 for serving on one 80GB card — and the version we would pull first for a firm that wants the 119B class without a multi-GPU node.
Mistral Medium 3.5 128B128B dense256k tokens~65GB at 4-bit quantisation; ~128GB at 8-bitA dense 128B flagship sized for an 80GB card or more and the release to choose if you want maximum capability on owned hardware, with the caveat that it is published under a Modified MIT licence carrying revenue-based exceptions.
Mistral Large 3 675B Instruct (2512)675B total / 41B active, granular mixture-of-experts256k tokens~340GB at 4-bit quantisation; ~680GB for the published FP8 weightsCapable and impractical for a 20-partner firm: it needs a multi-GPU node or a private tenancy, for a level of capability that legal workloads of this size rarely call on.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

Instruction obedience is Mistral's strongest legal-work trait, and the reason it appears in so many of the private deployments we have run. Hand it a drafting brief with a required structure — parties, recitals, operative clauses, a defined-terms block — and it produces all four, in order, in the register you asked for, at roughly the length you asked for. That sounds modest until you compare it with models that quietly drop the fourth element of a five-part instruction. Extraction behaves the same way: given a contract set and a fixed schema for dates, parties, obligations, notice periods and termination rights, the mid-size releases return clean rows that a paralegal validates rather than repairs, and the small Apache-licensed variants handle the same task at lower volume. Summarisation is where a fee-earner first notices the family's character: it is terse. A bundle summary comes back shorter than most models produce, with fewer hedges and less invented connective tissue, which senior lawyers tend to like and junior lawyers occasionally find too clipped to send on. The trade is that compression can read as certainty — a tight summary of an ambiguous document can look more decided than the document is. Formatting discipline is good but not absolute: it respects a schema, a heading hierarchy and a table when those are specified, and drifts back into prose when the instruction leaves room for interpretation. On reasoning, the dedicated reasoning releases behave differently enough from the general ones that they should be evaluated separately; the traces are useful for triage and for showing where an assumption entered, but they are not a substitute for reading the source. Over-assertion is the main behavioural risk. Ask a question the retrieved passages do not answer and Mistral will frequently answer it anyway, in the same confident register it uses when the passages do answer, which is precisely the failure that turns a tidy summary into a hallucinated file note. It responds unusually well to an explicit negative instruction — if the documents do not answer this, say so and stop — but the instruction has to be present on every call rather than assumed. Jurisdiction drift is specific to this publisher. Its training leans European, and in our tasks it writes English-law documents that read more naturally than most US-centric models manage; that is a real advantage. The drift runs the other way too: it imports civil-law habits of structure and phrasing into common-law drafting, particularly around definitions, agreements to negotiate and termination mechanics, and it tends to follow the user's framing rather than correct it. If your precedent set or your prompt has a civil-law tilt, the model leans into it instead of flagging the mismatch. Multilingual behaviour is a genuine strength and the reason several firms shortlist the family at all — French, German, Spanish and Italian material is handled with noticeably less register loss than in most open weights, and cross-language summarisation inside a single matter file is workable. What fee-earners report in the first week is consistent: quick to become useful, easy to steer with firm precedent, and prone to sounding more certain than the evidence supports. Budget supervision time for the second problem, not the first.

evidence and abstention

Quotation fidelity is good when the instruction is explicit about verbatim text and a page or paragraph reference, and it is the family's most reliable abstainer once a negative instruction is in the system prompt. Without that instruction it paraphrases fluently, and on adjacent-but-not-answering material it will bridge the gap rather than say the corpus is silent. Two habits improve results: ask for the quoted span before the comment on it, and require a separate list of questions the documents do not answer. Treat any quotation produced without a reference as unverified.

What we would use it for

  • First-pass drafting from firm precedent, in a specified structure and register
  • Extraction into a fixed schema across a contract set, with source references
  • Chronology and bundle summarisation where brevity is wanted
  • Cross-border and European-language matter support, including cross-language summaries
  • A small Apache-licensed pilot running on one workstation inside the firm's perimeter

What to watch

  • Licence varies by release — verify the tag, not the family page: current generation is Apache 2.0 or Modified MIT
  • Does not refuse reliably on its own; abstention must be enforced by the harness
  • European provenance is a selling point with some clients and a question with others — know which applies to you
  • The larger releases need server-class hardware, so the cheap pilot and the firm-wide deployment are different decisions
  • Rapid release cadence: pin the tag and re-run the evaluation set on any change

What it costs to run

We would serve Magistral Small 2509 at 8-bit on one 48GB card for a firm of 10–50 fee-earners: the 24B dense release keeps 8-bit quality, has unambiguous Apache 2.0 terms, and holds the register that fee-earners like, while Mistral Small 4 119B gives considerably higher throughput on the same card at 4-bit for extraction at volume. Latency is interactive for a single user and comfortable for a team on schema-bound work, provided you respect the context limits the publisher states rather than the headline window. We would not put the 675B flagship in a firm of this size at all: it needs a multi-GPU node and the operational overhead that comes with it.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$0.85 – $2.85$625 – $2,080Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$0.85 – $2.85$180 – $600The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$1.80 – $3.60$110 – $215Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~175 GB)$25Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 48GB class — L40S / RTX A6000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

Mistral's 24B releases are the easiest case in this index for buying rather than renting: a 48GB-class workstation at an indicative $8,000–15,000 runs the recommended release at 8-bit, and the hardware stays useful across several generations of a family that ships often. Renting still wins for a first pilot, and it wins outright for the 675B flagship, where a multi-GPU server costs substantially more than the 80GB class and would sit idle for most of the day in a firm of this size. Both bands are indicative rather than quotes and ignore your own power, support and hardware-refresh terms.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Extraction and classification: temperature 0–0.1, top_p 0.8, fixed schema, no creative latitude. Drafting from precedent: 0.3–0.4, with the precedent supplied in the prompt rather than trusted to memory. Summarisation of a bundle: 0.1, with a source-reference-per-paragraph rule. Analysis and triage: 0.2 on the general releases; keep reasoning releases at 0–0.2, since raising temperature buys rambling rather than variety. Pin the exact release in your evaluation record — the line refreshes often, and a licence or behavioural change between releases is entirely possible.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
Single workstation (small dense, quantised)Extraction, summarisation and drafting support for one teamThe realistic entry point, and the point at which the Apache licence makes client work defensible
One GPU server (mid-size dense or smaller mixture-of-experts)Firm-wide support with retrieval over precedentsCheck whether the release you are running is Apache-licensed or research-licensed before it touches client material
Larger mixture-of-experts release, multi-GPUHigher-throughput drafting and analysisJustify the step up with evaluation evidence, and expect the licence constraint to be the binding one

Who it suits

Good fit

Firms wanting permissively licensed, European-hosted capability for drafting, extraction and summarisation, with a pilot that fits one 48GB card.

Poor fit

Firms that need a model to police its own uncertainty, or that will not pin releases and re-evaluate.

Review history

DateChange
Sep 2026First entry. Assessed from published documentation and our standard legal evaluation task set.
Sep 2026Licence re-scored from 15 to 21 after reading the current model cards: the releases in this generation are Apache 2.0 or Modified MIT, and the research-licensed framing in our first pass reflects older guidance rather than the current tags. Total 77 → 83, tier unchanged at B.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.