AI Tools Index / Open models / OLMo
Allen Institute for AI · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

OLMo

The fully open OLMo line — open weights, open training data, open training code and published training logs

The reference case for genuine openness — weights, data, code and logs — with the best abstention behaviour we have measured and a capability ceiling a fee-earner will find on hard legal synthesis.

Our verdict

Tier B — conditional, and admitted on the strength of its provenance rather than its output. OLMo is the model to deploy when the question is 'can you prove where this came from', and the model to avoid when the question is 'summarise this bundle'. Its abstention behaviour is the best in this index and is genuinely useful in a triage workflow; its comprehension of hard legal synthesis is not. Use it where answers are checkable, and keep a larger model for the reading.

Specifications, as published

PublisherAllen Institute for AI
FamilyThe fully open OLMo line — open weights, open training data, open training code and published training logs
ParametersMid-size dense models in a small and a larger size, each published in an instruct variant and a longer-reasoning 'think' variant
ContextPublished context is modest by current standards and is treated here at family level; the family is trained and tuned principally on English
LicenceApache 2.0 across the releases we assessed, with training data, training code and checkpoints published alongside the weights
WeightsDownloadable, with the full training pipeline and intermediate checkpoints released for inspection
ReleaseGenerational releases rather than continuous point updates, so the previous generation stays in use and capability gaps persist between versions
Licence postureApache 2.0 across the releases we assessed, with training data, training code and checkpoints published alongside the weights

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
Olmo-3.1-32B-Instruct32B dense, instruct-tuned (the 3.1 point update)65,536 tokens~64GB at BF16, ~20GB at 4-bitThe release we would serve for firm-wide work, and the one a firm can point to when a client asks how the model was trained — Apache 2.0 with the training data, code and logs published alongside the weights — although at BF16 it needs one 80GB card, which is a server rather than a workstation.
Olmo-3.1-32B-Think32B dense, long chain-of-thought variant of the 3.1 release65,536 tokens~64GB at BF16, ~20GB at 4-bitThe reasoning sibling of the 32B release — useful for triage and for showing where an assumption entered, but slower and chattier, so evaluate it separately before letting it near a supervised workflow.
Olmo-3-7B-Instruct7B dense, instruct-tuned on the Dolci post-training sets65,536 tokens~15GB at BF16, ~5GB at 4-bitThe workstation release of the family, and the one to start with if a firm of 10–50 fee-earners wants a genuinely open pipeline for extraction and verification on hardware it already owns.
Olmo-3-7B-Think7B dense, long chain-of-thought variant65,536 tokens~15GB at BF16, ~5GB at 4-bitA reasoning release at a size a single team can host, and the sensible way to test whether readable reasoning traces are worth anything in your review process before you spend memory on a larger model.
Olmo-Hybrid-Instruct-DPO-7B7B hybrid recurrent model — 75% of layers use gated DeltaNet heads instead of attention heads, instruct-tuned65,536 tokens~15GB at BF16, ~5GB at 4-bit, with materially smaller long-context memory than the equivalent transformerThe interesting 2026 release in this family: the publisher reports roughly 75% better throughput and memory at long context, which is exactly the constraint a firm hits on bundle-sized inputs, and it runs on a workstation.
OLMoE-1B-7B-0125-Instruct7B total with ~1B active per token (mixture of experts, 64 experts, 8 active)4,096 tokens as published — the shortest window in this review~14GB at BF16, ~4GB at 4-bitA genuinely cheap sparse release for classification and short-passage work on a small card or even a CPU, but its 4K window rules it out of document-level legal work without a retrieval layer in front of it.
OLMo-2-0425-1B-Instruct1B dense, instruct-tuned (previous generation)4,096 tokens as published~2.5GB at BF16, under 1GB at 4-bitSmall enough to run on a laptop or a CPU-only machine for demos and teaching, and the cheapest way for a firm of 10–50 fee-earners to rehearse an evaluation harness before committing a budget to anything larger.
OLMo-2-0325-32B-Instruct32B dense, instruct-tuned (previous generation, superseded by Olmo 3.1)4,096 tokens as published~64GB at BF16, ~20GB at 4-bitKept in circulation because the family moves in generations rather than continuous point updates, but the 4K window and the arrival of Olmo 3.1 make it a poor place to start a new deployment.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

OLMo's distinguishing feature is not what it can do but what you can prove about it, and that changes the compliance conversation rather than the legal output. A firm whose client asks where a model came from can answer with a dataset name, a training recipe, a code repository and logs — a combination no other entry in this index can match end to end. The output quality, however, is mid-size dense model quality, and the shortfall is most visible in exactly the tasks that matter most. Summarisation is OLMo's weakest legal behaviour in our tasks, and the independent Swiss legal evaluation work we cite reached the same conclusion at the larger size, where the OLMo entry was among the models that did not reliably produce a usable summary at all on the hardest summary task in that evaluation. That is a serious caveat, because summarisation is the default ask. Instruct OLMo to compress a judgment into a headnote and you get something that reads plausibly and drops load-bearing qualifications. Ask for a chronology from a bundle and it will attribute a date to the wrong document unless you require a citation for every line. Where a task has a checkable answer, the picture improves sharply: schedule building, field extraction and cross-checking a draft against a source document are reliable enough to supervise rather than to redo. What OLMo does unusually well is decline. In the same evaluation it was the most willing of the assessed models to say it did not know rather than guess, and we see the same tendency in our own tasks: given a question the supplied documents do not answer, it stops rather than bridging the gap. For legal work that is a genuinely valuable property and it is rarer than it should be. It is not, however, a substitute for capability — a model that abstains reliably and summarises poorly will give you a clean 'I don't know' on a question it should have been able to answer from the file. Extraction into a schema is serviceable. Field-by-field extraction from a single well-structured document is fine; across a set with inconsistent formatting it is mid-pack, and it will occasionally drop a field rather than flag it, which argues for a schema with an explicit null value and a validation step downstream. Drafting is plain and correct: it writes like a careful junior who has read the file and is anxious not to over-claim. Tone is neutral to a fault, which suits internal work and reads as thin in a client letter. Formatting discipline is good at the start of an answer and looser at the end. Structural instructions usually survive, but a long requested table will sometimes degrade into prose paragraphs, and a requested word limit is treated as advisory. Fee-earners describe this in week one as 'it forgets what I asked halfway through'. Reasoning traces in the think variants are long and legible, useful for training and audit and tedious in a production interface. Week one therefore has two halves. The first is real enthusiasm about provenance: being able to show a client, an auditor or a regulator the exact data and code behind a model is a compliance advantage no closed alternative offers. The second is the discovery that it needs closer supervision than a larger general model, especially on summarisation, and that the honest way to use it is on tasks with a checkable answer rather than on synthesis. Firms that plan for that split get value; firms that treat it as a cheaper substitute for a frontier model conclude in a fortnight that open weights do not work.

evidence and abstention

OLMo is the best-abstaining model we have assessed: asked a question its documents do not answer, it says so, and it takes little prompting to get there. Abstention is not the whole of evidence discipline, though. Quotation fidelity is good when it is asked to quote and cite, and markedly worse when it is asked to summarise — it will paraphrase a clause into a different obligation without appearing to notice. The prompting that helps is mechanical: a rule that every asserted fact carries a document and page, an explicit null value instead of an omission, and a limit forbidding reference to anything outside the supplied material.

What we would use it for

  • Tasks with a checkable answer: field extraction, date and party verification, schedule building
  • Deployments where a client or auditor asks to see the training data and code
  • Abstention-first triage — routing to a fee-earner the questions the documents do not answer
  • Internal research and training material where provenance matters more than polish
  • An evaluation baseline: a model whose behaviour you can reason about end to end

What to watch

  • Weak summarisation for its size — the default legal ask is its weakest task
  • English-only focus; not a multilingual option
  • Modest context by current standards, and weaker than the figure suggests once retrieval is in play
  • Generational rather than continuous releases, so gaps persist between versions
  • Openness is not maintenance — check the release cadence before you build a workflow on it

What it costs to run

The 32B releases are the ones worth serving for firm-wide work, and at BF16 they need one 80GB card — which is the point of this family, because quantisation costs it more summarisation quality than the memory it saves. Expect a handful of concurrent users on that card with the 65,536-token window in play and second-scale output on bundle extraction, slower again on the Think variants that spend their time in the reasoning trace. If only one team needs triage and extraction, run Olmo 3.1 7B Instruct on a 24GB workstation and leave the 32B for the workload that justifies it.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$1.58 – $5.24$1,150 – $3,820Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$1.58 – $5.24$330 – $1,100The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$3.30 – $6.90$200 – $415Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~250 GB)$40Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 80GB class — A100 80GB / H100. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

Rent while you evaluate, for a reason specific to this family: it is released generationally, so hardware bought against today's checkpoint will outlive the release you benchmarked, and rental lets you move on without writing anything off. If you do buy, the 80GB server class is indicatively $25,000–60,000, and the honest justification is a workload that runs for most of every working day. The OLMo line's argument is transparency rather than economy — and these figures are indicative, so obtain quotes for your own configuration.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Extraction and verification: temperature 0, top_p 0.8, strict schema with explicit nulls. Chronology building: 0.1 with a citation required per line. Drafting: 0.3. Summarisation: 0.1, and check the omissions against the source rather than the output. The think variants should be run with their own recommended settings rather than borrowed ones, and you should record which variant you used — the instruct and think models are different systems on legal triage. Pin the exact checkpoint, not the family name.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
Single workstation, small variant at low precisionExtraction, verification and triage for one teamCheap to run; the think variants are slow rather than large, so budget time rather than memory
One GPU server, larger variant in full precisionFirm-wide extraction with retrieval over precedentsFull precision is worth the memory here — quantisation degrades the family's already-limited summarisation more than it degrades extraction
CPU or a modest accelerator, small variantEvaluation, teaching and internal demonstrationsUseful precisely because it is inexpensive enough to run repeatedly while you build and refine your own task set

Who it suits

Good fit

Firms with clients or auditors who demand full provenance, and teams building evaluation pipelines where an abstention-first model is an asset.

Poor fit

Summarisation-heavy work, multilingual matters, or any task where the model must synthesise rather than extract.

Review history

DateChange
Sep 2026First entry.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.