AI Tools Index / Open models / Granite
IBM · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

Granite

The Granite model line — small dense language models, earlier mixture-of-experts variants, and the surrounding safety, embedding and document-conversion releases

The least complicated entry in this index: Apache-licensed small dense models, safety and document tooling around them, and an enterprise support posture — with capability that is solid rather than spectacular.

Our verdict

Tier B — conditional, and the most straightforward licence in this index: Apache 2.0, no thresholds, no acceptable-use riders to negotiate, with a useful family of safety and document tooling around it. It is not the most capable model here and it will not win a drafting beauty contest. It is the model we would use to prove a concept cheaply, on a workstation, against your own precedents, before spending on anything larger.

Specifications, as published

PublisherIBM
FamilyThe Granite model line — small dense language models, earlier mixture-of-experts variants, and the surrounding safety, embedding and document-conversion releases
ParametersFamily spans small dense language models of a few billion parameters up to roughly 30B, alongside earlier mixture-of-experts variants carrying more total parameters with a fraction active per token
ContextLong-context extension is published for the current dense releases, advertised in the hundreds of thousands of tokens; effective context under retrieval is materially lower than the published figure
LicenceApache 2.0 across the current language releases — permissive, with no user threshold, no acceptable-use rider on the weights and no naming obligations beyond attribution practice
WeightsDownloadable, with publisher-published quantised formats and a broad set of community builds
ReleaseFast point-release cadence — several minor generations within a year — so pinning the exact build matters more than in most families
Licence postureApache 2.0 across the current language releases — permissive, with no user threshold, no acceptable-use rider on the weights and no naming obligations beyond attribution practice

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
Granite-4.0-1B1B dense128K tokens~1GB at 4-bit, ~2.5GB at 8-bitThe edge release, small enough to run without a GPU, and useful only for tagging, routing and screening — not for reading a document and not for deciding anything. It is the floor of a family whose appeal is that the floor and the ceiling both run on hardware a firm of 10–50 fee-earners already owns.
Granite-4.0-H-Micro3B total / 3B active (hybrid Mamba-2 and attention, no expert routing)128K tokens~2GB at 4-bit, ~4GB at 8-bitA 3B long-context instruct release explicitly built for constrained hardware, with the hybrid architecture chosen to keep memory low on modest silicon. Realistically runnable by a firm of 10–50 fee-earners on a small GPU for summarising single documents and schema-bound extraction, with the usual small-model ceiling on judgement.
Granite-4.2-3B3B dense (3.66B parameters in the released weights)128K natively, extensible to 512K~3GB at 4-bit, ~5GB at 8-bitThe compact reasoning release of the current generation, with a built-in thinking mode and selectable effort levels so latency can be traded against depth per call. Runs on a laptop-adjacent GPU, and is the smallest release in this table we would trust with anything that benefits from step-by-step working — capable for its size, and not a substitute for a large model on hard legal reasoning.
Granite-4.0-H-Tiny7B total / 1B active (MoE with shared experts, hybrid Mamba-2 and attention)128K tokens~4GB at 4-bit, ~8GB at 8-bitThe smallest genuine mixture-of-experts release here and a deliberate constrained-hardware design: 7B of weights to store with only 1B active per token, so it is quick on modest silicon. The knowledge stored is correspondingly thin, which matters more in legal work than speed does; realistically runnable on one small card by a mid-sized firm.
Granite-4.2-8B8B dense (8.79B parameters in the released weights)128K natively, extensible to 512K~6GB at 4-bit, ~11GB at 8-bitThe mid-size release of the current generation and the one we would recommend as a firm's default: Apache 2.0, a very long context extension, native reasoning, and a serving footprint small enough to leave room for a retrieval index and long context on one workstation card. Realistically runnable by any firm of 10–50 fee-earners, and the release most of them should start with.
Granite-4.0-H-Small32B total / 9B active (MoE with shared experts, hybrid Mamba-2 and attention)128K tokens~18GB at 4-bit, ~33GB at 8-bitThe largest mixture-of-experts release in the family: 32B of weights with 9B active per token, so it fits a 48GB-class card at 8-bit or a 24GB card at 4-bit with limited context headroom. The best quality per pound here for a firm with real concurrency, and a reminder that this family's MoE variants were designed to stay on-premise.
Granite-4.2-30B30B dense (29.3B parameters in the released weights)128K natively, extensible to 512K~17GB at 4-bit, ~31GB at 8-bitThe largest and most capable release in the current generation, and the one that shows how far this family's ceiling sits from a data centre: even the flagship fits a single card. A UK firm of 10–50 fee-earners can realistically run it, though a 30B dense model will not match a 350B mixture-of-experts release on a hard analytical question — nothing in this family is beyond a mid-sized firm's hardware, which is the opposite of every other family in this index.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

Granite's enterprise reputation is earned by obedience rather than brilliance, and obedience is what a fee-earner notices first on legal tasks. Give it a house format — clause heading, then a numbered summary, then a table of dates — and it reproduces that format on the tenth document as faithfully as on the first. For a firm with existing precedents, a style guide and a supervision model, that predictability is worth more than a point of raw reasoning. Extraction into a fixed schema is where it earns its place in a pilot: run over licence agreements and facility letters asking for parties, dates, notice periods, termination triggers and governing law as strict JSON, the output needs validation rather than repair. It rarely invents a field, it respects an 'unknown' convention if you define one, and it tolerates the layout damage that survives a scan-and-OCR pipeline better than a model of its size has any right to. Drafting is where the ceiling shows. Granite writes clean, plain, businesslike English and is genuinely good at the small courtesies of tone — 'we write further to your letter of 3 March' rather than the florid alternative. But it is a conservative drafter: asked for a first-pass letter of advice it produces something correct, slightly flat and shorter than the matter justifies, and a partner will add the argumentative spine it did not supply. Where it does noticeably well is the unglamorous drafting that consumes trainee time — attendance notes, file notes, covering emails, witness-summary tables, chronologies in a fixed template. Fee-earners stop rewriting those within a week, which is the honest test of whether an open model has paid for itself. Summarisation fidelity is good but conditional. Ask for a summary of a bundle and you get a competent one. Ask for a summary that uses only phrases appearing in the documents and it stays much closer to the source than most of this index, because the instruction-following is doing the work rather than the model's judgement. Without that constraint it compresses aggressively and drops the qualifications that make a legal summary safe — the 'subject to', the 'without prejudice', the sentence that turns a deadline into a target. Its reasoning traces, where a release offers them, are short and sober: less useful for audit than a large reasoning model's monologue, and also less prone to the confident dead end. Over-assertion is real but smaller than its peers'. Granite is more likely than the generalists above it to say that a document does not address something, particularly when the system prompt defines what 'not addressed' looks like. It also hedges in language a client may read as evasive — 'it may be that' — which is safe for review and irritating in a client-facing draft. In week one, three things strike a fee-earner. First, that a model of this size can do their document chores at all, which changes the economics of a pilot. Second, that the family around it — content-safety classifiers, embedding models, a document-conversion tool for PDFs and scans — is more useful than the language model alone, because retrieval quality dominates output quality in every legal workflow we have tested. Third, that the answers are unexciting, which is exactly the property you want when the alternative is eloquence you have to verify.

evidence and abstention

Granite quotes faithfully when the instruction is specific: name the clause, quote it verbatim, then give its page. Given that, we see few silent edits and no merging of passages. Left to summarise without the constraint it paraphrases, and paraphrase is where legal meaning drifts. Abstention is above average for its size — asked whether a bundle answers a question it will often say it does not — and improves further when the system prompt defines an explicit refusal string and supplies one worked example. Its evidence weakness is negative: it is a poor reporter of its own uncertainty, so it under-flags the moments when it has drifted from the supplied documents to general knowledge. Require a quotation for every assertion and that largely disappears.

What we would use it for

  • High-volume extraction into a fixed schema — parties, dates, notice and termination provisions
  • Chronologies, attendance notes and document schedules in a firm template
  • First-pass summaries where the house rule is quote-or-omit
  • Document conversion and OCR pipelines feeding a retrieval layer, using the publisher's own tooling
  • A low-cost first private deployment on a single workstation, before committing to larger hardware

What to watch

  • Conservative drafting that under-argues compared with a frontier-class model
  • Aggressive compression when no source-quoting rule is set
  • Hedging that reads as evasive in client-facing correspondence
  • Under-reporting of its own uncertainty when it drifts off the supplied documents
  • Rapid point releases — pin the exact build and re-run your task set after each bump

What it costs to run

One 24GB card serving Granite-4.2-8B at 8-bit is the configuration we would recommend to a firm of 10–50 fee-earners, and it is the cheapest credible private deployment in this index. That covers extraction, summarisation and drafting support for a team at concurrency that suits asynchronous work, with the retrieval index resident on the same machine, and the 30B release can be swapped in at 4-bit on the same card when a task needs more capacity. Latency is good on realistic prompts because this is a dense 8B model rather than a large mixture-of-experts, so the practical ceiling is quality rather than speed. The moment the model visibly fails your task set is the moment to move up a size, not to buy more GPUs.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$0.58 – $1.65$425 – $1,205Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$0.58 – $1.65$125 – $345The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$1.03 – $1.95$60 – $115Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~150 GB)$20Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 24GB class — RTX 4090 / L4 / A5000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

This is the family where buying genuinely competes with renting: a 24GB-class workstation is an indicative $3,000–6,000, and a firm that keeps the model resident through the working day will usually recover that against the equivalent rented time within a year of full-time use. Renting a 24GB-class GPU is still the right way to run the pilot, because evaluation work is bursty and the pilot is what tells you which release is adequate. Even the mixture-of-experts and 30B releases run on a 48GB-class workstation, an indicative $8,000–15,000, so a firm can grow within this family without ever approaching server hardware. All of these figures are indicative.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Extraction and classification: temperature 0, top_p 0.8, a strict JSON schema and an explicit 'unknown' value. Summarisation under a quote-or-omit rule: 0.1. Drafting file notes and correspondence: 0.3. Do not push temperature above 0.5 anywhere in a legal workflow — this family does not buy quality with it, it buys drift. Pin the exact release in your evaluation record; the line moves in minor versions rather than major ones, and a minor bump has changed our extraction results before.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
Single workstation, small dense variant at low precisionExtraction, summarisation and file-note drafting for one teamThe lowest-cost credible private pilot we have assessed; a 4-bit build leaves room for a small retrieval index on the same machine
One GPU server, mid-size dense variantFirm-wide extraction and drafting support with retrieval over precedentsGood throughput per pound for schema-bound work; quantise the weights but keep the retrieval index in full precision
Larger deployment, or an earlier mixture-of-experts variantHigher volume and concurrent usersOnly justify the step up once the smaller model has visibly failed your task set — MoE quantisation has cost us more calibration than dense quantisation in our testing

Who it suits

Good fit

Firms that want a defensible, low-cost private deployment first — extraction, chronologies, document chores — with the least licence risk and the least hardware.

Poor fit

Matters that turn on persuasive drafting or multi-document reasoning, where a frontier-class model justifies the extra infrastructure.

Review history

DateChange
Sep 2026First entry. Assessed from published documentation and our standard legal evaluation task set.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.