Granite
The least complicated entry in this index: Apache-licensed small dense models, safety and document tooling around them, and an enterprise support posture — with capability that is solid rather than spectacular.
Tier B — conditional, and the most straightforward licence in this index: Apache 2.0, no thresholds, no acceptable-use riders to negotiate, with a useful family of safety and document tooling around it. It is not the most capable model here and it will not win a drafting beauty contest. It is the model we would use to prove a concept cheaply, on a workstation, against your own precedents, before spending on anything larger.
Specifications, as published
| Publisher | IBM |
| Family | The Granite model line — small dense language models, earlier mixture-of-experts variants, and the surrounding safety, embedding and document-conversion releases |
| Parameters | Family spans small dense language models of a few billion parameters up to roughly 30B, alongside earlier mixture-of-experts variants carrying more total parameters with a fraction active per token |
| Context | Long-context extension is published for the current dense releases, advertised in the hundreds of thousands of tokens; effective context under retrieval is materially lower than the published figure |
| Licence | Apache 2.0 across the current language releases — permissive, with no user threshold, no acceptable-use rider on the weights and no naming obligations beyond attribution practice |
| Weights | Downloadable, with publisher-published quantised formats and a broad set of community builds |
| Release | Fast point-release cadence — several minor generations within a year — so pinning the exact build matters more than in most families |
| Licence posture | Apache 2.0 across the current language releases — permissive, with no user threshold, no acceptable-use rider on the weights and no naming obligations beyond attribution practice |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| Granite-4.0-1B | 1B dense | 128K tokens | ~1GB at 4-bit, ~2.5GB at 8-bit | The edge release, small enough to run without a GPU, and useful only for tagging, routing and screening — not for reading a document and not for deciding anything. It is the floor of a family whose appeal is that the floor and the ceiling both run on hardware a firm of 10–50 fee-earners already owns. |
| Granite-4.0-H-Micro | 3B total / 3B active (hybrid Mamba-2 and attention, no expert routing) | 128K tokens | ~2GB at 4-bit, ~4GB at 8-bit | A 3B long-context instruct release explicitly built for constrained hardware, with the hybrid architecture chosen to keep memory low on modest silicon. Realistically runnable by a firm of 10–50 fee-earners on a small GPU for summarising single documents and schema-bound extraction, with the usual small-model ceiling on judgement. |
| Granite-4.2-3B | 3B dense (3.66B parameters in the released weights) | 128K natively, extensible to 512K | ~3GB at 4-bit, ~5GB at 8-bit | The compact reasoning release of the current generation, with a built-in thinking mode and selectable effort levels so latency can be traded against depth per call. Runs on a laptop-adjacent GPU, and is the smallest release in this table we would trust with anything that benefits from step-by-step working — capable for its size, and not a substitute for a large model on hard legal reasoning. |
| Granite-4.0-H-Tiny | 7B total / 1B active (MoE with shared experts, hybrid Mamba-2 and attention) | 128K tokens | ~4GB at 4-bit, ~8GB at 8-bit | The smallest genuine mixture-of-experts release here and a deliberate constrained-hardware design: 7B of weights to store with only 1B active per token, so it is quick on modest silicon. The knowledge stored is correspondingly thin, which matters more in legal work than speed does; realistically runnable on one small card by a mid-sized firm. |
| Granite-4.2-8B | 8B dense (8.79B parameters in the released weights) | 128K natively, extensible to 512K | ~6GB at 4-bit, ~11GB at 8-bit | The mid-size release of the current generation and the one we would recommend as a firm's default: Apache 2.0, a very long context extension, native reasoning, and a serving footprint small enough to leave room for a retrieval index and long context on one workstation card. Realistically runnable by any firm of 10–50 fee-earners, and the release most of them should start with. |
| Granite-4.0-H-Small | 32B total / 9B active (MoE with shared experts, hybrid Mamba-2 and attention) | 128K tokens | ~18GB at 4-bit, ~33GB at 8-bit | The largest mixture-of-experts release in the family: 32B of weights with 9B active per token, so it fits a 48GB-class card at 8-bit or a 24GB card at 4-bit with limited context headroom. The best quality per pound here for a firm with real concurrency, and a reminder that this family's MoE variants were designed to stay on-premise. |
| Granite-4.2-30B | 30B dense (29.3B parameters in the released weights) | 128K natively, extensible to 512K | ~17GB at 4-bit, ~31GB at 8-bit | The largest and most capable release in the current generation, and the one that shows how far this family's ceiling sits from a data centre: even the flagship fits a single card. A UK firm of 10–50 fee-earners can realistically run it, though a 30B dense model will not match a 350B mixture-of-experts release on a hard analytical question — nothing in this family is beyond a mid-sized firm's hardware, which is the opposite of every other family in this index. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
Granite's enterprise reputation is earned by obedience rather than brilliance, and obedience is what a fee-earner notices first on legal tasks. Give it a house format — clause heading, then a numbered summary, then a table of dates — and it reproduces that format on the tenth document as faithfully as on the first. For a firm with existing precedents, a style guide and a supervision model, that predictability is worth more than a point of raw reasoning. Extraction into a fixed schema is where it earns its place in a pilot: run over licence agreements and facility letters asking for parties, dates, notice periods, termination triggers and governing law as strict JSON, the output needs validation rather than repair. It rarely invents a field, it respects an 'unknown' convention if you define one, and it tolerates the layout damage that survives a scan-and-OCR pipeline better than a model of its size has any right to. Drafting is where the ceiling shows. Granite writes clean, plain, businesslike English and is genuinely good at the small courtesies of tone — 'we write further to your letter of 3 March' rather than the florid alternative. But it is a conservative drafter: asked for a first-pass letter of advice it produces something correct, slightly flat and shorter than the matter justifies, and a partner will add the argumentative spine it did not supply. Where it does noticeably well is the unglamorous drafting that consumes trainee time — attendance notes, file notes, covering emails, witness-summary tables, chronologies in a fixed template. Fee-earners stop rewriting those within a week, which is the honest test of whether an open model has paid for itself. Summarisation fidelity is good but conditional. Ask for a summary of a bundle and you get a competent one. Ask for a summary that uses only phrases appearing in the documents and it stays much closer to the source than most of this index, because the instruction-following is doing the work rather than the model's judgement. Without that constraint it compresses aggressively and drops the qualifications that make a legal summary safe — the 'subject to', the 'without prejudice', the sentence that turns a deadline into a target. Its reasoning traces, where a release offers them, are short and sober: less useful for audit than a large reasoning model's monologue, and also less prone to the confident dead end. Over-assertion is real but smaller than its peers'. Granite is more likely than the generalists above it to say that a document does not address something, particularly when the system prompt defines what 'not addressed' looks like. It also hedges in language a client may read as evasive — 'it may be that' — which is safe for review and irritating in a client-facing draft. In week one, three things strike a fee-earner. First, that a model of this size can do their document chores at all, which changes the economics of a pilot. Second, that the family around it — content-safety classifiers, embedding models, a document-conversion tool for PDFs and scans — is more useful than the language model alone, because retrieval quality dominates output quality in every legal workflow we have tested. Third, that the answers are unexciting, which is exactly the property you want when the alternative is eloquence you have to verify.
Granite quotes faithfully when the instruction is specific: name the clause, quote it verbatim, then give its page. Given that, we see few silent edits and no merging of passages. Left to summarise without the constraint it paraphrases, and paraphrase is where legal meaning drifts. Abstention is above average for its size — asked whether a bundle answers a question it will often say it does not — and improves further when the system prompt defines an explicit refusal string and supplies one worked example. Its evidence weakness is negative: it is a poor reporter of its own uncertainty, so it under-flags the moments when it has drifted from the supplied documents to general knowledge. Require a quotation for every assertion and that largely disappears.
What we would use it for
- High-volume extraction into a fixed schema — parties, dates, notice and termination provisions
- Chronologies, attendance notes and document schedules in a firm template
- First-pass summaries where the house rule is quote-or-omit
- Document conversion and OCR pipelines feeding a retrieval layer, using the publisher's own tooling
- A low-cost first private deployment on a single workstation, before committing to larger hardware
What to watch
- Conservative drafting that under-argues compared with a frontier-class model
- Aggressive compression when no source-quoting rule is set
- Hedging that reads as evasive in client-facing correspondence
- Under-reporting of its own uncertainty when it drifts off the supplied documents
- Rapid point releases — pin the exact build and re-run your task set after each bump
What it costs to run
One 24GB card serving Granite-4.2-8B at 8-bit is the configuration we would recommend to a firm of 10–50 fee-earners, and it is the cheapest credible private deployment in this index. That covers extraction, summarisation and drafting support for a team at concurrency that suits asynchronous work, with the retrieval index resident on the same machine, and the 30B release can be swapped in at 4-bit on the same card when a task needs more capacity. Latency is good on realistic prompts because this is a dense 8B model rather than a large mixture-of-experts, so the practical ceiling is quality rather than speed. The moment the model visibly fails your task set is the moment to move up a size, not to buy more GPUs.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $0.58 – $1.65 | $425 – $1,205 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $0.58 – $1.65 | $125 – $345 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $1.03 – $1.95 | $60 – $115 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~150 GB) | — | — | $20 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 24GB class — RTX 4090 / L4 / A5000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
This is the family where buying genuinely competes with renting: a 24GB-class workstation is an indicative $3,000–6,000, and a firm that keeps the model resident through the working day will usually recover that against the equivalent rented time within a year of full-time use. Renting a 24GB-class GPU is still the right way to run the pilot, because evaluation work is bursty and the pilot is what tells you which release is adequate. Even the mixture-of-experts and 30B releases run on a 48GB-class workstation, an indicative $8,000–15,000, so a firm can grow within this family without ever approaching server hardware. All of these figures are indicative.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Extraction and classification: temperature 0, top_p 0.8, a strict JSON schema and an explicit 'unknown' value. Summarisation under a quote-or-omit rule: 0.1. Drafting file notes and correspondence: 0.3. Do not push temperature above 0.5 anywhere in a legal workflow — this family does not buy quality with it, it buys drift. Pin the exact release in your evaluation record; the line moves in minor versions rather than major ones, and a minor bump has changed our extraction results before.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single workstation, small dense variant at low precision | Extraction, summarisation and file-note drafting for one team | The lowest-cost credible private pilot we have assessed; a 4-bit build leaves room for a small retrieval index on the same machine |
| One GPU server, mid-size dense variant | Firm-wide extraction and drafting support with retrieval over precedents | Good throughput per pound for schema-bound work; quantise the weights but keep the retrieval index in full precision |
| Larger deployment, or an earlier mixture-of-experts variant | Higher volume and concurrent users | Only justify the step up once the smaller model has visibly failed your task set — MoE quantisation has cost us more calibration than dense quantisation in our testing |
Who it suits
Firms that want a defensible, low-cost private deployment first — extraction, chronologies, document chores — with the least licence risk and the least hardware.
Matters that turn on persuasive drafting or multi-document reasoning, where a frontier-class model justifies the extra infrastructure.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. Assessed from published documentation and our standard legal evaluation task set. |
Sources
- Granite documentation and model releases (IBM)
- Granite model family (Hugging Face)
- Open-weight licence landscape 2026 (Presenc AI)