Gemma
A small-model family that fits a single workstation, and whose current generation is Apache 2.0 — the most permissive licence position of any large publisher in this index.
Tier B — conditional, and now for capability reasons rather than licence ones. Nothing here needs a data-centre budget: the releases we would serve run on one 24GB card at 4-bit, which makes this the cheapest credible starting point for a firm that wants AI it owns. The caveat is ceiling: there is no Gemma release in the data-centre class, so hard reasoning and long-context work will need a different family eventually — and the licence improves as you move up to the current generation, which is worth knowing before you standardise on an older tag.
Specifications, as published
| Publisher | Google DeepMind |
| Family | The Gemma line of small dense open-weights models, including multimodal and multilingual releases |
| Parameters | Small dense models across several sizes, sized for a single workstation or a modest shared server rather than for datacentre-scale deployment |
| Context | Long context on the larger sizes and shorter on the smallest; the moderate sizes hold retrieval-grounded tasks better than their size suggests |
| Licence | Mixed by generation, and this is the field to check first. The Gemma 4 releases are published under Apache 2.0 — as permissive as open weights get, with no acceptable-use rider attached to the weights themselves — while Gemma 3 remains under Google's own Gemma Terms of Use, which carry use restrictions. Confirm which generation you are downloading before you serve it on client matters. |
| Weights | Downloadable, with strong quantised community support across the range |
| Release | Regular generations with the previous generation still widely deployed |
| Licence posture | Mixed by generation, and this is the field to check first. The Gemma 4 releases are published under Apache 2.0 — as permissive as open weights get, with no acceptable-use rider attached to the weights themselves — while Gemma 3 remains under Google's own Gemma Terms of Use, which carry use restrictions. Confirm which generation you are downloading before you serve it on client matters. |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| Gemma 3 270M IT | 268M dense | 32k tokens as published | ~0.3GB at 4-bit quantisation; ~0.5GB at 8-bit | A previous-generation hyper-efficient release for on-device classification and formatting work — a component rather than a legal-work model, and still published under the Gemma Terms of Use rather than an Apache licence. |
| Gemma 4 E2B IT | 2.3B effective (5.1B including embeddings), text, image and audio input | 128k tokens | ~3GB at 4-bit quantisation; ~6GB at 8-bit, embeddings included | The nano end of the current generation, small enough for a laptop — useful for an individual fee-earner's summarisation and first-pass translation, and not for client-facing output without review. |
| Gemma 4 E4B IT | 4.5B effective (8B including embeddings), text, image and audio input | 128k tokens | ~5GB at 4-bit quantisation; ~9GB at 8-bit, embeddings included | The best-value release in this family for a firm of 10–50 fee-earners: one 24GB card serves it at 8-bit alongside a document index, which is a great deal for the summarisation and multilingual work this family does well. |
| Gemma 4 12B IT (Unified) | 11.95B dense, encoder-free multimodal | 256k tokens | ~7GB at 4-bit quantisation; ~13GB at 8-bit | The step up for firms that want longer context and better drafting on a single 24GB card, with images and audio projected straight into the same decoder rather than through separate encoders. |
| Gemma 4 12B IT QAT q4_0 | 11.95B dense, quantisation-aware training | 256k tokens | ~7GB as published in the q4_0 build; the quantised checkpoints are the ones the publisher ships for edge and workstation deployment | A quantised-first release of the 12B model, trained for quantisation rather than quantised afterwards, and the build to choose if you want 12B-class behaviour on a 16GB or 24GB card. |
| Gemma 4 26B A4B IT | 25.2B total / 3.8B active, 8 active of 128 experts plus one shared | 256k tokens | ~14GB at 4-bit quantisation; ~26GB at 8-bit | The release we would actually serve: 3.8B active parameters make it quick on a 24GB card, and it is the largest Gemma that a firm of 10–50 fee-earners can run without buying a server. |
| Gemma 4 31B IT | 30.7B dense, multimodal | 256k tokens | ~17GB at 4-bit quantisation; ~32GB at 8-bit | The family's largest release and still a single-card deployment — 4-bit fits a 24GB card and 8-bit wants a 48GB card — so, unlike the other families in this index, no Gemma release is out of reach for a 20-partner firm on hardware grounds alone. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
For legal work Gemma's appeal is that it gets a great deal done per pound of hardware. In our evaluation tasks the moderate sizes produce tidy, well-structured summaries of witness statements, attendance notes and correspondence bundles, keep a section structure when asked for one, and write plain-English explanations that a non-lawyer client could read without a covering letter. That last point is the family's most commercially useful trait for a small firm: it is good at explaining, and it is good at doing so briefly. Extraction into a fixed schema works well on short, clean documents — one letter, one order, one invoice — and degrades as documents get longer or denser, which is the pattern you would expect from models of this size and worth planning around. If your extraction task is a hundred short documents rather than one agreement, the family is efficient; if it is a single 200-page bundle, split the task rather than hoping. Formatting discipline is better than most small models and short of the frontier-class entries here: it holds a requested structure, produces usable tables, and will occasionally add an unrequested summary or a closing pleasantry that you then have to strip. The helpfulness bias is the behavioural problem to plan for. Gemma is trained to be agreeable and useful, so when the retrieved documents do not answer the question it tends to offer something anyway — general background, a plausible inference, an answer shaped like the one that was asked for. It is not unusual in that, but it is more polite about it than the frontier models and therefore easier to miss on review. A negative instruction helps a great deal, and so does asking for the supporting quotation before the conclusion, which turns an unverifiable claim into a checkable one. On reasoning, there is no dedicated reasoning line in this family, so multi-step analysis is the weakest use: allocations, conflict analysis and argument-strength questions come back as fluent prose with the working hidden, and the fluency can be mistaken for analysis. Where the family is genuinely differentiated is language. Coverage extends far beyond the languages most open weights handle well, and in our tasks the register holds up in German, French, Spanish, Italian, Portuguese, Dutch and the Nordic languages to a standard that small models from other publishers do not reach. Translation of incoming correspondence into English, summarisation in the language of the document, and same-language reply drafting are all realistic first-week uses. Two things fee-earners notice quickly. The first is speed and predictability — the moderate sizes respond fast enough to sit in a normal document workflow, and quantised builds are stable. The second is tone: default output is helpful, slightly formal and inclined to soften conclusions, which produces acceptable client correspondence and slightly mushy internal notes. Instruct it to state findings flatly and it complies. The realistic picture is a capable, private, low-footprint assistant for summarisation, explanation, translation and light drafting, with the evidential questions and the analytical work left to a supervised reviewer or a larger model.
Quotation fidelity is decent when asked directly and weaker when the request is implicit: without a verbatim instruction the model paraphrases, and paraphrased obligation language is where meaning shifts. Its abstention is its weak point — the helpfulness bias means it prefers a plausible answer to a refusal, so a negative instruction, an explicit empty-answer format and a requirement to quote first are all needed. With those three in place, in our tasks the moderate sizes flag silent corpora reliably enough to be supervised rather than distrusted. Verify every quotation against the source regardless.
What we would use it for
- Summarising correspondence, attendance notes and short bundles in plain English
- Client-facing explanations of process and next steps, reviewed before sending
- Translation and cross-language summarisation of incoming foreign-language material
- Extraction from short documents into a fixed schema at volume
- Drafting support on routine correspondence where a firm precedent is supplied
What to watch
- Licence differs by generation — Gemma 4 is Apache 2.0; Gemma 3 is under the Gemma Terms of Use with use restrictions
- No data-centre-class release: hard reasoning and very long context will need another family
- Small models are persuasive summarisers — dates and party names should be verified against the source, not the summary
- Quantisation choices change behaviour more here than on larger models; re-run your evaluation set after any precision change
- Tune the instruction-following to a fixed schema or output drifts between releases
What it costs to run
One 24GB card is enough for the release we would serve — Gemma 4 26B A4B at 4-bit — whose 3.8B active parameters per token make it noticeably faster than its 25B total suggests, and which holds the 256k window the medium sizes claim for retrieval-grounded work. For a firm of 10–50 fee-earners that is a shared summarisation, translation and light-extraction capability with room for several concurrent users, and the E4B release runs comfortably for an individual fee-earner on a laptop or a desktop card. Latency is this family's strength at these sizes; analytical depth is the limit, so we would pair it with a larger model rather than ask it to carry reasoning-heavy work.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $0.58 – $1.65 | $425 – $1,205 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $0.58 – $1.65 | $125 – $345 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $1.03 – $1.95 | $60 – $115 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~100 GB) | — | — | $15 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 24GB class — RTX 4090 / L4 / A5000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
Gemma 4 at these sizes is the one family in this index where buying is straightforwardly sensible: the 24GB class is the least expensive hardware here — indicatively a few thousand dollars for a workstation rather than the $8,000–15,000 that a 48GB-class machine implies — and it runs both the 26B A4B and the 12B releases without a further hardware decision. Renting the same class is still cheaper for occasional use, but at this capital level the choice is closer than for the 48GB and 80GB classes, and a firm that wants client material to stay inside its own perimeter may prefer the certainty of owning. These bands are indicative rather than quotes and exclude power, support, storage and the retrieval layer.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Extraction from short documents: temperature 0, top_p 0.8, strict schema, no commentary. Summarisation: 0.1–0.2, with a sentence-per-source rule to stop it merging documents. Client-facing explanation: 0.3. Translation: 0.2, with the target register named explicitly. Avoid the family for multi-step analysis rather than trying to tune it into one. Pin the exact release and the quantisation you evaluated — quantised community builds of the same model are not interchangeable, and small models are more sensitive to quantisation than large ones.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single workstation (smallest quantised variant) | Summarisation, translation and light extraction for an individual fee-earner | The cheapest credible private deployment in this index, and light enough to keep entirely in the firm's own environment |
| One GPU server (moderate size, or a smaller size at full precision) | Shared firm capability over published precedent and know-how, with retrieval | The sweet spot for a firm of 10–30 fee-earners; capacity is limited by concurrent users rather than document size |
| Two-GPU server or private tenancy | Firm-wide summarisation and translation, longer contexts | Worth it only for throughput; a larger model may be the better purchase if the workload includes analysis |
Who it suits
Firms of 5–50 fee-earners wanting a genuinely private first deployment on hardware they already own, under an Apache-2.0 licence, for extraction and summarisation.
Work that depends on hard multi-step legal reasoning or very long context — pick a larger family for that, or run two models.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. |
| Sep 2026 | Licence re-scored from 17 to 22 and the verdict reframed after reading the current model cards: the Gemma 4 generation is Apache 2.0, while Gemma 3 remains under the Gemma Terms of Use. Total 76 → 81, tier unchanged at B. |
Sources
- Gemma model releases (Google DeepMind on Hugging Face)
- Gemma Terms of Use (Google)
- Open source LLM landscape 2026