AI Tools Index / Open models / Gemma
Google DeepMind · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

Gemma

The Gemma line of small dense open-weights models, including multimodal and multilingual releases

A small-model family that fits a single workstation, and whose current generation is Apache 2.0 — the most permissive licence position of any large publisher in this index.

Our verdict

Tier B — conditional, and now for capability reasons rather than licence ones. Nothing here needs a data-centre budget: the releases we would serve run on one 24GB card at 4-bit, which makes this the cheapest credible starting point for a firm that wants AI it owns. The caveat is ceiling: there is no Gemma release in the data-centre class, so hard reasoning and long-context work will need a different family eventually — and the licence improves as you move up to the current generation, which is worth knowing before you standardise on an older tag.

Specifications, as published

PublisherGoogle DeepMind
FamilyThe Gemma line of small dense open-weights models, including multimodal and multilingual releases
ParametersSmall dense models across several sizes, sized for a single workstation or a modest shared server rather than for datacentre-scale deployment
ContextLong context on the larger sizes and shorter on the smallest; the moderate sizes hold retrieval-grounded tasks better than their size suggests
LicenceMixed by generation, and this is the field to check first. The Gemma 4 releases are published under Apache 2.0 — as permissive as open weights get, with no acceptable-use rider attached to the weights themselves — while Gemma 3 remains under Google's own Gemma Terms of Use, which carry use restrictions. Confirm which generation you are downloading before you serve it on client matters.
WeightsDownloadable, with strong quantised community support across the range
ReleaseRegular generations with the previous generation still widely deployed
Licence postureMixed by generation, and this is the field to check first. The Gemma 4 releases are published under Apache 2.0 — as permissive as open weights get, with no acceptable-use rider attached to the weights themselves — while Gemma 3 remains under Google's own Gemma Terms of Use, which carry use restrictions. Confirm which generation you are downloading before you serve it on client matters.

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
Gemma 3 270M IT268M dense32k tokens as published~0.3GB at 4-bit quantisation; ~0.5GB at 8-bitA previous-generation hyper-efficient release for on-device classification and formatting work — a component rather than a legal-work model, and still published under the Gemma Terms of Use rather than an Apache licence.
Gemma 4 E2B IT2.3B effective (5.1B including embeddings), text, image and audio input128k tokens~3GB at 4-bit quantisation; ~6GB at 8-bit, embeddings includedThe nano end of the current generation, small enough for a laptop — useful for an individual fee-earner's summarisation and first-pass translation, and not for client-facing output without review.
Gemma 4 E4B IT4.5B effective (8B including embeddings), text, image and audio input128k tokens~5GB at 4-bit quantisation; ~9GB at 8-bit, embeddings includedThe best-value release in this family for a firm of 10–50 fee-earners: one 24GB card serves it at 8-bit alongside a document index, which is a great deal for the summarisation and multilingual work this family does well.
Gemma 4 12B IT (Unified)11.95B dense, encoder-free multimodal256k tokens~7GB at 4-bit quantisation; ~13GB at 8-bitThe step up for firms that want longer context and better drafting on a single 24GB card, with images and audio projected straight into the same decoder rather than through separate encoders.
Gemma 4 12B IT QAT q4_011.95B dense, quantisation-aware training256k tokens~7GB as published in the q4_0 build; the quantised checkpoints are the ones the publisher ships for edge and workstation deploymentA quantised-first release of the 12B model, trained for quantisation rather than quantised afterwards, and the build to choose if you want 12B-class behaviour on a 16GB or 24GB card.
Gemma 4 26B A4B IT25.2B total / 3.8B active, 8 active of 128 experts plus one shared256k tokens~14GB at 4-bit quantisation; ~26GB at 8-bitThe release we would actually serve: 3.8B active parameters make it quick on a 24GB card, and it is the largest Gemma that a firm of 10–50 fee-earners can run without buying a server.
Gemma 4 31B IT30.7B dense, multimodal256k tokens~17GB at 4-bit quantisation; ~32GB at 8-bitThe family's largest release and still a single-card deployment — 4-bit fits a 24GB card and 8-bit wants a 48GB card — so, unlike the other families in this index, no Gemma release is out of reach for a 20-partner firm on hardware grounds alone.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

For legal work Gemma's appeal is that it gets a great deal done per pound of hardware. In our evaluation tasks the moderate sizes produce tidy, well-structured summaries of witness statements, attendance notes and correspondence bundles, keep a section structure when asked for one, and write plain-English explanations that a non-lawyer client could read without a covering letter. That last point is the family's most commercially useful trait for a small firm: it is good at explaining, and it is good at doing so briefly. Extraction into a fixed schema works well on short, clean documents — one letter, one order, one invoice — and degrades as documents get longer or denser, which is the pattern you would expect from models of this size and worth planning around. If your extraction task is a hundred short documents rather than one agreement, the family is efficient; if it is a single 200-page bundle, split the task rather than hoping. Formatting discipline is better than most small models and short of the frontier-class entries here: it holds a requested structure, produces usable tables, and will occasionally add an unrequested summary or a closing pleasantry that you then have to strip. The helpfulness bias is the behavioural problem to plan for. Gemma is trained to be agreeable and useful, so when the retrieved documents do not answer the question it tends to offer something anyway — general background, a plausible inference, an answer shaped like the one that was asked for. It is not unusual in that, but it is more polite about it than the frontier models and therefore easier to miss on review. A negative instruction helps a great deal, and so does asking for the supporting quotation before the conclusion, which turns an unverifiable claim into a checkable one. On reasoning, there is no dedicated reasoning line in this family, so multi-step analysis is the weakest use: allocations, conflict analysis and argument-strength questions come back as fluent prose with the working hidden, and the fluency can be mistaken for analysis. Where the family is genuinely differentiated is language. Coverage extends far beyond the languages most open weights handle well, and in our tasks the register holds up in German, French, Spanish, Italian, Portuguese, Dutch and the Nordic languages to a standard that small models from other publishers do not reach. Translation of incoming correspondence into English, summarisation in the language of the document, and same-language reply drafting are all realistic first-week uses. Two things fee-earners notice quickly. The first is speed and predictability — the moderate sizes respond fast enough to sit in a normal document workflow, and quantised builds are stable. The second is tone: default output is helpful, slightly formal and inclined to soften conclusions, which produces acceptable client correspondence and slightly mushy internal notes. Instruct it to state findings flatly and it complies. The realistic picture is a capable, private, low-footprint assistant for summarisation, explanation, translation and light drafting, with the evidential questions and the analytical work left to a supervised reviewer or a larger model.

evidence and abstention

Quotation fidelity is decent when asked directly and weaker when the request is implicit: without a verbatim instruction the model paraphrases, and paraphrased obligation language is where meaning shifts. Its abstention is its weak point — the helpfulness bias means it prefers a plausible answer to a refusal, so a negative instruction, an explicit empty-answer format and a requirement to quote first are all needed. With those three in place, in our tasks the moderate sizes flag silent corpora reliably enough to be supervised rather than distrusted. Verify every quotation against the source regardless.

What we would use it for

  • Summarising correspondence, attendance notes and short bundles in plain English
  • Client-facing explanations of process and next steps, reviewed before sending
  • Translation and cross-language summarisation of incoming foreign-language material
  • Extraction from short documents into a fixed schema at volume
  • Drafting support on routine correspondence where a firm precedent is supplied

What to watch

  • Licence differs by generation — Gemma 4 is Apache 2.0; Gemma 3 is under the Gemma Terms of Use with use restrictions
  • No data-centre-class release: hard reasoning and very long context will need another family
  • Small models are persuasive summarisers — dates and party names should be verified against the source, not the summary
  • Quantisation choices change behaviour more here than on larger models; re-run your evaluation set after any precision change
  • Tune the instruction-following to a fixed schema or output drifts between releases

What it costs to run

One 24GB card is enough for the release we would serve — Gemma 4 26B A4B at 4-bit — whose 3.8B active parameters per token make it noticeably faster than its 25B total suggests, and which holds the 256k window the medium sizes claim for retrieval-grounded work. For a firm of 10–50 fee-earners that is a shared summarisation, translation and light-extraction capability with room for several concurrent users, and the E4B release runs comfortably for an individual fee-earner on a laptop or a desktop card. Latency is this family's strength at these sizes; analytical depth is the limit, so we would pair it with a larger model rather than ask it to carry reasoning-heavy work.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$0.58 – $1.65$425 – $1,205Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$0.58 – $1.65$125 – $345The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$1.03 – $1.95$60 – $115Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~100 GB)$15Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 24GB class — RTX 4090 / L4 / A5000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

Gemma 4 at these sizes is the one family in this index where buying is straightforwardly sensible: the 24GB class is the least expensive hardware here — indicatively a few thousand dollars for a workstation rather than the $8,000–15,000 that a 48GB-class machine implies — and it runs both the 26B A4B and the 12B releases without a further hardware decision. Renting the same class is still cheaper for occasional use, but at this capital level the choice is closer than for the 48GB and 80GB classes, and a firm that wants client material to stay inside its own perimeter may prefer the certainty of owning. These bands are indicative rather than quotes and exclude power, support, storage and the retrieval layer.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Extraction from short documents: temperature 0, top_p 0.8, strict schema, no commentary. Summarisation: 0.1–0.2, with a sentence-per-source rule to stop it merging documents. Client-facing explanation: 0.3. Translation: 0.2, with the target register named explicitly. Avoid the family for multi-step analysis rather than trying to tune it into one. Pin the exact release and the quantisation you evaluated — quantised community builds of the same model are not interchangeable, and small models are more sensitive to quantisation than large ones.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
Single workstation (smallest quantised variant)Summarisation, translation and light extraction for an individual fee-earnerThe cheapest credible private deployment in this index, and light enough to keep entirely in the firm's own environment
One GPU server (moderate size, or a smaller size at full precision)Shared firm capability over published precedent and know-how, with retrievalThe sweet spot for a firm of 10–30 fee-earners; capacity is limited by concurrent users rather than document size
Two-GPU server or private tenancyFirm-wide summarisation and translation, longer contextsWorth it only for throughput; a larger model may be the better purchase if the workload includes analysis

Who it suits

Good fit

Firms of 5–50 fee-earners wanting a genuinely private first deployment on hardware they already own, under an Apache-2.0 licence, for extraction and summarisation.

Poor fit

Work that depends on hard multi-step legal reasoning or very long context — pick a larger family for that, or run two models.

Review history

DateChange
Sep 2026First entry.
Sep 2026Licence re-scored from 17 to 22 and the verdict reframed after reading the current model cards: the Gemma 4 generation is Apache 2.0, while Gemma 3 remains under the Gemma Terms of Use. Total 76 → 81, tier unchanged at B.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.