AI Tools Index / Open models / Llama
Meta · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

Llama

Llama 4 family (Scout, Maverick) and the Llama 3.x line still in production use

The default answer to 'can we run something capable ourselves?' — broad ecosystem support, predictable tooling, and a licence that a law firm's compliance team will want to read closely before signing anything.

Our verdict

Tier B — conditional, and the safest first private deployment for most firms because the ecosystem is mature and the paths are well trodden. The caveat is not capability, it is assertion: Llama will answer questions it should decline, so your deployment needs abstention instructions, source-quoting by default, and a reviewer who checks dates and names rather than tone.

Specifications, as published

PublisherMeta
FamilyLlama 4 family (Scout, Maverick) and the Llama 3.x line still in production use
ParametersFamily spans small single-GPU models to mixture-of-experts releases in the hundreds of billions of total parameters with a fraction active per token
ContextLong context on the current generation; effective context is materially lower than the advertised figure once retrieval is in play
LicenceMeta Llama Community Licence — permissive for most commercial use, with a monthly active user threshold that triggers additional terms, acceptable-use restrictions, and naming/attribution obligations
WeightsDownloadable, including quantised community builds
ReleaseCurrent generation released 2025, with the previous generation still widely deployed
Licence postureMeta Llama Community Licence — permissive for most commercial use, with a monthly active user threshold that triggers additional terms, acceptable-use restrictions, and naming/attribution obligations

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
Llama 3.2 1B Instruct1.23B dense128k tokens as published~1GB at 4-bit quantisation; ~2.5GB at 8-bitA classification and routing model rather than a drafter — useful as the cheap gate in front of a larger model, and small enough to run on hardware a firm already owns.
Llama 3.2 3B Instruct3.21B dense128k tokens as published~2.5GB at 4-bit quantisation; ~6GB at 8-bitSuits bounded extraction and formatting work on a single 24GB card, but a firm of 10–50 fee-earners should treat its output as a draft for a paralegal to validate rather than anything client-facing.
Llama 3.1 8B Instruct8B dense128k tokens as published~6GB at 4-bit quantisation; ~10GB at 8-bitThe workhorse small release: one 24GB card runs it for summarisation and schema-bound extraction with several users at once, and it is the smallest Llama we would put in front of a real matter file.
Llama 3.3 70B Instruct70.6B dense128k tokens as published~40GB at 4-bit quantisation; ~73GB at 8-bitOur recommended default for this family at firm scale — it fits one 48GB card at 4-bit and is the smallest Llama that fee-earners tend to accept as a drafting assistant.
Llama 4 Scout (17Bx16E)109B total / 17B active, 16 experts10M tokens as published~55GB at 4-bit quantisation; ~109GB at 8-bit; ~218GB for the published BF16 weightsThe long-context release — a whole disclosure bundle in one prompt on a single 80GB card — but most firms will use a fraction of the advertised window and should read the Llama 4 Community Licence before deploying it.
Llama 4 Maverick (17Bx128E)400B total / 17B active, 128 experts1M tokens as published~200GB at 4-bit quantisation; ~400GB for the published FP8 build; ~800GB at BF16Capable, and impractical for a 20-partner firm: a 400B release needs a multi-GPU node or a private tenancy, and in our legal tasks the gain over Llama 3.3 70B does not pay for that hardware.
Llama 3.1 405B Instruct405.9B dense128k tokens as published~205GB at 4-bit quantisation; ~405GB for the published FP8 checkpoint; ~810GB at BF16Meta's densest text release and a genuine data-centre deployment — no firm of 10–50 fee-earners should plan to serve it themselves, and it appears here mainly to show how far the Llama name reaches from a workstation.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

Llama's default behaviour is the closest thing to a general-purpose house style among open weights: it follows formatting instructions reliably, holds a requested structure across a long answer, and produces drafting that a fee-earner can edit rather than rewrite. Where it needs supervision in legal work is assertion. Ask it a question that the retrieved passages do not answer and it will often construct a plausible answer from surrounding context instead of stopping, which is exactly the failure mode that makes a hallucinated citation dangerous in a file note. It is also a confident summariser: give it a bundle and it will produce a tidy chronology, but you should treat every date and party name as unverified unless you asked it to quote the source line. Its reasoning traces are useful for review — when it is asked to work through a question, the intermediate steps expose where an assumption entered, which is more than most legal AI products show. In drafting-from-precedent work it is strong on structure and tone, weaker on jurisdiction-specific nuance: it will write an English-law letter that reads well and quietly imports concepts from US practice. Test it on your own precedents before you trust the register.

evidence and abstention

Quotation fidelity is good when the instruction is explicit ('quote the clause verbatim, with its page reference'). Without that instruction it paraphrases, and paraphrased legal language is how an obligation acquires a meaning nobody agreed. Negative instruction ('if the documents do not answer this, say so') improves abstention markedly and should be in every system prompt.

What we would use it for

  • Chronology and bundle summarisation, with source references insisted on
  • First-pass drafting from firm precedent, reviewed line by line
  • Extraction into a fixed schema (dates, parties, obligations, termination rights)
  • Internal know-how search over published materials — non-confidential

What to watch

  • Confident answers where the retrieved material is silent
  • Occasional US-law drift in drafting tasks
  • Licence thresholds and acceptable-use riders — read them, do not skim them
  • Advertised context well above what survives a real matter file

What it costs to run

One 48GB card runs the release we would actually serve — Llama 3.3 70B Instruct at 4-bit — and carries a firm of 10–50 fee-earners for drafting support, extraction and bundle summarisation at something like eight to twelve concurrent users before queueing becomes visible. Latency is the honest constraint with a dense 70B: expect reading speed on long generations and near-instant responses on short extraction calls, which is why we would keep schema-bound work and drafting on separate queues. Llama 4 Scout needs an 80GB card and Maverick a multi-GPU node, and neither is worth buying until evaluation shows 3.3 70B failing on work that matters.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$0.85 – $2.85$625 – $2,080Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$0.85 – $2.85$180 – $600The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$1.80 – $3.60$110 – $215Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~150 GB)$20Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 48GB class — L40S / RTX A6000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

For the 70B release the comparison is a single 48GB-class workstation against renting the same class: a 48GB-class workstation is indicatively $8,000–15,000, and a firm using the model for only part of each working day will not recover that against rented capacity quickly. Renting also keeps the refresh problem — quantisation formats, inference engines, the next Llama generation — with the provider rather than with your own IT. These bands are indicative rather than quotes and exclude power, support and the retrieval layer; the sensible pattern for most 10–50 fee-earner firms is to rent for a pilot and buy only once the workload is proven and steady.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Extraction and classification tasks: temperature 0–0.2, top_p 0.9, with a strict output schema. Drafting: 0.3–0.5 for tone variety, never above 0.7 for anything client-facing. Summarisation of a bundle: 0.1. Always pin the model version in your evaluation record — behaviour moves between point releases.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
Single workstation (quantised small variant)Summarisation, extraction, drafting support for one teamThe realistic starting point for a firm of 10–30 fee-earners
One GPU server (mid-size variant)Shared firm capability with retrieval over precedentsCost is dominated by the retrieval layer and evaluation time, not the hardware
MoE releases, larger deploymentFirm-wide deployment, higher throughputOnly worth it once evaluation shows the smaller model failing on tasks that matter

Who it suits

Good fit

Firms wanting a capable general model they own, with the largest pool of tooling, documentation and community experience behind it.

Poor fit

Firms looking for a legal-specialist model out of the box, or anyone who will not read the licence terms.

Review history

DateChange
Sep 2026First entry. Assessed from published documentation and our standard legal evaluation task set.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.