Llama
The default answer to 'can we run something capable ourselves?' — broad ecosystem support, predictable tooling, and a licence that a law firm's compliance team will want to read closely before signing anything.
Tier B — conditional, and the safest first private deployment for most firms because the ecosystem is mature and the paths are well trodden. The caveat is not capability, it is assertion: Llama will answer questions it should decline, so your deployment needs abstention instructions, source-quoting by default, and a reviewer who checks dates and names rather than tone.
Specifications, as published
| Publisher | Meta |
| Family | Llama 4 family (Scout, Maverick) and the Llama 3.x line still in production use |
| Parameters | Family spans small single-GPU models to mixture-of-experts releases in the hundreds of billions of total parameters with a fraction active per token |
| Context | Long context on the current generation; effective context is materially lower than the advertised figure once retrieval is in play |
| Licence | Meta Llama Community Licence — permissive for most commercial use, with a monthly active user threshold that triggers additional terms, acceptable-use restrictions, and naming/attribution obligations |
| Weights | Downloadable, including quantised community builds |
| Release | Current generation released 2025, with the previous generation still widely deployed |
| Licence posture | Meta Llama Community Licence — permissive for most commercial use, with a monthly active user threshold that triggers additional terms, acceptable-use restrictions, and naming/attribution obligations |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| Llama 3.2 1B Instruct | 1.23B dense | 128k tokens as published | ~1GB at 4-bit quantisation; ~2.5GB at 8-bit | A classification and routing model rather than a drafter — useful as the cheap gate in front of a larger model, and small enough to run on hardware a firm already owns. |
| Llama 3.2 3B Instruct | 3.21B dense | 128k tokens as published | ~2.5GB at 4-bit quantisation; ~6GB at 8-bit | Suits bounded extraction and formatting work on a single 24GB card, but a firm of 10–50 fee-earners should treat its output as a draft for a paralegal to validate rather than anything client-facing. |
| Llama 3.1 8B Instruct | 8B dense | 128k tokens as published | ~6GB at 4-bit quantisation; ~10GB at 8-bit | The workhorse small release: one 24GB card runs it for summarisation and schema-bound extraction with several users at once, and it is the smallest Llama we would put in front of a real matter file. |
| Llama 3.3 70B Instruct | 70.6B dense | 128k tokens as published | ~40GB at 4-bit quantisation; ~73GB at 8-bit | Our recommended default for this family at firm scale — it fits one 48GB card at 4-bit and is the smallest Llama that fee-earners tend to accept as a drafting assistant. |
| Llama 4 Scout (17Bx16E) | 109B total / 17B active, 16 experts | 10M tokens as published | ~55GB at 4-bit quantisation; ~109GB at 8-bit; ~218GB for the published BF16 weights | The long-context release — a whole disclosure bundle in one prompt on a single 80GB card — but most firms will use a fraction of the advertised window and should read the Llama 4 Community Licence before deploying it. |
| Llama 4 Maverick (17Bx128E) | 400B total / 17B active, 128 experts | 1M tokens as published | ~200GB at 4-bit quantisation; ~400GB for the published FP8 build; ~800GB at BF16 | Capable, and impractical for a 20-partner firm: a 400B release needs a multi-GPU node or a private tenancy, and in our legal tasks the gain over Llama 3.3 70B does not pay for that hardware. |
| Llama 3.1 405B Instruct | 405.9B dense | 128k tokens as published | ~205GB at 4-bit quantisation; ~405GB for the published FP8 checkpoint; ~810GB at BF16 | Meta's densest text release and a genuine data-centre deployment — no firm of 10–50 fee-earners should plan to serve it themselves, and it appears here mainly to show how far the Llama name reaches from a workstation. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
Llama's default behaviour is the closest thing to a general-purpose house style among open weights: it follows formatting instructions reliably, holds a requested structure across a long answer, and produces drafting that a fee-earner can edit rather than rewrite. Where it needs supervision in legal work is assertion. Ask it a question that the retrieved passages do not answer and it will often construct a plausible answer from surrounding context instead of stopping, which is exactly the failure mode that makes a hallucinated citation dangerous in a file note. It is also a confident summariser: give it a bundle and it will produce a tidy chronology, but you should treat every date and party name as unverified unless you asked it to quote the source line. Its reasoning traces are useful for review — when it is asked to work through a question, the intermediate steps expose where an assumption entered, which is more than most legal AI products show. In drafting-from-precedent work it is strong on structure and tone, weaker on jurisdiction-specific nuance: it will write an English-law letter that reads well and quietly imports concepts from US practice. Test it on your own precedents before you trust the register.
Quotation fidelity is good when the instruction is explicit ('quote the clause verbatim, with its page reference'). Without that instruction it paraphrases, and paraphrased legal language is how an obligation acquires a meaning nobody agreed. Negative instruction ('if the documents do not answer this, say so') improves abstention markedly and should be in every system prompt.
What we would use it for
- Chronology and bundle summarisation, with source references insisted on
- First-pass drafting from firm precedent, reviewed line by line
- Extraction into a fixed schema (dates, parties, obligations, termination rights)
- Internal know-how search over published materials — non-confidential
What to watch
- Confident answers where the retrieved material is silent
- Occasional US-law drift in drafting tasks
- Licence thresholds and acceptable-use riders — read them, do not skim them
- Advertised context well above what survives a real matter file
What it costs to run
One 48GB card runs the release we would actually serve — Llama 3.3 70B Instruct at 4-bit — and carries a firm of 10–50 fee-earners for drafting support, extraction and bundle summarisation at something like eight to twelve concurrent users before queueing becomes visible. Latency is the honest constraint with a dense 70B: expect reading speed on long generations and near-instant responses on short extraction calls, which is why we would keep schema-bound work and drafting on separate queues. Llama 4 Scout needs an 80GB card and Maverick a multi-GPU node, and neither is worth buying until evaluation shows 3.3 70B failing on work that matters.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $0.85 – $2.85 | $625 – $2,080 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $0.85 – $2.85 | $180 – $600 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $1.80 – $3.60 | $110 – $215 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~150 GB) | — | — | $20 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 48GB class — L40S / RTX A6000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
For the 70B release the comparison is a single 48GB-class workstation against renting the same class: a 48GB-class workstation is indicatively $8,000–15,000, and a firm using the model for only part of each working day will not recover that against rented capacity quickly. Renting also keeps the refresh problem — quantisation formats, inference engines, the next Llama generation — with the provider rather than with your own IT. These bands are indicative rather than quotes and exclude power, support and the retrieval layer; the sensible pattern for most 10–50 fee-earner firms is to rent for a pilot and buy only once the workload is proven and steady.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Extraction and classification tasks: temperature 0–0.2, top_p 0.9, with a strict output schema. Drafting: 0.3–0.5 for tone variety, never above 0.7 for anything client-facing. Summarisation of a bundle: 0.1. Always pin the model version in your evaluation record — behaviour moves between point releases.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single workstation (quantised small variant) | Summarisation, extraction, drafting support for one team | The realistic starting point for a firm of 10–30 fee-earners |
| One GPU server (mid-size variant) | Shared firm capability with retrieval over precedents | Cost is dominated by the retrieval layer and evaluation time, not the hardware |
| MoE releases, larger deployment | Firm-wide deployment, higher throughput | Only worth it once evaluation shows the smaller model failing on tasks that matter |
Who it suits
Firms wanting a capable general model they own, with the largest pool of tooling, documentation and community experience behind it.
Firms looking for a legal-specialist model out of the box, or anyone who will not read the licence terms.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. Assessed from published documentation and our standard legal evaluation task set. |
Sources
- Llama licence and acceptable use (Meta)
- Open-weight licence landscape 2026 (Presenc AI)
- Probative Co: private AI in law — the 2026 guide