DeepSeek
Exceptional reasoning per pound under a permissive licence, with a provenance and hosting question that a law firm's client due-diligence conversation will eventually have to answer.
Tier B — conditional, and the best value in open weights for analysis. It is not the model we would choose for citation-heavy work, and it is the model in this index most likely to trigger a client question about provenance. Use it for reasoning and triage, keep it away from final drafting, and document your answers to the due-diligence questions before you are asked.
Specifications, as published
| Publisher | DeepSeek |
| Family | DeepSeek V-series general models and the R-series reasoning line |
| Parameters | Large mixture-of-experts general models plus a reasoning-focused line |
| Context | Long context advertised; effective context under retrieval is mid-pack in our tasks |
| Licence | MIT for the principal releases — permissive, with a model-specific acceptable-use policy layered on top |
| Weights | Downloadable and widely re-hosted |
| Release | Rapid iteration between general and reasoning lines |
| Licence posture | MIT for the principal releases — permissive, with a model-specific acceptable-use policy layered on top |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| DeepSeek-R1-Distill-Qwen-1.5B | 1.5B dense, distilled from R1 | 128k tokens as configured in the published release; served far shorter in practice | ~1GB at 4-bit quantisation; ~1.8GB at 8-bit | A distillation release for experiment and edge routing — MIT-licensed and trivially small — and not a model any firm should put in front of client material. |
| DeepSeek-R1-Distill-Llama-8B | 8B dense, distilled from R1 | 128k tokens as configured in the published release | ~6GB at 4-bit quantisation; ~10GB at 8-bit | The small reasoning release most firms actually start with — one 24GB card runs it comfortably and shows the visible chain of reasoning the R-series is chosen for — though its weights derive from Llama 3.1, so Meta's community licence terms travel with it. |
| DeepSeek-R1-Distill-Qwen-32B | 32B dense, distilled from R1 | 128k tokens as configured in the published release | ~18GB at 4-bit quantisation; ~34GB at 8-bit | Our recommended default for this family at mid-size: a 48GB card serves it at 4-bit with room for the long reasoning traces, which gives a firm the reasoning line's behaviour without a data-centre node. |
| DeepSeek-V4-Flash-0731 | 284B total / 13B active (the published V4-Flash structure), with a speculative decoding module attached | 1M tokens | ~167GB for the published FP4 and FP8 mixed weights; any further quantisation is a research project rather than a deployment | The efficient frontier release — strong analysis at 13B active parameters — but it needs a 192GB-class node, so for a firm of 10–50 fee-earners it is shared infrastructure rather than a first private model. |
| DeepSeek-V4.1-Flash | 552B backbone plus 196B of conditional memory (763B on disk); 8B active per token at prefill, 16B at decode | 1M tokens | ~510GB for the published mixed-precision weights; the publisher reports a KV cache footprint of roughly 890 bytes per token | A million-token multimodal release with a deliberately small active footprint, and still out of reach for a 20-partner firm — treat it as evidence of where the family is heading rather than a procurement option. |
| DeepSeek-R1-0528 | 671B total / 37B active | 128k tokens as published | ~340GB at 4-bit quantisation; ~670GB for the published FP8 build | The full reasoning model that the distils are copied from — capable enough to change how a firm triages a matter, and impractical to serve without a multi-GPU node or a private tenancy. |
| DeepSeek-V4-Pro-0813 | 1.6T total / 49B active | 1M tokens | ~1.8TB for the published weights, at roughly one byte per billion parameters in mixed precision | Capable and impractical for any law firm of this size: 1.6T parameters is a data-centre deployment, and the honest advice is to use a hosted endpoint or a private tenancy instead of buying for it. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
The reasoning line is the interesting one for legal work. Give it an analytical task — does this clause conflict with that one, what does this chronology imply, which of these three arguments is weakest — and it works through the problem visibly and often arrives somewhere useful. That makes it a strong triage and analysis assistant and a decent second opinion generator. The trade is verbosity and over-confidence in equal measure: reasoning traces are long, they contain dead ends, and the model rarely signals when it has exhausted what the evidence supports. For drafting it is serviceable rather than elegant, and its formatting discipline is weaker than Qwen's — you will spend more time reshaping output. The practical obstacle for a UK law firm is not the model, it is the surrounding questions: where it was trained, on what, by whom, under which jurisdiction's access rules, and what your client's compliance team will make of a matter document touching those weights. That is a conversation you can win with documentation and a private tenancy, but you should have it before you deploy rather than after a client asks.
Citation behaviour is the weakest of the frontier-class models here: it quotes approximately, which is dangerous in legal work, and it will merge two passages into one quotation without flagging it. Prompting for verbatim quotation with page references improves this materially, but our evaluation still shows paraphrase leakage. Treat every quotation as needing verification.
What we would use it for
- Analytical second opinions on clause interactions and argument strength
- Chronology reasoning and inconsistency spotting across a document set
- Drafting support where the output will be substantially rewritten
- Non-confidential research and internal training material
What to watch
- Approximate quotation — verify every quote character by character
- Long reasoning traces that obscure where certainty ends
- Provenance and hosting questions in client due diligence
- Mid-pack effective context for its advertised window
What it costs to run
For a firm of 10–50 fee-earners the release worth serving is DeepSeek-R1-Distill-Qwen-32B at 4-bit on one 48GB card: it is the only current member of this family that runs on a single GPU, and the long reasoning traces that are the point of the line are exactly what consumes memory during generation. Expect useful analysis throughput for a handful of concurrent users and clearly worse latency than a non-reasoning model, because the model writes many more tokens before it answers. DeepSeek-V4-Flash-0731 is a different proposition — 284B parameters and a 192GB-class node at minimum — and we would only recommend it to a firm that has already proved a firm-wide analysis workload.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $0.85 – $2.85 | $625 – $2,080 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $0.85 – $2.85 | $180 – $600 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $1.80 – $3.60 | $110 – $215 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~250 GB) | — | — | $40 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 48GB class — L40S / RTX A6000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
For the 32B distil the arithmetic favours renting for longer than in most families here: a 48GB-class workstation is indicatively $8,000–15,000, and this family's generation turns over within months, which is the strongest argument against buying for DeepSeek specifically. If you go for DeepSeek-V4-Flash-0731 instead, the hardware is a multi-GPU server costing substantially more than the 80GB class, and renting is the only sensible entry point; buying that class to run weights that may be superseded twice a year is hard to justify to a partnership. All bands are indicative rather than quotes, and a private tenancy of a re-hosted build should be compared on the same basis.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Reasoning mode with low temperature (0–0.2) for analysis; 0.3 for drafting. Do not raise temperature on the reasoning line: it does not buy variety, it buys rambling. Pin the exact release — the family moves faster than most.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single high-memory workstation (heavily quantised) | Individual analysis tasks, not shared throughput | Quantisation costs you more reasoning quality here than on smaller dense models |
| Multi-GPU server | Firm-wide analysis and triage | Realistic deployment for the general models; check throughput before promising firm-wide access |
| Private tenancy of a re-hosted build | Firms wanting the capability without running hardware | The provenance conversation moves to the host — get it in writing |
Who it suits
Analytical tasks — clause conflicts, chronology reasoning, second opinions — where a visible chain of reasoning is more useful than polished prose.
Citation-dependent work, client-facing drafting without heavy review, or firms whose clients impose provenance restrictions they cannot satisfy.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. |
Sources
- DeepSeek model releases and licences
- Open-weight legal LLMs 2026 (Presenc AI)
- Probative Co: the 2026 private AI in law guide