Command
A retrieval-tuned family that cites what it was given — under a licence position that changes by release: the research weights are non-commercial, and the newest release is Apache 2.0.
Tier C — pilot only, and only on the Apache-licensed release. On behaviour alone this family would sit near the top of the index: its instinct to answer from the passages it was given, and to name them, is closer to what legal work needs than several models we score higher. It stays conditional for a practical reason instead: the commercially usable release is new and unproven on legal tasks, and the rest of the family cannot lawfully touch client data. So pilot it — scoped to command-a-plus-05-2026, with the licence tag checked at the point of download rather than assumed.
Specifications, as published
| Publisher | Cohere |
| Family | The Command R line of open-weight retrieval-oriented models, alongside the publisher's larger hosted Command models |
| Parameters | Large dense releases rather than a small-model range; the open weights are sized for server deployment, not a workstation |
| Context | Long context advertised and designed around retrieved passages; effective context in our tasks is mid-pack, with the usual loss in the middle of very long prompts |
| Licence | Split by release, and the split is the story. The research line (R, R+, A, A-reasoning) is published under CC-BY-NC-4.0 with the publisher's acceptable-use policy layered on top and a gated non-commercial download — a law firm cannot serve those weights on client matters, because client work is commercial use. The newest release, command-a-plus-05-2026, is published under Apache 2.0, which makes it the first Command open weight a firm may lawfully run on client work. A family whose licence changes with the tag is a family where downloading the wrong file is a compliance incident. |
| Weights | Downloadable, though re-hosting is constrained by the licence terms |
| Release | The open-weight line lags the publisher's hosted models, which are the better-performing and commercially available option |
| Licence posture | Split by release, and the split is the story. The research line (R, R+, A, A-reasoning) is published under CC-BY-NC-4.0 with the publisher's acceptable-use policy layered on top and a gated non-commercial download — a law firm cannot serve those weights on client matters, because client work is commercial use. The newest release, command-a-plus-05-2026, is published under Apache 2.0, which makes it the first Command open weight a firm may lawfully run on client work. A family whose licence changes with the tag is a family where downloading the wrong file is a compliance incident. |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| c4ai-command-r7b-12-2024 | 7B dense (8.0B total parameters in the released weights) | 128K tokens | ~5GB at 4-bit, ~9GB at 8-bit | The small release in the family, sized for one workstation, and published under the same CC-BY-NC licence as the rest of the research line: a UK firm of 10–50 fee-earners could run it comfortably, but only for non-commercial evaluation. Client work is outside its terms, and no deployment design converts that. |
| c4ai-command-r-v01 | 35B dense | 128K tokens | ~20GB at 4-bit, ~37GB at 8-bit | The retrieval-oriented release that set the family's reputation for grounding answers in supplied passages. Licence: CC-BY-NC-4.0 with a further acceptable-use policy on top — non-commercial only, so a firm may benchmark it and may not put client material through it, whatever the hardware allows. |
| c4ai-command-r-plus-08-2024 | 104B dense | 128K tokens | ~55GB at 4-bit, ~105GB at 8-bit | The larger partner of the R line and the release most firms first met this family through. Licence: CC-BY-NC-4.0 — non-commercial and acceptable-use restricted, so it is an evaluation and comparison point only. At 4-bit it fits a single 141GB-class card, which is a data-centre purchase for a model a firm is not permitted to use commercially. |
| c4ai-command-a-03-2025 | 111B dense | 256K tokens | ~60GB at 4-bit, ~112GB at 8-bit | The current-generation flagship of the research line, with a 256K window designed around retrieved material. Licence: CC-BY-NC-4.0 — non-commercial, and the acceptable-use policy applies in addition. Capable, and for a 20-partner firm doubly out of reach: the licence bars client work and the 4-bit footprint needs a 141GB-class card. |
| command-a-reasoning-08-2025 | 111B dense, with switchable reasoning mode | 256K tokens, 32K output | ~60GB at 4-bit, ~112GB at 8-bit | A reasoning build of the same 111B model, with reasoning that can be turned off for lower latency. Licence: CC-BY-NC-4.0, non-commercial with an acceptable-use policy — the reasoning behaviour is interesting to benchmark against models you can actually deploy, and unusable for client work on these terms. |
| command-a-plus-05-2026 | 218B total / 25B active (MoE — 128 experts, 8 active per token plus one shared expert) | 128K input, 64K output | ~110GB at 4-bit, ~218GB at 8-bit | The current release and the exception that matters: published under Apache 2.0, which makes it the only Command open weights a firm may lawfully run on client work — a different position from every release above it, and one worth re-reading per release rather than assuming. It also carries vision input and is the largest model in this table, needing a 141GB-class card at 4-bit: capable, and impractical for a 20-partner firm without a deliberate infrastructure decision. The publisher's hosted line is a separate commercial product on its own terms. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
Command is the entry that shows why the licence criterion carries the weight it does, because on behaviour it would otherwise be one of the more interesting models in this index. It is built for retrieval, and it shows. Given a question and a set of passages, its default instinct is to ground the answer in the passages, structure the response around them, and offer citations alongside the claims — the behaviour that most models in this index have to be instructed into and that Command does partly by default. In our evaluation tasks that made it the fastest model to produce a defensible first-pass research note: it kept to the retrieved material more reliably than the larger general models, it produced a cleaner separation between what the documents said and what the model inferred, and it was noticeably better at keeping a document's own terminology rather than substituting a synonym. Summarisation of a document set is competent and moderately concise. Its citation behaviour is a qualified strength rather than a clean one: the passages it selects are usually the right ones, but the span it attaches to a claim is often wider than the proposition requires, so the citation points at the correct page and not precisely at the supporting sentence. A reviewer checking the reference will find the topic and will have to read on to find the words. That is much better than citing nothing, and it is not the same as a verified quotation. Multilingual coverage is broad and, in our tasks, better than most open weights on European and Asian languages in summarisation and question answering, with the caveat that register holds up better than legal terminology: it will translate a French commercial letter into natural English and lose the defined terms in the process, so a firm working across languages should expect to reattach the terminology by hand. Where it is weaker is drafting. Given a brief, it produces orderly prose that is a little flat, and its structure can read as a template — the same headings with the content swapped. It also has a mild tendency toward consultant register, restating the question before answering it, which a fee-earner will edit out of every output. That register is survivable in an internal note and less so in a client-facing letter, where it reads as padding and occasionally as evasion, and it is not something a prompt entirely removes. On reasoning, the open weights show their age against the current frontier: multi-step analysis is serviceable on triage and allocation, and thin on questions of construction or weighing competing arguments. But none of this is the finding that matters. The finding is the licence. The open weights are published on non-commercial terms. Client work is commercial use. A UK law firm deploying these weights on a client matter, or on internal work that supports a commercial practice, is outside the terms of the grant, and no amount of internal policy, private hosting or prompt discipline changes that. The publisher's commercially licensed hosted models are a different product on different terms; the open weights are the free version with the free version's restrictions, and the difference between the two is not a technicality a firm can waive. Read the licence before you read the scorecard.
For a model this size the abstention behaviour is comparatively good: given passages that do not answer the question, Command is more likely than most open weights to say so, particularly when it is asked to answer only from the supplied passages. Its citations are correctly aimed but loosely bounded — the span attached to a claim is often broader than the claim — so quotations need verifying character by character even when the reference is right. Requiring the exact supporting sentence, quoted, before the conclusion tightens this materially. All of which is academic until the licence question is resolved.
What we would use it for
- Contract and clause extraction on client matters — once the Apache-licensed release has been evaluated on your own documents
- Retrieval-grounded question answering over a closed document set, where naming its sources is the point
- Chronology and bundle summarisation with citations back to the record
- Non-confidential evaluation of the research-line weights (never client material)
- Not: any use of the CC-BY-NC releases on client work, however good the output looks
What to watch
- Licence differs by release — the research line is CC-BY-NC-4.0 and cannot be used on client work; only command-a-plus-05-2026 is Apache 2.0
- Verify the licence tag on the exact file you serve, not on the family page
- The Apache release is new: it has no track record on legal tasks and should be evaluated before it touches client documents
- Its genuinely strong retrieval behaviour is the reason to do that evaluation rather than skipping the family entirely
- Mixed-licence families punish loose model management — know which weights are on which server
What it costs to run
The only release in this family we would recommend serving is command-a-plus-05-2026, because it is the only one published on terms that permit client work; at 4-bit it fits a single 141GB-class card, and that is the minimum sensible configuration for a firm of 10–50 fee-earners. Expect modest concurrency from one card on 128K-input prompts — enough for a handful of fee-earners queuing research notes, not for a department calling it per paragraph. Latency is governed by prefill on long retrieved contexts, so the first answer on a large bundle is slow and later extraction calls are quicker. If a firm wants this behaviour commercially without operating hardware, the publisher's hosted service is the simpler route, and note that the licence finding applies to the downloadable weights rather than to that service.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $3.75 – $6.88 | $2,740 – $5,025 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $3.75 – $6.88 | $790 – $1,445 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $5.10 – $9.00 | $305 – $540 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~500 GB) | — | — | $75 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 141GB class — H200. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
For this family the buying-against-renting question is secondary to the licence question, because only one release is open to client work and it is large. A 141GB-class card is a data-centre part and we do not publish an indicative capex band for that class here; the nearest indicative figure we give is $25,000–60,000 for an 80GB-class server, and a 141GB-class machine sits above that. Renting is the sensible pattern while the workload is unproven, and for most firms of this size the honest conclusion is that this family is a hosted purchase rather than a private deployment. All of these figures are indicative.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Extraction from supplied passages: temperature 0, top_p 0.8, with an instruction to answer only from the context. Grounded question answering: 0.1–0.2, requiring the supporting sentence quoted before the answer. Summarisation: 0.1. Drafting: 0.3–0.4, expecting structural editing. As with every entry, pin the exact release and the licence revision you reviewed — for this model the licence is the field that matters most, and it is the field most likely to change. Note that all of these settings are for evaluation work only while the non-commercial terms stand. Confirm the licence tag on the exact release you are serving: in this family the sampling settings are the same across releases, but the legal position is not.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single high-memory workstation (heavily quantised) | Individual evaluation and prototype work | Quantisation is a real cost here; evaluate at the precision you intend to run |
| One or two GPU servers | The configuration the model was designed for — retrieval-grounded answering at firm scale | The deployment that would make sense if the licence permitted it, which on the published terms it does not |
| Publisher's hosted service instead of the weights | Firms wanting this behaviour commercially | A different product on different terms — the licence finding applies to the open weights, not to the hosted line |
Who it suits
Firms willing to evaluate a single Apache-licensed release on their own matters, and to keep the rest of the family out of client work.
Firms that cannot control which weights are deployed — a mixed-licence family punishes loose model management.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. Scored Tier D on the published non-commercial licence for the open weights, notwithstanding competitive retrieval behaviour. |
| Sep 2026 | Re-scored from 55 (tier D) to 71 (tier C) after the publisher's August 2026 Apache-2.0 release. The research line remains non-commercial and unusable for client work; the family now contains one release that is not. Licence raised from 2 to 18 — the deduction reflects mixed-licence deployment risk rather than the Apache release itself. |
Sources
- Cohere open-weight releases (Cohere For AI on Hugging Face)
- Command model card and licence (Cohere For AI)
- Open-weight licence landscape 2026 (Presenc AI)
- Open-weight legal LLMs 2026 (Presenc AI)