Kimi
A long-context mixture-of-experts family that holds up across very large matter files and follows complex drafting briefs closely, with an attribution-flavoured licence and a hardware appetite that make it a deliberate purchase.
Tier B — conditional, and the pick when the document will not fit anywhere else. The long-context behaviour is the strongest in this index and the licence is permissive enough for most firms today, with an attribution condition to track as you grow. The conditions are verbosity, which you control by asking, reasoning that is confident beyond its evidence, which you control by supervising, and a hardware appetite that removes the cheaper configurations from consideration.
Specifications, as published
| Publisher | Moonshot AI |
| Family | The Kimi K-series mixture-of-experts models |
| Parameters | Large mixture-of-experts releases with a small fraction of parameters active per token, sized for multi-GPU server deployment |
| Context | Very long context is the family's design centre, and in our tasks the long window survives longer than most; performance still degrades in the middle of the largest prompts, so retrieval structure still matters |
| Licence | A modified MIT licence: permissive in substance, with an added attribution condition that applies to large-scale commercial products — more permissive than the community-style licences, less permissive than plain MIT |
| Weights | Downloadable, with quantised builds available though the model is large enough that quantisation is a compromise rather than a solution |
| Release | New generations on a fast cycle, with the mixture-of-experts architecture carried across the line |
| Licence posture | A modified MIT licence: permissive in substance, with an added attribution condition that applies to large-scale commercial products — more permissive than the community-style licences, less permissive than plain MIT |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| Kimi-VL-A3B-Instruct | 16B total / 3B active (MoE, with a vision encoder) | 128K tokens | ~9GB at 4-bit, ~17GB at 8-bit | The small end of the family and the only Kimi release that sits comfortably on one workstation card. MIT-licensed and genuinely runnable by a firm of 10–50 fee-earners for extraction from images and short documents, but 3B active parameters is a modest engine and it is not the reason to choose this family. |
| Kimi-Linear-48B-A3B-Instruct | 48B total / 3B active (MoE, hybrid linear attention) | 1M tokens | ~27GB at 4-bit, ~50GB at 8-bit | The release we would actually serve: MIT-licensed, 48B total with 3B active, and a linear-attention design that cuts the key-value cache by roughly three quarters so a very long window becomes affordable. A single 48GB-class card carries it at 4-bit, which makes it the cheapest credible route to long-document review inside a 20-partner firm's own perimeter, provided context is capped well below the maximum. |
| Kimi-Dev-72B | 72B dense | as published | ~40GB at 4-bit, ~73GB at 8-bit | A dense 72B release tuned for software issue-resolution rather than for prose, extraction or document work, and the family's only single-card dense option at a useful size. MIT-licensed and realistically runnable by a mid-sized firm on a 48GB-class card, but its purpose is code and we would not deploy it for legal tasks. |
| Kimi-K2-Instruct | 1T total / 32B active (MoE, 384 experts) | 128K tokens | ~500GB at 4-bit, ~1TB at 8-bit — multi-node, not single-server | The release that established this family's reputation for holding a large matter file in view while following elaborate instructions. Licence: a modified MIT licence that is permissive in substance with an attribution condition once usage reaches commercial scale — the version to track as a firm grows. At roughly 500GB of weights even at 4-bit it is capable and impractical for a 20-partner firm. |
| Kimi-K2-Thinking | 1T total / 32B active (MoE) | 256K tokens | ~500GB at 4-bit, ~1TB at 8-bit — multi-node, not single-server | The reasoning build of the K2 architecture, with a doubled 256K window and visible traces that are useful for triage but confident beyond their evidence. Modified MIT licence, same attribution condition. Multi-node hardware, so beyond any firm of 10–50 fee-earners without its own infrastructure. |
| Kimi-K2.6 | 1T total / 32B active (MoE, with a 400M vision encoder) | 256K tokens | ~500GB at 4-bit, ~1TB at 8-bit — multi-node, not single-server | The current general release of the K2 line, multimodal and carrying the family's long-context behaviour forward. Modified MIT licence with an attribution condition that applies once a product passes a very high usage or revenue threshold — a firm of any realistic size will never trigger it, but it is a term and should be recorded. Hardware remains data-centre scale, so this is a comparison point rather than a deployment for a 20-partner firm. |
| Kimi-K3 | 2.8T total / 104B active (MoE — 16 of 896 experts active per token) | 1M tokens | ~1.4TB at 4-bit, ~2.8TB at 8-bit — data-centre scale | The frontier release and the newest in the family: natively multimodal across text, images and video, with a 1M-token window and the largest parameter count in this index. Note that the licence has moved — it is published under the bespoke Kimi K3 Licence rather than the modified MIT of the K2 line, so the terms must be read again rather than assumed. Roughly 1.4TB of weights at 4-bit means a multi-node cluster: capable beyond question, and impractical for any firm without its own data centre. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
Kimi's proposition for legal work is narrow and genuinely useful: it keeps more of a large matter file in view than anything else we have assessed, and it follows elaborate instructions while doing so. In our evaluation tasks it was the model most likely to hold a long document set coherently — cross-referencing an agreement against a schedule, a chronology and a set of correspondence without losing the thread — and the least likely to answer from the instruction rather than the material when the prompt grew very long. That is the behaviour that matters most in disclosure-heavy and transaction-heavy work, and it is worth paying for if the workload contains it. Instruction following on drafting is strong and detailed: given a precedent, a house style, a defined-terms list and several constraints, it satisfied them consistently, including negative constraints like do not introduce facts not in the record, which many models quietly ignore in the third act of a long document. Output is long. Where Mistral is terse and Gemma is helpful, Kimi is thorough, and a fee-earner's first week is largely spent learning to ask for shorter answers. Ask for a summary of a bundle and you get a competent summary with a structure you did not request, examples you did not need and a closing section restating the introduction. Ask explicitly for a page limit and it complies, but the instruction has to be given. Extraction is accurate on long documents and slower than the smaller models here, which is a straightforward trade: it is the model to use when the document will not fit anywhere else rather than the model to use for a thousand short letters. Throughput matters for review work, where the model is called once per document set rather than once per paragraph, and the family's speed is adequate for that pattern rather than for bulk processing. Its reasoning is fluent and well organised rather than rigorous; the traces — where a reasoning mode is used — are more readable than most and show their working, but they are confident throughout, including in the passages where the model is guessing, and it does not reliably mark where the evidence stops. Bilingual behaviour is strong, with Chinese and English handled to a high standard and cross-language work comparable to the best entries here, and European languages adequate rather than excellent. The two practical obstacles are the licence and the hardware. The licence is permissive but not unconditional: the attribution condition bites at scale, which for a growing firm is a forward-looking question rather than a present one, and it should be recorded in the model register alongside the release. The hardware requirement is real — these are multi-GPU models, and the quantised builds that fit on a single server lose the long-context reliability that is the entire reason to choose the family. A firm evaluating Kimi should therefore test the quantised configuration it can actually afford, on its own long documents, before concluding anything from the full-precision documentation.
Quotation fidelity is the best of the frontier-class entries here. Given a long document set and an instruction to quote verbatim with a reference, Kimi reproduces passages closely and attaches the correct reference, and it keeps a quotation intact when assembling several into one answer rather than merging them. Abstention is good on silent corpora with an explicit refusal rule, and merely average on adjacent material, where it will extend a nearby passage to cover the gap. Requiring the supporting quotation before the conclusion, and a separate list of unanswered questions, is the combination that worked best in our tasks.
What we would use it for
- Long-document review: agreement against schedule, chronology and correspondence in one prompt
- Cross-referencing and inconsistency spotting across a large matter file
- Complex drafting briefs with a precedent, house style and several constraints
- Summarisation of substantial bundles where completeness matters more than brevity
- Bilingual and cross-border work, including Chinese-language documents
What to watch
- Verbose output that must be constrained explicitly at every call
- Confident reasoning throughout, including where the evidence stops
- An attribution condition in the licence that applies at commercial scale — track it as the firm grows
- Multi-GPU requirements, with quantised builds losing the long-context reliability that justifies the choice
- Effective context still degrading in the middle of the very largest prompts
What it costs to run
The realistic configuration for a firm of 10–50 fee-earners is one 48GB-class card serving Kimi-Linear-48B-A3B-Instruct at 4-bit with context deliberately capped below the maximum, and the retrieval layer doing the work of finding the right passages rather than the window holding everything. That gives a few fee-earners concurrent long-document work, which matches how review tasks are queued in practice. The K2 line needs roughly 500GB of weights at 4-bit and K3 roughly 1.4TB, so neither is a single-server option however good the throughput looks on paper. Latency on long prompts is dominated by prefill rather than by active parameters, so expect the first answer on a large bundle to take noticeably longer than the extraction calls that follow, and size the deployment for queueing rather than for interactive chat.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $0.85 – $2.85 | $625 – $2,080 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $0.85 – $2.85 | $180 – $600 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $1.80 – $3.60 | $110 – $215 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~700 GB) | — | — | $105 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 48GB class — L40S / RTX A6000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
Renting is the right pattern for this family, because the release we would serve needs a 48GB-class card — an indicative $8,000–15,000 as a workstation — while the releases the family is known for need a multi-card server well above the 80GB-class band of $25,000–60,000, for which we do not publish a figure here. A firm that buys a 48GB-class workstation can serve the linear-attention release through the working day and stop paying for rented time, which is the point at which owning starts to win. We would not advise buying anything larger for this family until an evaluation on the firm's own longest real documents shows the smaller release failing. All of these figures are indicative.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Long-document review: temperature 0–0.1, with the document set ordered deliberately and the instruction placed after the material. Extraction: 0, top_p 0.8, schema-bound, one document class per call. Drafting: 0.3 with an explicit length limit and a list of prohibited additions. Summarisation: 0.1 with a stated maximum length, because this family will otherwise triple it. Keep reasoning modes at low temperature for analysis. Pin the exact release and the quantisation, and record the licence version — the attribution condition is a term, not a footnote.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single high-memory workstation (heavily quantised) | Evaluation and individual long-document tasks, not shared throughput | Expect the long-context reliability that defines this family to suffer; test this configuration before you rely on it |
| Multi-GPU server | Firm-wide long-document review and complex drafting with retrieval | The configuration the family is designed for, and the throughput pattern matters more than raw speed for review work |
| Larger multi-GPU deployment or private tenancy | Transaction and disclosure-heavy workloads across several teams | Justify it with evaluation evidence on your longest real documents, and check the licence position as usage scales |
Who it suits
Firms with large transaction or disclosure files that need one model to hold the whole picture, and those with bilingual Chinese and English workloads.
Firms wanting a single-workstation deployment, high-volume processing of short documents, or anyone who will not constrain output length and enforce citation rules.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. |
Sources
- Kimi K-series releases (Moonshot AI)
- Kimi model cards and licence (Moonshot AI on Hugging Face)
- Kimi K2 model card (Moonshot AI)
- Open-weight licence landscape 2026 (Presenc AI)