AI Tools Index / Open models / Kimi
Moonshot AI · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

Kimi

The Kimi K-series mixture-of-experts models

A long-context mixture-of-experts family that holds up across very large matter files and follows complex drafting briefs closely, with an attribution-flavoured licence and a hardware appetite that make it a deliberate purchase.

Our verdict

Tier B — conditional, and the pick when the document will not fit anywhere else. The long-context behaviour is the strongest in this index and the licence is permissive enough for most firms today, with an attribution condition to track as you grow. The conditions are verbosity, which you control by asking, reasoning that is confident beyond its evidence, which you control by supervising, and a hardware appetite that removes the cheaper configurations from consideration.

Specifications, as published

PublisherMoonshot AI
FamilyThe Kimi K-series mixture-of-experts models
ParametersLarge mixture-of-experts releases with a small fraction of parameters active per token, sized for multi-GPU server deployment
ContextVery long context is the family's design centre, and in our tasks the long window survives longer than most; performance still degrades in the middle of the largest prompts, so retrieval structure still matters
LicenceA modified MIT licence: permissive in substance, with an added attribution condition that applies to large-scale commercial products — more permissive than the community-style licences, less permissive than plain MIT
WeightsDownloadable, with quantised builds available though the model is large enough that quantisation is a compromise rather than a solution
ReleaseNew generations on a fast cycle, with the mixture-of-experts architecture carried across the line
Licence postureA modified MIT licence: permissive in substance, with an added attribution condition that applies to large-scale commercial products — more permissive than the community-style licences, less permissive than plain MIT

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
Kimi-VL-A3B-Instruct16B total / 3B active (MoE, with a vision encoder)128K tokens~9GB at 4-bit, ~17GB at 8-bitThe small end of the family and the only Kimi release that sits comfortably on one workstation card. MIT-licensed and genuinely runnable by a firm of 10–50 fee-earners for extraction from images and short documents, but 3B active parameters is a modest engine and it is not the reason to choose this family.
Kimi-Linear-48B-A3B-Instruct48B total / 3B active (MoE, hybrid linear attention)1M tokens~27GB at 4-bit, ~50GB at 8-bitThe release we would actually serve: MIT-licensed, 48B total with 3B active, and a linear-attention design that cuts the key-value cache by roughly three quarters so a very long window becomes affordable. A single 48GB-class card carries it at 4-bit, which makes it the cheapest credible route to long-document review inside a 20-partner firm's own perimeter, provided context is capped well below the maximum.
Kimi-Dev-72B72B denseas published~40GB at 4-bit, ~73GB at 8-bitA dense 72B release tuned for software issue-resolution rather than for prose, extraction or document work, and the family's only single-card dense option at a useful size. MIT-licensed and realistically runnable by a mid-sized firm on a 48GB-class card, but its purpose is code and we would not deploy it for legal tasks.
Kimi-K2-Instruct1T total / 32B active (MoE, 384 experts)128K tokens~500GB at 4-bit, ~1TB at 8-bit — multi-node, not single-serverThe release that established this family's reputation for holding a large matter file in view while following elaborate instructions. Licence: a modified MIT licence that is permissive in substance with an attribution condition once usage reaches commercial scale — the version to track as a firm grows. At roughly 500GB of weights even at 4-bit it is capable and impractical for a 20-partner firm.
Kimi-K2-Thinking1T total / 32B active (MoE)256K tokens~500GB at 4-bit, ~1TB at 8-bit — multi-node, not single-serverThe reasoning build of the K2 architecture, with a doubled 256K window and visible traces that are useful for triage but confident beyond their evidence. Modified MIT licence, same attribution condition. Multi-node hardware, so beyond any firm of 10–50 fee-earners without its own infrastructure.
Kimi-K2.61T total / 32B active (MoE, with a 400M vision encoder)256K tokens~500GB at 4-bit, ~1TB at 8-bit — multi-node, not single-serverThe current general release of the K2 line, multimodal and carrying the family's long-context behaviour forward. Modified MIT licence with an attribution condition that applies once a product passes a very high usage or revenue threshold — a firm of any realistic size will never trigger it, but it is a term and should be recorded. Hardware remains data-centre scale, so this is a comparison point rather than a deployment for a 20-partner firm.
Kimi-K32.8T total / 104B active (MoE — 16 of 896 experts active per token)1M tokens~1.4TB at 4-bit, ~2.8TB at 8-bit — data-centre scaleThe frontier release and the newest in the family: natively multimodal across text, images and video, with a 1M-token window and the largest parameter count in this index. Note that the licence has moved — it is published under the bespoke Kimi K3 Licence rather than the modified MIT of the K2 line, so the terms must be read again rather than assumed. Roughly 1.4TB of weights at 4-bit means a multi-node cluster: capable beyond question, and impractical for any firm without its own data centre.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

Kimi's proposition for legal work is narrow and genuinely useful: it keeps more of a large matter file in view than anything else we have assessed, and it follows elaborate instructions while doing so. In our evaluation tasks it was the model most likely to hold a long document set coherently — cross-referencing an agreement against a schedule, a chronology and a set of correspondence without losing the thread — and the least likely to answer from the instruction rather than the material when the prompt grew very long. That is the behaviour that matters most in disclosure-heavy and transaction-heavy work, and it is worth paying for if the workload contains it. Instruction following on drafting is strong and detailed: given a precedent, a house style, a defined-terms list and several constraints, it satisfied them consistently, including negative constraints like do not introduce facts not in the record, which many models quietly ignore in the third act of a long document. Output is long. Where Mistral is terse and Gemma is helpful, Kimi is thorough, and a fee-earner's first week is largely spent learning to ask for shorter answers. Ask for a summary of a bundle and you get a competent summary with a structure you did not request, examples you did not need and a closing section restating the introduction. Ask explicitly for a page limit and it complies, but the instruction has to be given. Extraction is accurate on long documents and slower than the smaller models here, which is a straightforward trade: it is the model to use when the document will not fit anywhere else rather than the model to use for a thousand short letters. Throughput matters for review work, where the model is called once per document set rather than once per paragraph, and the family's speed is adequate for that pattern rather than for bulk processing. Its reasoning is fluent and well organised rather than rigorous; the traces — where a reasoning mode is used — are more readable than most and show their working, but they are confident throughout, including in the passages where the model is guessing, and it does not reliably mark where the evidence stops. Bilingual behaviour is strong, with Chinese and English handled to a high standard and cross-language work comparable to the best entries here, and European languages adequate rather than excellent. The two practical obstacles are the licence and the hardware. The licence is permissive but not unconditional: the attribution condition bites at scale, which for a growing firm is a forward-looking question rather than a present one, and it should be recorded in the model register alongside the release. The hardware requirement is real — these are multi-GPU models, and the quantised builds that fit on a single server lose the long-context reliability that is the entire reason to choose the family. A firm evaluating Kimi should therefore test the quantised configuration it can actually afford, on its own long documents, before concluding anything from the full-precision documentation.

evidence and abstention

Quotation fidelity is the best of the frontier-class entries here. Given a long document set and an instruction to quote verbatim with a reference, Kimi reproduces passages closely and attaches the correct reference, and it keeps a quotation intact when assembling several into one answer rather than merging them. Abstention is good on silent corpora with an explicit refusal rule, and merely average on adjacent material, where it will extend a nearby passage to cover the gap. Requiring the supporting quotation before the conclusion, and a separate list of unanswered questions, is the combination that worked best in our tasks.

What we would use it for

  • Long-document review: agreement against schedule, chronology and correspondence in one prompt
  • Cross-referencing and inconsistency spotting across a large matter file
  • Complex drafting briefs with a precedent, house style and several constraints
  • Summarisation of substantial bundles where completeness matters more than brevity
  • Bilingual and cross-border work, including Chinese-language documents

What to watch

  • Verbose output that must be constrained explicitly at every call
  • Confident reasoning throughout, including where the evidence stops
  • An attribution condition in the licence that applies at commercial scale — track it as the firm grows
  • Multi-GPU requirements, with quantised builds losing the long-context reliability that justifies the choice
  • Effective context still degrading in the middle of the very largest prompts

What it costs to run

The realistic configuration for a firm of 10–50 fee-earners is one 48GB-class card serving Kimi-Linear-48B-A3B-Instruct at 4-bit with context deliberately capped below the maximum, and the retrieval layer doing the work of finding the right passages rather than the window holding everything. That gives a few fee-earners concurrent long-document work, which matches how review tasks are queued in practice. The K2 line needs roughly 500GB of weights at 4-bit and K3 roughly 1.4TB, so neither is a single-server option however good the throughput looks on paper. Latency on long prompts is dominated by prefill rather than by active parameters, so expect the first answer on a large bundle to take noticeably longer than the extraction calls that follow, and size the deployment for queueing rather than for interactive chat.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$0.85 – $2.85$625 – $2,080Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$0.85 – $2.85$180 – $600The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$1.80 – $3.60$110 – $215Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~700 GB)$105Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 48GB class — L40S / RTX A6000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

Renting is the right pattern for this family, because the release we would serve needs a 48GB-class card — an indicative $8,000–15,000 as a workstation — while the releases the family is known for need a multi-card server well above the 80GB-class band of $25,000–60,000, for which we do not publish a figure here. A firm that buys a 48GB-class workstation can serve the linear-attention release through the working day and stop paying for rented time, which is the point at which owning starts to win. We would not advise buying anything larger for this family until an evaluation on the firm's own longest real documents shows the smaller release failing. All of these figures are indicative.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Long-document review: temperature 0–0.1, with the document set ordered deliberately and the instruction placed after the material. Extraction: 0, top_p 0.8, schema-bound, one document class per call. Drafting: 0.3 with an explicit length limit and a list of prohibited additions. Summarisation: 0.1 with a stated maximum length, because this family will otherwise triple it. Keep reasoning modes at low temperature for analysis. Pin the exact release and the quantisation, and record the licence version — the attribution condition is a term, not a footnote.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
Single high-memory workstation (heavily quantised)Evaluation and individual long-document tasks, not shared throughputExpect the long-context reliability that defines this family to suffer; test this configuration before you rely on it
Multi-GPU serverFirm-wide long-document review and complex drafting with retrievalThe configuration the family is designed for, and the throughput pattern matters more than raw speed for review work
Larger multi-GPU deployment or private tenancyTransaction and disclosure-heavy workloads across several teamsJustify it with evaluation evidence on your longest real documents, and check the licence position as usage scales

Who it suits

Good fit

Firms with large transaction or disclosure files that need one model to hold the whole picture, and those with bilingual Chinese and English workloads.

Poor fit

Firms wanting a single-workstation deployment, high-volume processing of short documents, or anyone who will not constrain output length and enforce citation rules.

Review history

DateChange
Sep 2026First entry.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.