AI Tools Index / Open models / Nemotron
NVIDIA · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

Nemotron

The Nemotron line — small mixture-of-experts releases, mid-size models in the low hundreds of billions of total parameters, and a very large flagship in the same pattern

Frontier-class reasoning and the strongest summarisation we have measured from open weights, with disclosed training datasets — and a publisher-specific licence that a compliance team will want read rather than skimmed.

Our verdict

Tier B — conditional, and the strongest summariser in this index. The caveats are structural rather than behavioural: a licence that is open in substance but publisher-defined, and hardware requirements that scale steeply with the releases you actually want. Pilot it on reading-heavy work, have the licence read properly, and keep the reasoning monologue out of anything a client sees.

Specifications, as published

PublisherNVIDIA
FamilyThe Nemotron line — small mixture-of-experts releases, mid-size models in the low hundreds of billions of total parameters, and a very large flagship in the same pattern
ParametersA family rather than a single model: small mixture-of-experts releases with a few billion active parameters, mid-size releases in the low hundreds of billions of total parameters with a small active fraction, and a much larger flagship in the same architecture
ContextPublished context of up to a million tokens on the larger releases — the longest claim in this index, and one we treat with the usual scepticism until it survives a real matter file
LicenceNVIDIA Nemotron Open Model License — a publisher-specific open model licence rather than a general-purpose permissive one; read the terms in full before deployment
WeightsDownloadable in several precisions, including quantised formats published by the vendor itself
ReleaseFast-moving family with overlapping generations in circulation; version pinning is not optional
Licence postureNVIDIA Nemotron Open Model License — a publisher-specific open model licence rather than a general-purpose permissive one; read the terms in full before deployment

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
NVIDIA-Nemotron-3-Nano-4B-BF163.97B dense — Mamba-2 and attention hybrid, no experts262,144 tokens, published as a context length of up to 262K~8GB at BF16, ~3GB at 4-bitAn edge-targeted small model built for on-device agents rather than casework, so a firm of 10–50 fee-earners could run it on one workstation card but should use it for classification, triage and formatting chores rather than drafting.
NVIDIA-Nemotron-3-Nano-30B-A3B-BF1630B total, ~3.5B active per token (128 routed experts, 6 active per token)262,144 tokens in the published configuration; the card describes a default of 256K, extendable to 1M at higher memory cost~60GB at BF16; the sibling NVFP4 checkpoint of the same release fits in ~20GBThe realistic entry point to this family for a mid-sized firm: a sparse mixture of experts whose small active fraction keeps it quick on one 80GB card, served under NVIDIA's own Nemotron Open Model License rather than Apache-2.0.
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF1630B total, 3B active per token — mixture-of-experts hybrid (Mamba-2 and attention) with multi-token prediction262,144 tokens default and up to 1M; the publisher validates 256K on a single 80GB H100 in BF16 and 1M in the NVFP4 build~60GB at BF16; the NVFP4 build is published as a single-GPU deployment on one 80GB H100 or a DGX SparkThe release we would actually serve for a firm of this size — current, fast for its class, and published under the OpenMDW-1.1 licence rather than NVIDIA's own Nemotron terms, so it is one more licence to read before deployment.
NVIDIA-Nemotron-3-Super-120B-A12B-BF16120B total, 12B active per token — hybrid LatentMoE with multi-token predictionUp to 1M tokens as published, with 256K in the default configuration~240GB at BF16 (8× 80GB is the publisher's stated minimum); ~66GB in the NVFP4 checkpoint, published as a single-GPU deployment on one B200 or DGX SparkPublished under the NVIDIA Nemotron Open Model License, this is the mid-size flagship and needs a multi-GPU node at BF16, so it is beyond what a 10–50 fee-earner firm should run on-premise.
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16550B total, 55B active per token — hybrid LatentMoE with multi-token predictionUp to 1M tokens as published, with 256K in the default configuration~1.1TB at BF16, published with a minimum of 8× GB200/B200/GB300/B300, 16× H100 or 8× H200; ~300GB in the NVFP4 checkpointThe flagship of the line and outside the reach of a mid-sized UK firm on-premise, since it needs a multi-node cluster whatever licence it is published under.
Llama-3_3-Nemotron-Super-49B-v149B dense decoder-only, derived from Llama-3.3-70B-Instruct by neural architecture search131,072 tokens (published as 128K)~98GB at BF16, published for 2× 80GB GPUs or a single high-memory card at reduced precisionA previous-generation release that carries the Llama 3.3 Community Licence on top of NVIDIA's own terms, so a firm of 10–50 fee-earners can run it on one 80GB card but inherits two licences to read and an older capability ceiling.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

Nemotron's standout legal behaviour is summarisation. Compressing a long judgment, a bundle or a disclosure set into something accurate and readable is the task where open models most often fall away from frontier quality, and the larger Nemotron releases are the strongest we have measured at it. That matters because summarisation is the single most common thing a fee-earner asks a model to do. The same strength shows up in the independent Swiss legal evaluation work we cite, where a Nemotron release was the strongest summariser in the field. It is also unusually good at holding a long document in view: given a full set of disclosed documents and asked for a chronology with cross-references, it produces something usable and invents noticeably less than the models below it. Instruction following is strong and tolerant of complexity. Multi-clause instructions — extract these six fields, flag any clause that conflicts with the schedule, and produce a table containing only the conflicts — survive contact with the model rather than collapsing into the first clause. Extraction into a schema is reliable at temperature 0 with a defined output shape. The difference from a dense mid-size model is that Nemotron will sometimes add a field you did not ask for because it noticed something relevant; that is a feature for triage and an irritation in a fixed pipeline. Drafting is capable but has a distinctive tic: it reasons in unmarked prose. The Swiss evaluation work notes this explicitly — its internal monologue is not wrapped in the reasoning tags other families use, so a pipeline author cannot strip it with a parser. In practice you get an internal-sounding preamble attached to a good answer, and you have to remove it by instruction rather than by post-processing. Ask for a letter of advice and you get a well-organised, slightly discursive draft with the working still visible in front of it. That narration is genuinely useful for review, because it shows the route to an answer and exposes where an assumption entered. It also cuts both ways. Nemotron will occasionally talk itself into a position and then state it more firmly than the material supports, and unlike a dedicated reasoning model it does not reliably announce when it has run out of evidence. Over-assertion in this family presents as fluency rather than as error: the answer reads well and the weakest link is unmarked. The licence is the reason this is not the top entry in this index for a UK firm. It is a publisher-specific open model licence with the familiar asymmetry — broad permission to use, conditions defined by the publisher, and no community process standing behind the terms. Nothing here is hostile to a law firm, but a firm that has signed up to a client's AI policy needs to read the terms rather than assume they match Apache 2.0. Week one, three observations. The smaller mixture-of-experts release is the one that changes minds, because it runs on hardware a firm may already own and summarises well enough to justify doing the retrieval work properly. The larger releases are visibly better at hard reasoning and materially harder to serve, so the step up is a procurement decision rather than a configuration change. And the multilingual coverage — German, French, Italian, Japanese and more — matters more to firms with European or Japanese clients than the charts suggest, because it makes fewer register errors on translated and codeswitched documents than English-focused models of similar size. Measure serving cost on your own throughput before you commit to the flagship.

evidence and abstention

Abstention is good when retrieval returns nothing and middling on near-misses: it will bridge adjacent but non-answering material rather than flag the gap. Quotation fidelity is strong when you ask for verbatim quotes with references; because its narration runs into its output, instruct it to place quoted text on its own line with a document and page reference. The prompting that lifts it most is a refusal rule with a defined string plus a requirement that every asserted fact carry a source line. Its reasoning run-on means you should also instruct it to answer before it explains, so a reader sees the conclusion first.

What we would use it for

  • Summarising long judgments, bundles and disclosure sets into house-format headnotes
  • Chronology and cross-reference work across a large document population
  • Multi-clause extraction and conflict flagging in a single pass
  • Multilingual matters — German, French, Italian and Japanese documents alongside English
  • Analytical second opinions where you want to see the route to the answer

What to watch

  • Publisher-specific licence terms — open in substance but publisher-defined, and not a permissive licence
  • Reasoning written in unmarked prose, which you cannot strip with a parser
  • Fluency standing in for flagged uncertainty
  • A very large advertised context that shrinks under real retrieval
  • Hardware-hungry flagship releases — serving cost rises faster than capability

What it costs to run

Serve one of the 30B-A3B releases — Nemotron 3.5 Lightning or Nemotron 3 Nano — rather than the flagship: the publisher validates them on a single 80GB H100-class card, at 256K context in BF16 and 1M in the NVFP4 build. One such card supports a firm of 10–50 fee-earners at a handful of concurrent requests if you accept second-scale first-token latency on long documents, and the 4B release on a 24GB workstation is the better answer for a single team doing classification and extraction. We would serve the NVFP4 checkpoint of Nemotron 3.5 Lightning on one 80GB card and keep the BF16 weights for evaluation and requantisation checks.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$1.58 – $5.24$1,150 – $3,820Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$1.58 – $5.24$330 – $1,100The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$3.30 – $6.90$200 – $415Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~300 GB)$45Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 80GB class — A100 80GB / H100. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

Rent first for this family: a 30B-class mixture of experts is small enough to test on one card and large enough that buying for it is a real decision, and rental lets you find out whether long-document summarisation on your own matters justifies the hardware. Buying into the 80GB server class is indicatively $25,000–60,000 once chassis, storage, networking and the operational care a client-confidential deployment needs are included, and these figures are indicative only — obtain quotes for your own specification. The flagship releases sit outside that band altogether, because they need multi-GPU nodes.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Summarisation: temperature 0.1–0.2, top_p 0.9, with a hard rule that quotations sit on their own line with a reference. Extraction: temperature 0 with a strict schema. Conflict analysis: 0.2. Drafting: 0.3–0.4. Where a release exposes a reasoning-effort or thinking setting, record it alongside the model version — a low-effort run is not comparable to a high-effort run, and we have seen the difference change extraction quality on dense schedules. Pin the exact release and its precision; the vendor ships several.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
One workstation, small mixture-of-experts release at low precisionSummarisation, extraction and triage for a single teamActive parameters, not total, drive your speed; a 4-bit build of a small MoE is the cheapest way to test the family
Two to four GPUs, mid-size release in vendor-quantised precisionFirm-wide summarisation and long-context reviewStart from the vendor's own quantised weights — requantising an already-quantised release has cost us more quality than memory it saved
Larger deployment, flagship releaseThe heaviest reasoning and the longest documentsJustify it against the mid-size model on your own task set; the flagship's summarisation advantage is real but narrower than its paper lead

Who it suits

Good fit

Firms whose bottleneck is reading volume — disclosure, bundles, long judgments — and who need multilingual coverage beyond English.

Poor fit

Firms unwilling to read a bespoke licence, or wanting a commoditised permissive model with no conditions attached.

Review history

DateChange
Sep 2026First entry.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.