AI Tools Index / Open models / Apertus
Swiss AI Initiative / ETH Zürich and EPFL · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

Apertus

A fully open, transparently documented European model line, published in two sizes with open training data and published recipes

A narrow, genuinely excellent translator of legal material with a documented provenance story, Apache-licensed weights behind an acceptable-use gate, and weak summarisation and reasoning beyond its specialty.

Our verdict

Tier B — conditional, and the most interesting entry in this index for one specific job: multilingual legal work with a documented provenance story. It is a translator and a terminology engine, not a general assistant, and the acceptable-use terms add a compliance task that has to be assigned rather than assumed. Pilot it on foreign-language matters and judge it there, not on the general tasks it was never built for.

Specifications, as published

PublisherSwiss AI Initiative / ETH Zürich and EPFL
FamilyA fully open, transparently documented European model line, published in two sizes with open training data and published recipes
ParametersTwo sizes: one small enough for a workstation and one requiring a real server
ContextLong context as published, and the family is trained for very broad multilingual coverage — the developers publish support for well over a thousand languages
LicenceApache 2.0 weights, gated behind an acceptable-use policy that must be accepted before download and that includes an indemnity in favour of the developing institutions and obligations around personal data in model output
WeightsDownloadable after accepting the terms, with a published hash list of data-protection deletion requests to apply as an output filter; training data and recipes are published separately
ReleaseVersioned releases with a point update to the current line, each accompanied by a public technical report
Licence postureApache 2.0 weights, gated behind an acceptable-use policy that must be accepted before download and that includes an indemnity in favour of the developing institutions and obligations around personal data in model output

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
Apertus-v1.5-8B8B dense; multimodal input (images, and experimental speech) with an optional thinking mode262,144 tokens, a four-fold increase on Apertus 1.0~17GB at BF16, ~5GB at 4-bitThe recommended default for a firm of 10–50 fee-earners: it fits one 24GB workstation card, reads scanned documents as images, and translates legal material well — with the caveat that the weights are gated behind an acceptable-use policy that must be accepted, and that the policy carries personal-data obligations and an indemnity in favour of the developing institutions.
Apertus-v1.5-70B70B dense, same multimodal and thinking additions as the 8B release262,144 tokens~140GB at BF16 (2× 80GB or a 141GB-class card), ~40GB at 4-bitThe strongest translation model in this family and reported as best-in-field for translating Swiss court decisions, but at 140GB in BF16 it needs server hardware a mid-sized firm should rent rather than buy.
Apertus-8B-Instruct-25098B dense, instruct-tuned (Apertus 1.0)65,536 tokens as published~17GB at BF16, ~5GB at 4-bitThe previous-generation 8B instruct release, still a perfectly runnable option on one workstation card where the 1.5 multimodal additions are not needed, and published under Apache 2.0 with the acceptable-use policy attached to the download.
Apertus-70B-Instruct-250970B dense, instruct-tuned (Apertus 1.0)65,536 tokens as published~140GB at BF16, ~40GB at 4-bitThe 1.0 flagship, now superseded by the 1.5 release on context and multimodal input, and still a server-class deployment rather than something a 10–50 fee-earner firm should host for its own sake.
Apertus-8B-25098B dense, pretrained base checkpoint (not instruction-tuned)65,536 tokens as published~17GB at BF16The base checkpoint is the starting point for firms that want to continue training on their own precedents and matter files, which is the only route here to a model that sounds like your practice — and it will not follow instructions until you post-train it.
Apertus-70B-250970B dense, pretrained base checkpoint (not instruction-tuned)65,536 tokens as published~140GB at BF16The same continued-training route at 70B, which means multi-GPU hardware and a real training capability before any of it is useful to a mid-sized firm.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

Apertus is the entry in this index where the compliance story is as interesting as the model. The weights are Apache-licensed, which sounds like the simplest position available until you reach the gate: access requires accepting an acceptable-use policy that includes an indemnity in favour of the two Swiss federal institutes behind it, and obligations to treat personal data appearing in output as your own controller responsibility, helped by a published hash list of deletion requests that the developers advise applying as an output filter and refreshing periodically. For a UK firm already acting as a controller that is not a deal-breaker, but it is a task: someone has to own applying that filter, and a duty to refresh it is exactly the sort of thing that lapses quietly. Either build it into the pipeline or do not deploy. On legal tasks Apertus is specialised in an unusual and genuinely useful direction. The independent Swiss legal evaluation work we cite found it the best model in the field at translating court decisions, ahead of models many times its size, while placing it well down the overall ranking because summarisation and multiple-choice reasoning were weak. That profile maps onto a real need in a UK firm more often than it first appears: cross-border matters, documents arriving in a language nobody in the team reads fluently, and the constant temptation to run a machine translation that flattens legal register. Apertus holds register in legal text noticeably better than a general model of its size, and it does so across an implausible range of languages. Everything outside translation is more ordinary. Summarisation is the weak spot: ask it to condense a judgment into a headnote and you get something readable that has lost the part of the reasoning that mattered, which is the worst kind of summary failure because it is invisible on a quick read. Extraction into a schema is fine on clean documents and needs an explicit null convention on untidy ones. Drafting in English is competent and slightly stiff — it writes like a careful non-native speaker, which is precisely what it is, and that is a small problem in a letter about a delicate point. Reasoning traces are terse and, unlike some families, properly marked, so a pipeline can strip them cleanly. Instruction following is good in the way that matters for translation: it respects do-not-translate lists covering defined terms, party names and statutory references, keeps a glossary consistent across a long document, and does not silently tidy up a poorly drafted original. Fee-earners notice that last point in week one, because the failure they expect from machine translation — a rendering that is better written than the source and therefore misleading about what the source said — is largely absent here. What a fee-earner also notices, less happily, is that Apertus is not a general workhorse. It is a model you point at a specific problem — translation, multilingual triage, terminology consistency — and then put down. Teams that try to make it the firm's general assistant get mediocre drafting and poor summaries and conclude the model is weak, when the accurate conclusion is that it is narrow. The transparency is not marketing either: training data, pipelines and a technical report are published, the line was built with an explicit stance on respecting data-owner opt-outs, and that documentation is the artefact a client's compliance function will accept. For a firm whose clients or internal AI policy ask for documented provenance, that is worth more than a few points of general capability.

evidence and abstention

Quotation fidelity is good, and it is markedly better at preserving source wording in translation than in summary — a reminder that the two are separate skills and should be evaluated separately. Abstention is solid: pointed at material that does not answer the question, it says so without elaborate prompting. The prompting that lifts it most is a proper translation brief — a glossary, a defined do-not-translate list, the original in brackets for defined terms, and an instruction to flag anything it could not render confidently rather than smooth it over. Add a refusal rule for extraction and a requirement of a page reference per assertion.

What we would use it for

  • Translating foreign-language judgments, contracts and correspondence into English
  • Terminology consistency across a bundle — defined terms, party names, statutory references
  • Multilingual triage: working out which documents in a foreign set matter
  • Deployments where the client or the firm's own policy demands documented training data
  • Cross-checking a machine translation against the source on a register-sensitive passage
  • Non-confidential research where provenance and on-premises running are requirements

What to watch

  • Acceptable-use gate and indemnity — have the terms read by someone whose job that is
  • An output filter for personal-data deletion requests that must be applied and refreshed
  • Weak summarisation and multiple-choice reasoning for its size
  • Stiff, slightly non-native English in client-facing drafting
  • Narrow strength — deploying it as a general assistant will disappoint

What it costs to run

Serve the 8B release: at BF16 it sits on a single 24GB card, which is what makes this family affordable for a firm of 10–50 fee-earners, and translation and terminology work is where its legal value is. One card carries a small team at a few concurrent requests with comfortable latency on single documents, but a 262,144-token input in Apertus 1.5 consumes most of the available cache, so keep long-document work to one or two calls at a time. The 70B release is a different proposition altogether: better translation on hard material, but it needs an 80GB card at BF16 or a 48GB card in a 4-bit build before it is worth discussing.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$0.58 – $1.65$425 – $1,205Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$0.58 – $1.65$125 – $345The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$1.03 – $1.95$60 – $115Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~200 GB)$30Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 24GB class — RTX 4090 / L4 / A5000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

This is the family where buying is defensible, because the release we recommend runs on a workstation-class card: a 24GB-class machine is indicatively $3,000–6,000, a capital decision a firm of this size can take without a business case. Renting a 24GB instance is still the better first step while you establish whether a Swiss-trained multilingual model beats your current translation route on your own documents. If you later want the 70B release, rent the 80GB class rather than buying, because that is the size at which utilisation decides the answer — and all of these figures are indicative only.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Translation: temperature 0.1–0.2 with a glossary and a defined do-not-translate list in the prompt; consistency breaks down if you raise it. Extraction: temperature 0 with an explicit null. Summarisation: 0.1, and check omissions against the source. Drafting: 0.3. Pin the exact release and note whether you are running the base release or the point update — the two behave differently enough on translation that a comparison across them is not a comparison at all. Record the filter version you applied alongside the model version.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
Single workstation, smaller variant at low precisionTranslation and terminology work for one teamThe practical entry point, and where most of the family's legal value sits; the smaller variant is the one most firms should evaluate first
One GPU server, larger variantFirm-wide translation and multilingual triageQuantise for memory but re-check translation quality afterwards — that is the capability most sensitive to compression in our tasks
Private tenancy of a re-hosted buildTeams that want the model without operating hardwareThe acceptable-use obligations follow the deployment, so confirm who accepts the terms and who applies the output filter

Who it suits

Good fit

Firms with cross-border and foreign-language matters, and anyone whose clients ask to see the training data behind the model.

Poor fit

Summarisation-heavy work, English drafting that has to persuade, or teams that will not operationalise an output filter and a set of acceptable-use terms.

Review history

DateChange
Sep 2026First entry.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.