Apertus
A narrow, genuinely excellent translator of legal material with a documented provenance story, Apache-licensed weights behind an acceptable-use gate, and weak summarisation and reasoning beyond its specialty.
Tier B — conditional, and the most interesting entry in this index for one specific job: multilingual legal work with a documented provenance story. It is a translator and a terminology engine, not a general assistant, and the acceptable-use terms add a compliance task that has to be assigned rather than assumed. Pilot it on foreign-language matters and judge it there, not on the general tasks it was never built for.
Specifications, as published
| Publisher | Swiss AI Initiative / ETH Zürich and EPFL |
| Family | A fully open, transparently documented European model line, published in two sizes with open training data and published recipes |
| Parameters | Two sizes: one small enough for a workstation and one requiring a real server |
| Context | Long context as published, and the family is trained for very broad multilingual coverage — the developers publish support for well over a thousand languages |
| Licence | Apache 2.0 weights, gated behind an acceptable-use policy that must be accepted before download and that includes an indemnity in favour of the developing institutions and obligations around personal data in model output |
| Weights | Downloadable after accepting the terms, with a published hash list of data-protection deletion requests to apply as an output filter; training data and recipes are published separately |
| Release | Versioned releases with a point update to the current line, each accompanied by a public technical report |
| Licence posture | Apache 2.0 weights, gated behind an acceptable-use policy that must be accepted before download and that includes an indemnity in favour of the developing institutions and obligations around personal data in model output |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| Apertus-v1.5-8B | 8B dense; multimodal input (images, and experimental speech) with an optional thinking mode | 262,144 tokens, a four-fold increase on Apertus 1.0 | ~17GB at BF16, ~5GB at 4-bit | The recommended default for a firm of 10–50 fee-earners: it fits one 24GB workstation card, reads scanned documents as images, and translates legal material well — with the caveat that the weights are gated behind an acceptable-use policy that must be accepted, and that the policy carries personal-data obligations and an indemnity in favour of the developing institutions. |
| Apertus-v1.5-70B | 70B dense, same multimodal and thinking additions as the 8B release | 262,144 tokens | ~140GB at BF16 (2× 80GB or a 141GB-class card), ~40GB at 4-bit | The strongest translation model in this family and reported as best-in-field for translating Swiss court decisions, but at 140GB in BF16 it needs server hardware a mid-sized firm should rent rather than buy. |
| Apertus-8B-Instruct-2509 | 8B dense, instruct-tuned (Apertus 1.0) | 65,536 tokens as published | ~17GB at BF16, ~5GB at 4-bit | The previous-generation 8B instruct release, still a perfectly runnable option on one workstation card where the 1.5 multimodal additions are not needed, and published under Apache 2.0 with the acceptable-use policy attached to the download. |
| Apertus-70B-Instruct-2509 | 70B dense, instruct-tuned (Apertus 1.0) | 65,536 tokens as published | ~140GB at BF16, ~40GB at 4-bit | The 1.0 flagship, now superseded by the 1.5 release on context and multimodal input, and still a server-class deployment rather than something a 10–50 fee-earner firm should host for its own sake. |
| Apertus-8B-2509 | 8B dense, pretrained base checkpoint (not instruction-tuned) | 65,536 tokens as published | ~17GB at BF16 | The base checkpoint is the starting point for firms that want to continue training on their own precedents and matter files, which is the only route here to a model that sounds like your practice — and it will not follow instructions until you post-train it. |
| Apertus-70B-2509 | 70B dense, pretrained base checkpoint (not instruction-tuned) | 65,536 tokens as published | ~140GB at BF16 | The same continued-training route at 70B, which means multi-GPU hardware and a real training capability before any of it is useful to a mid-sized firm. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
Apertus is the entry in this index where the compliance story is as interesting as the model. The weights are Apache-licensed, which sounds like the simplest position available until you reach the gate: access requires accepting an acceptable-use policy that includes an indemnity in favour of the two Swiss federal institutes behind it, and obligations to treat personal data appearing in output as your own controller responsibility, helped by a published hash list of deletion requests that the developers advise applying as an output filter and refreshing periodically. For a UK firm already acting as a controller that is not a deal-breaker, but it is a task: someone has to own applying that filter, and a duty to refresh it is exactly the sort of thing that lapses quietly. Either build it into the pipeline or do not deploy. On legal tasks Apertus is specialised in an unusual and genuinely useful direction. The independent Swiss legal evaluation work we cite found it the best model in the field at translating court decisions, ahead of models many times its size, while placing it well down the overall ranking because summarisation and multiple-choice reasoning were weak. That profile maps onto a real need in a UK firm more often than it first appears: cross-border matters, documents arriving in a language nobody in the team reads fluently, and the constant temptation to run a machine translation that flattens legal register. Apertus holds register in legal text noticeably better than a general model of its size, and it does so across an implausible range of languages. Everything outside translation is more ordinary. Summarisation is the weak spot: ask it to condense a judgment into a headnote and you get something readable that has lost the part of the reasoning that mattered, which is the worst kind of summary failure because it is invisible on a quick read. Extraction into a schema is fine on clean documents and needs an explicit null convention on untidy ones. Drafting in English is competent and slightly stiff — it writes like a careful non-native speaker, which is precisely what it is, and that is a small problem in a letter about a delicate point. Reasoning traces are terse and, unlike some families, properly marked, so a pipeline can strip them cleanly. Instruction following is good in the way that matters for translation: it respects do-not-translate lists covering defined terms, party names and statutory references, keeps a glossary consistent across a long document, and does not silently tidy up a poorly drafted original. Fee-earners notice that last point in week one, because the failure they expect from machine translation — a rendering that is better written than the source and therefore misleading about what the source said — is largely absent here. What a fee-earner also notices, less happily, is that Apertus is not a general workhorse. It is a model you point at a specific problem — translation, multilingual triage, terminology consistency — and then put down. Teams that try to make it the firm's general assistant get mediocre drafting and poor summaries and conclude the model is weak, when the accurate conclusion is that it is narrow. The transparency is not marketing either: training data, pipelines and a technical report are published, the line was built with an explicit stance on respecting data-owner opt-outs, and that documentation is the artefact a client's compliance function will accept. For a firm whose clients or internal AI policy ask for documented provenance, that is worth more than a few points of general capability.
Quotation fidelity is good, and it is markedly better at preserving source wording in translation than in summary — a reminder that the two are separate skills and should be evaluated separately. Abstention is solid: pointed at material that does not answer the question, it says so without elaborate prompting. The prompting that lifts it most is a proper translation brief — a glossary, a defined do-not-translate list, the original in brackets for defined terms, and an instruction to flag anything it could not render confidently rather than smooth it over. Add a refusal rule for extraction and a requirement of a page reference per assertion.
What we would use it for
- Translating foreign-language judgments, contracts and correspondence into English
- Terminology consistency across a bundle — defined terms, party names, statutory references
- Multilingual triage: working out which documents in a foreign set matter
- Deployments where the client or the firm's own policy demands documented training data
- Cross-checking a machine translation against the source on a register-sensitive passage
- Non-confidential research where provenance and on-premises running are requirements
What to watch
- Acceptable-use gate and indemnity — have the terms read by someone whose job that is
- An output filter for personal-data deletion requests that must be applied and refreshed
- Weak summarisation and multiple-choice reasoning for its size
- Stiff, slightly non-native English in client-facing drafting
- Narrow strength — deploying it as a general assistant will disappoint
What it costs to run
Serve the 8B release: at BF16 it sits on a single 24GB card, which is what makes this family affordable for a firm of 10–50 fee-earners, and translation and terminology work is where its legal value is. One card carries a small team at a few concurrent requests with comfortable latency on single documents, but a 262,144-token input in Apertus 1.5 consumes most of the available cache, so keep long-document work to one or two calls at a time. The 70B release is a different proposition altogether: better translation on hard material, but it needs an 80GB card at BF16 or a 48GB card in a 4-bit build before it is worth discussing.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $0.58 – $1.65 | $425 – $1,205 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $0.58 – $1.65 | $125 – $345 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $1.03 – $1.95 | $60 – $115 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~200 GB) | — | — | $30 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 24GB class — RTX 4090 / L4 / A5000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
This is the family where buying is defensible, because the release we recommend runs on a workstation-class card: a 24GB-class machine is indicatively $3,000–6,000, a capital decision a firm of this size can take without a business case. Renting a 24GB instance is still the better first step while you establish whether a Swiss-trained multilingual model beats your current translation route on your own documents. If you later want the 70B release, rent the 80GB class rather than buying, because that is the size at which utilisation decides the answer — and all of these figures are indicative only.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Translation: temperature 0.1–0.2 with a glossary and a defined do-not-translate list in the prompt; consistency breaks down if you raise it. Extraction: temperature 0 with an explicit null. Summarisation: 0.1, and check omissions against the source. Drafting: 0.3. Pin the exact release and note whether you are running the base release or the point update — the two behave differently enough on translation that a comparison across them is not a comparison at all. Record the filter version you applied alongside the model version.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single workstation, smaller variant at low precision | Translation and terminology work for one team | The practical entry point, and where most of the family's legal value sits; the smaller variant is the one most firms should evaluate first |
| One GPU server, larger variant | Firm-wide translation and multilingual triage | Quantise for memory but re-check translation quality afterwards — that is the capability most sensitive to compression in our tasks |
| Private tenancy of a re-hosted build | Teams that want the model without operating hardware | The acceptable-use obligations follow the deployment, so confirm who accepts the terms and who applies the output filter |
Who it suits
Firms with cross-border and foreign-language matters, and anyone whose clients ask to see the training data behind the model.
Summarisation-heavy work, English drafting that has to persuade, or teams that will not operationalise an output filter and a set of acceptable-use terms.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. |
Sources
- Apertus (Swiss AI Initiative)
- Apertus models and documentation (Hugging Face)
- The state of the art in open-source AI for Swiss legal tasks
- ISO/IEC 42001 — AI management systems