GLM
A permissively licensed mixture-of-experts family with strong bilingual capability and capable tool use, whose verbose, over-confident reasoning and heavy hardware footprint keep it a supervised tool rather than a system of record.
Tier B — conditional, with the cleanest licence position among the frontier-class entries. It is capable, genuinely strong in Chinese and English, and useful for the analysis and triage work where a visible chain of reasoning helps a lawyer decide. The conditions are that the reasoning is supervised rather than trusted, that citations are verified rather than assumed, that the hardware is real, and that the firm can answer provenance questions from its own clients without improvising.
Specifications, as published
| Publisher | Zhipu AI |
| Family | The GLM series of mixture-of-experts general models and the accompanying reasoning line |
| Parameters | Large mixture-of-experts models with a fraction of parameters active per token, alongside a reasoning-focused line |
| Context | Long context advertised across the current generation; effective context under retrieval is mid-pack, with the familiar degradation in the middle of long prompts |
| Licence | MIT for the recent principal releases, published alongside the model cards — permissive and OSI-recognised, subject to a model-specific acceptable-use policy |
| Weights | Downloadable, with quantised community builds available although quality loss on the larger mixture-of-experts releases is noticeable |
| Release | Rapid iteration, with general and reasoning lines refreshing on separate cycles |
| Licence posture | MIT for the recent principal releases, published alongside the model cards — permissive and OSI-recognised, subject to a model-specific acceptable-use policy |
Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.
Releases and variants
| Release | Size | Context | Serving footprint | What it is for |
|---|---|---|---|---|
| GLM-4.7-Flash | 31B total / 3B active (30B-A3B MoE — 4 of 64 experts active per token) | 200K tokens (202,752 as published in the release configuration) | ~17GB at 4-bit, ~32GB at 8-bit | The lightweight release in the current line and the only one a firm of 10–50 fee-earners can serve on a single 24GB-class card. MIT-licensed and genuinely runnable on-premise, with a 200K window that suits document sets; 3B active parameters is a small engine for hard legal reasoning, so treat it as extraction and summarisation capability rather than analysis. |
| GLM-4.5-Air | 106B total / 12B active (MoE) | 128K tokens | ~55GB at 4-bit, ~106GB at 8-bit | The mid-size release and the one we would actually serve: it fits a single 80GB-class card at 4-bit, which is the largest configuration a mid-sized firm should contemplate for a general assistant, and it is the smallest release in this family whose reasoning we would put in front of a fee-earner as a check on their own view. Older than the current flagship line, MIT-licensed, realistically runnable. |
| GLM-5.3-Flash | 320B total / 18B active (MoE, natively multimodal) | as published — the model card reports evaluation at up to 300K tokens | ~165GB at 4-bit, ~320GB at 8-bit | The current fast flagship and the first natively multimodal release in the GLM-5 series, with text, image and video input. At 4-bit it needs more memory than any single card holds, so it is a two-card server at minimum: capable, MIT-licensed, and past the point where a 20-partner firm should be buying hardware. |
| GLM-4.6 | 355B total / 32B active (MoE — 160 experts, 8 active per token) | 200K tokens | ~185GB at 4-bit, ~356GB at 8-bit | The generation that established the family's 200K working window and its tool-use behaviour, and still a common production build. MIT-licensed but multi-GPU only: roughly 185GB of weights at 4-bit means two 141GB-class cards. Capable, and impractical for a 20-partner firm. |
| GLM-4.7 | 358B total (MoE — 160 experts, 8 active per token; the predecessor GLM-4.6 publishes 355B total / 32B active) | 200K tokens | ~185GB at 4-bit, ~360GB at 8-bit | The last flagship of the 4.x line, with improved tool use and reasoning over GLM-4.6 and the same 200K window. MIT-licensed and a sensible target for a firm that already operates a two-card server; the footprint, not the licence, is what keeps it out of a 20-partner firm's reach. |
| GLM-5 | 744B total / 40B active (MoE — 256 experts, 8 active per token) | 200K tokens | ~380GB at 4-bit, ~744GB at 8-bit | The release that scaled the line from 355B total and 32B active to 744B and 40B, and the first to use sparse attention to contain deployment cost. MIT-licensed, and a four-card data-centre deployment at 4-bit — capable well beyond what a 20-partner firm can buy or would need to buy. |
| GLM-5.2 | 753B total (MoE — 256 experts, 8 active per token; the model card does not state an active-parameter total) | 1M tokens | ~385GB at 4-bit, ~753GB at 8-bit | The current MIT-licensed flagship and the first in the line with a solid 1M-token window, which is the point at which a whole transaction file could sit in one prompt. Roughly 385GB of weights at 4-bit puts it behind multiple 141GB-class cards: capable beyond question, and impractical for a 20-partner firm, which should read this row as context for the smaller ones rather than as a purchase option. |
| GLM-5.3 | 753B total (MoE — 256 experts, 8 active per token; the model card does not state an active-parameter total) | 1M tokens as published in the release configuration; the model card reports evaluation at up to 300K | ~385GB at 4-bit, ~753GB at 8-bit | The newest release in the line, and the licence is the field that changed: it is published under the bespoke GLM-5.3 Licence rather than MIT. The permissions are MIT-equivalent in substance, with one added condition that bites only on operators running the model as a service whose group revenue exceeds a stated threshold — a term most law firms will never approach, but a term nonetheless, and one to record. Its hardware position is unchanged from GLM-5.2, so it too is a multi-card data-centre deployment and impractical for a 20-partner firm. |
Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.
How it behaves on legal work
GLM's most useful trait in legal work is that it is a competent generalist under a permissive licence, and its least useful trait is that it rarely knows when to stop. In our evaluation tasks the mixture-of-experts releases handled drafting briefs well: given a structure, a register and a precedent, they produced documents that needed editing rather than rewriting, and they were better than most open weights at following several constraints at once — length, formality, defined terms and a prohibition on new facts. Extraction into a fixed schema is solid where the document is well formed and weaker where it is not, which puts it behind the strongest extraction models here but comfortably inside usable range for supervised work. Summarisation is thorough to the point of being overweight: it produces complete, well-organised summaries that a fee-earner will cut, and it has a habit of preserving the document's own muddle rather than resolving it, so an ambiguous letter stays ambiguous in the summary. Whether that is a virtue depends on the task; for chronology work it usually is. Formatting discipline is good, including tables and nested structures, though the reasoning line will sometimes place working inside the answer where it does not belong. The reasoning behaviour is the family's selling point and its main risk. Give it an analytical question — competing constructions of a clause, the strength of three arguments, what a chronology implies — and it works visibly, considers alternatives and often reaches a defensible position. The traces are long, they contain reversals, and the model states intermediate conclusions with far more confidence than they deserve, so a reviewer skimming the trace can mistake an abandoned hypothesis for a finding. Read the conclusion, then read the last third of the reasoning, rather than trusting the middle. Bilingual behaviour is a genuine strength: Chinese and English material is handled with less register loss than almost anything else in this index, cross-language summarisation works, and terminology stays consistent within a document. European languages are serviceable rather than strong, and the family is not the first choice for a firm whose cross-border work is German or Spanish. Mixed-language output is a real failure mode, though — ask a question in English over Chinese sources and the answer can switch mid-paragraph, which is easy to miss and awkward to explain to a client. Over-assertion is a tendency in the same direction as the rest of the family: it fills gaps in the retrieved material with confident background rather than flagging them. On the deployment side, the current releases are large and want multiple GPUs; quantised builds run and are noticeably weaker on exactly the reasoning tasks people want them for. That combination — a permissive licence, credible analysis, heavy hardware, and a publisher whose jurisdiction will prompt questions in client due diligence — shapes how a firm should use this family. Deploy it where the reasoning trace is an input to a human decision, keep it away from final drafting and from citation-heavy work, and have the provenance conversation with your own compliance function before a client raises it.
Quotation fidelity is average: with an explicit verbatim instruction and a reference requirement it reproduces passages closely, and without them it paraphrases while retaining approximate wording, which is the most dangerous form of paraphrase because it looks like a quotation. Abstention is better than most frontier-class open weights — asked to answer only from supplied passages, it will usually concede that the material is insufficient — but the reasoning line will supplement silence with general background and present it in the same register as sourced material. Requiring a source reference on every sentence is the single most effective control.
What we would use it for
- Analytical second opinions on competing constructions and argument strength, with the reasoning reviewed
- Chronology reasoning and inconsistency spotting across a document set
- Bilingual and cross-border matter support, including Chinese-language material
- Extraction into a fixed schema from well-formed documents, with human validation
- Internal research and training material on non-confidential sources
What to watch
- Reasoning traces that present abandoned hypotheses as findings
- Language switching mid-answer in mixed-language matters
- Confident background filling where retrieved documents are silent
- Heavy hardware requirements, with quantisation degrading the reasoning quality people want
- Provenance and jurisdiction questions in client due diligence that should be answered in advance
What it costs to run
One 80GB-class card serving GLM-4.5-Air at 4-bit is the configuration we would recommend for a firm of 10–50 fee-earners: roughly 55GB of weights leaves headroom for retrieved context and for several fee-earners working asynchronously. That is a deliberate trade, because the current GLM-5.2 and GLM-5.3 flagships are stronger models whose 4-bit footprint is roughly 385GB — four 141GB-class cards and a data-centre pattern no 20-partner firm should enter for a general assistant. Quantisation also costs more quality on this family's mixture-of-experts releases than on dense models, and the loss shows up first in reasoning, so fix the release and the precision together before you evaluate. If the workload is bulk classification rather than analysis, GLM-4.7-Flash on a 24GB card is the more economical starting point.
| Basis | GPU hours | Hourly (USD) | Monthly (USD) | When this is the right pattern |
|---|---|---|---|---|
| Always-on server (24/7) | 730 h | $1.58 – $5.24 | $1,150 – $3,820 | Firm-wide access, no cold starts, predictable latency |
| Business hours (10 h × 21 days) | 210 h | $1.58 – $5.24 | $330 – $1,100 | The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying |
| Bursty / autoscaled endpoints | 60 h | $3.30 – $6.90 | $200 – $415 | Occasional analysis and pilots; you pay only for the seconds the model is working |
| Storage — weights, index and evaluation sets (~400 GB) | — | — | $60 | Billed whether the model is running or not — the quiet line on the invoice |
Indicative GPU class: 80GB class — A100 80GB / H100. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.
Renting is clearly the right first move, because the step from a rentable single card to the hardware the current flagships want is large: an 80GB-class server is an indicative $25,000–60,000, and the 141GB-class machines those releases need sit above that band, for which we do not publish a figure here. Buying becomes sensible only when the model is in use for most of every working day, and for most firms we would expect the workload to justify one 80GB-class card rather than the multi-card configuration. A firm that wants GLM's bilingual strength at the lowest entry point should pilot GLM-4.7-Flash on a 24GB-class workstation first, because that is where the cost of being wrong is smallest. All of these figures are indicative.
Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.
Sampling and prompt settings
Analysis and triage: temperature 0–0.2 on the reasoning line, and let it finish — truncating traces produces confident fragments. Extraction: 0, top_p 0.8, schema-bound, one document class per call. Summarisation: 0.1–0.2, with an instruction to preserve ambiguity rather than resolve it. Drafting: 0.3–0.4 with precedent in the prompt. For bilingual matters, state the output language explicitly and require the whole answer in it. Pin the exact release across both the general and reasoning lines; they iterate separately and an evaluation that names the family rather than the release cannot be reproduced.
Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.
Hardware and quantisation
| Deployment profile | What it fits | What to know |
|---|---|---|
| Single high-memory workstation (heavily quantised) | Individual analysis tasks and evaluation, not shared throughput | Quantisation costs more here than on dense models, and the loss shows up first in reasoning quality |
| Multi-GPU server | Firm-wide triage, analysis and extraction support with retrieval | The realistic configuration; measure throughput against the firm's concurrent-user pattern before promising access |
| Private tenancy of a re-hosted build | Firms wanting the capability without operating the hardware | Moves the provenance question to the host rather than answering it — get the answers in writing, and record them |
Who it suits
Firms with cross-border or Chinese-language work wanting permissive-licence capability, and those who want an analytical second opinion whose reasoning they can inspect.
Citation-dependent work, client-facing final drafting, single-GPU deployments expecting frontier-class reasoning, or firms that cannot support multi-GPU hardware and the accompanying provenance conversation.
Review history
| Date | Change |
|---|---|
| Sep 2026 | First entry. |
Sources
- GLM model releases and licences (Zhipu AI / Z.ai)
- GLM model cards (Z.ai on Hugging Face)
- ISO/IEC 42001 — AI management systems
- Probative Co: private AI in law — the 2026 guide