AI Tools Index / Open models / Phi
Microsoft · reviewed Sep 2026 · assessed from published documentation and our own evaluation tasks

Phi

The Phi line of small dense models, including compact and mini releases

The cleanest licence in the index wrapped around the weakest legal behaviour we have assessed: excellent for narrow, verifiable tasks and short documents, unreliable the moment a matter file needs judgement.

Our verdict

Tier C — pilot only, and only for narrow, verifiable work. The MIT licence is the best in this index and the deployment realities are trivial, but the model's legal behaviour is the weakest we have assessed: it drops parts of instructions, drifts from the documents as context grows, asserts confidently on points of English law, and almost never declines. Do not put client data through it until you have evaluated it on your own tasks and satisfied yourself that the failure mode is visible to your reviewer.

Specifications, as published

PublisherMicrosoft
FamilyThe Phi line of small dense models, including compact and mini releases
ParametersSmall dense models, sized to run on a laptop or a single modest GPU rather than a server
ContextShort by the standards of this index, and effective context is shorter still; retrieval-grounded performance falls away well before the advertised window is reached
LicenceMIT for the principal releases — the most permissive position in this index, with no acceptable-use rider layered on top
WeightsDownloadable, with extensive quantised and low-memory community builds
ReleaseSequential small-model generations, each superseding the last quickly
Licence postureMIT for the principal releases — the most permissive position in this index, with no acceptable-use rider layered on top

Specifications are as published by the publisher at the review date and change frequently. Confirm them in your own evaluation before you procure.

Releases and variants

ReleaseSizeContextServing footprintWhat it is for
Phi-4-mini-instruct3.8B dense128K tokens~3GB at 4-bit, ~5–6GB at 8-bitThe general-purpose small release and the only Phi we would put in production: it is a competent classifier and tagger of short documents and a UK firm of 10–50 fee-earners can run it on one modest GPU. Capable for its size, and not a substitute for a large model on hard legal reasoning.
Phi-4-mini-reasoning3.8B dense128K tokens~3GB at 4-bit, ~5–6GB at 8-bitThe reasoning-tuned build of the same 3.8B base, with visible working that helps on narrow step-by-step tasks such as date arithmetic and ordering a short chronology. Realistically runnable by a firm of 10–50 fee-earners, but the family's habit of answering when the documents are silent is unchanged.
Phi-4-mini-flash-reasoning3.8B dense, hybrid SambaY architecture with differential attention64K tokens~3GB at 4-bit, ~5–6GB at 8-bitTuned for low latency on short reasoning calls, which makes it the cheapest option for tagging a large volume of short documents. The 64K window is the shortest in the family, so it is realistically runnable for bulk classification and not usable for bundle work.
Phi-4-multimodal-instruct5.6B128K tokens~4–5GB at 4-bit, ~8GB at 8-bit including the vision and speech encodersThe one release in this family with a distinct legal use case — reading scanned and photographed documents — and the smallest credible private deployment for that task, well within a 10–50 fee-earner firm's hardware. Every extracted figure needs checking against the image, because the family's quotation fidelity is weak.
Phi-414B dense16K tokens~9GB at 4-bit, ~16GB at 8-bitThe best-known release in the family and still its strongest reasoning model per pound of memory. The 16K window is what rules it out for a real matter file rather than the hardware: a firm of 10–50 fee-earners can run it easily, but chunking a bundle into 16K prompts invites exactly the drift from the material that this family already shows.
Phi-4-reasoning14B dense32K tokens~9GB at 4-bit, ~16GB at 8-bitA reasoning-tuned 14B with twice the window of Phi-4, so a single medium-length document or a small set of short ones will fit. Useful for structure and arithmetic inside one document; not a model to put in front of a hard question of English law, and still comfortably runnable by a mid-sized firm.
Phi-4-reasoning-vision-15B15B16,384 tokens~10GB at 4-bit, ~17GB at 8-bitThe newest release in the family and its most capable vision model, intended for document-image reasoning rather than for general drafting. Capable for its size on extraction from scans, but the short window and weak abstention keep it a supervised tool; a firm of 10–50 fee-earners can run it on one card without difficulty.
Phi-3.5-MoE-instruct41.9B total / 6.6B active (MoE — 16 experts with 2 active per token)128K tokens~22GB at 4-bit, ~42GB at 8-bitThe only mixture-of-experts release in this family and its largest by total parameters, carrying more stored knowledge than the small dense models — which is precisely where Phi is weakest. Still servable on one workstation card, so a mid-sized firm can realistically run it, but Microsoft has not refreshed it since 2024 and the family's abstention problem is the binding constraint, not the footprint.

Sizes, context windows and licences are as published by the publisher at the review date. The variant you pick matters more than the family name: a small dense release that fits one workstation and a large mixture-of-experts release that needs a multi-GPU server are not the same product, whatever the marketing says.

How it behaves on legal work

Phi is the clearest illustration in this index of why licence freedom and legal usefulness are separate questions. On paper it is ideal for a cautious firm: MIT-licensed, small enough to run on a laptop, cheap enough to deploy without a business case. In our evaluation tasks it performs well on narrow, verifiable work and poorly on anything that requires holding a whole matter in view. Set it to something tightly defined with an obvious right answer — classify this clause, pull the party names and dates from this one-page order, decide whether this letter is a complaint or a request for information, tag these 200 short documents into one of four categories — and it is quick, stable and cheap, and the output is checkable in seconds. Ask more of it and the picture changes. Multi-part instructions are the first casualty: give it five requirements and it will satisfy three, drop one without comment and quietly reinterpret the fifth, and it does this without any of the hedging a larger model would use to signal uncertainty. Long documents are the second: as the prompt grows, the model increasingly answers from the instruction rather than the material, so a summary of a long bundle drifts toward what a bundle of that kind usually says. That is a serious failure in legal work, and it is precisely the failure a fee-earner is least likely to catch, because the output is fluent and the drift is invisible. Formatting discipline is inconsistent. It will produce valid JSON when the schema is simple and will break the schema when the schema is nested; it will hold a heading structure and then append an unrequested commentary block; it is prone to an earnest, slightly textbook register that reads as student work in client-facing correspondence. Over-assertion is high. Ask it about a case, a section or a statutory provision and it will answer with the same confidence whether it knows the area or not, and in our tasks it produced confident statements about points of English procedure that were simply wrong. Language coverage is predominantly English, with limited reliable capability elsewhere, so it is not a candidate for cross-border material — a limitation that removes it from a large share of the work firms actually buy open models for, since incoming correspondence, overseas filings and cross-border instructions are common even in a domestic practice. The reasoning behaviour is best described as procedural rather than legal: it works through arithmetic, logic and structured puzzles competently for its size, and that competence does not transfer to questions of construction, precedence or commercial judgement. It will, however, follow the surface form of legal reasoning convincingly, producing numbered issues, sub-issues and conclusions in the shape a lawyer expects, which makes its weak answers harder to spot than a plainly garbled response would be. Asked about an area it does not know, it fills the shape with plausible content rather than leaving it empty. What a fee-earner notices in the first week is that Phi is genuinely pleasant to work with on small things and dangerous on big ones, and that the boundary between the two is not obvious until it is crossed. Our recommendation follows from that: use it for narrow, schema-bound, single-document tasks where a human can verify the output at a glance, never as the model that reads the file, and keep it out of any workflow where a plausible but wrong answer would survive review.

evidence and abstention

Abstention is the weakest in this index. Phi rarely says the documents do not answer the question; it produces an answer in the shape the question implied, and it does so without the hedging that would prompt a reviewer to check. Quotation is also approximate: it reproduces the gist of a passage with altered wording, and it will not reliably admit that it has paraphrased. An explicit verbatim-only instruction improves fidelity modestly, and a required empty-answer token when the corpus is silent improves abstention more, but neither brings it to the standard of the frontier-class entries. Never rely on a Phi quotation without opening the source.

What we would use it for

  • Classification and tagging of short documents into a small, fixed set of categories
  • Extraction of names, dates and amounts from single-page documents for a human to validate
  • Internal triage labelling where the output is reviewed in bulk rather than relied on
  • Non-confidential drafting of internal notes and training material
  • A low-cost sandbox for staff to learn how these models behave on legal text

What to watch

  • Silent omission and reinterpretation of multi-part instructions
  • Answering from the prompt rather than the documents as context grows
  • Confident wrong statements on points of English law and procedure
  • Near-total absence of abstention when the materials do not answer the question
  • Predominantly English capability, with little reliable support for other languages

What it costs to run

One 24GB card covers this entire family at the sizes a firm of 10–50 fee-earners would use, and no release here needs a second GPU. We would serve Phi-4-mini-instruct at 8-bit for classification and single-document extraction, because that narrow, verifiable work is what this family can be trusted with; Phi-4 at 4-bit is the alternative if a fee-earner wants the stronger reasoning and accepts a 16K window. Latency and concurrency are both comfortable for asynchronous one-document-per-call work, so the ceiling on this deployment is the model's judgement rather than the hardware. Keep the evaluation set resident on the same volume, since re-scoring a pinned release is the only way to notice when a small quantised model has drifted.

BasisGPU hoursHourly (USD)Monthly (USD)When this is the right pattern
Always-on server (24/7)730 h$0.58 – $1.65$425 – $1,205Firm-wide access, no cold starts, predictable latency
Business hours (10 h × 21 days)210 h$0.58 – $1.65$125 – $345The realistic pattern for a firm of 10–50 fee-earners: power it up, use it, stop paying
Bursty / autoscaled endpoints60 h$1.03 – $1.95$60 – $115Occasional analysis and pilots; you pay only for the seconds the model is working
Storage — weights, index and evaluation sets (~120 GB)$20Billed whether the model is running or not — the quiet line on the invoice

Indicative GPU class: 24GB class — RTX 4090 / L4 / A5000. Every figure above includes a 50% buffer on the underlying cloud rates — for encrypted storage, egress, idle capacity between requests, cold starts, operational overhead, and the plain fact that these are estimates rather than quotes. Rates move weekly and vary by region, tier and commitment.

the comparison that decides it

This is the simplest buying case in the index, because the hardware is ordinary: a 24GB-class workstation is an indicative $3,000–6,000, and a firm keeping a model of this size resident all day is not paying for a data-centre part. Renting a 24GB-class GPU is still the better first move, since the pilot workload is bursty and the hardware decision can wait until the task set is settled. We would treat any Phi deployment as a pilot rather than an investment: the licence is MIT and the running cost is trivial, so the constraint is evaluation effort. All of these figures are indicative.

Two rules of thumb that hold across the models we have deployed: renting beats buying until a firm is using the model more than about half of every working day, and stopping the instance matters more than the hourly rate — an idle server, and an idle storage volume attached to it, are where private AI budgets quietly go.

Sampling and prompt settings

Extraction and classification: temperature 0, top_p 0.8, one task per call and one short document per call — do not batch long material into a single prompt. Drafting: 0.3, and expect to rewrite the register. Avoid summarisation of anything longer than a few pages: the failure mode is drift rather than obvious error, and sampling settings will not fix it. If you use the model at all, pin the exact release and its quantisation, because small models lose more to quantisation than large ones and the difference shows up first in instruction following.

Pin the exact model release in your evaluation record. Behaviour moves between point releases, and an evaluation that does not name a version cannot be reproduced.

Hardware and quantisation

Deployment profileWhat it fitsWhat to know
Laptop or single small workstation GPU (quantised)Classification, tagging and short-document extraction for one userRuns almost anywhere; the appeal is cost and licence freedom rather than capability
One modest GPU serverBulk classification of large volumes of short documentsThroughput is good and the licence imposes no obligations, which is the family's real advantage
Private tenancy of a re-hosted buildFirms that want the capability without managing hardwareSeldom worth it for a model this small; the value is in the MIT terms, which a host cannot improve on

Who it suits

Good fit

High-volume, low-judgement tasks on short documents — classification, tagging and single-field extraction — where every output is checked at a glance and licence certainty matters more than depth.

Poor fit

Anything requiring a whole-matter view, long-document summarisation, multi-part instructions, cross-border material, or evidential work where a plausible wrong answer could survive review.

Review history

DateChange
Sep 2026First entry.

Sources

Published under our rubric. Specifications are as published by the model publisher at the review date; licences and capabilities change without notice, so verify before you procure. Scores are editorial opinion formed from published documentation and our own evaluation tasks — not a benchmark result and not a vendor statement. No publisher pays for placement, sees a score before publication, or can have an entry withdrawn. Nothing here is legal advice; test any model on your own matters before you put client data through it.