AI Tools Index / How we score
the method · v0.9 · September 2026

How we score legal AI

One hundred points across five weighted criteria, four tiers, and an evidence standard that says what we looked at. The method is published so you can disagree with a specific number rather than the whole exercise.

01The five criteria

Each tool is scored out of 100. The weights are not equal because the risks are not equal: a tool that leaks client data is a different category of problem from a tool with an ugly interface.

25

Data protection & location

Where data rests, where it is processed, who the subprocessors are, what the retention defaults say, and whether your content trains anything. Full marks require published, specific answers — not a security page with badges. A tool that answers vaguely loses points here even when we suspect the answer is fine, because your DPIA cannot cite a suspicion.

25

Confidentiality & privilege

Whether client confidentiality and legal privilege survive normal use, and how the architecture helps or hurts. Tools that inherit your existing permissions — DMS-native ones especially — score highest because the confidentiality answer is structural rather than contractual. Tools that require pasting material into a shared third-party service score lowest.

20

Evidential transparency

Can you show your work? Citations that resolve to primary sources, logs of what the tool changed, outputs a supervisor can verify in minutes, and an audit trail that survives being asked about six months later. This is the criterion most tools fail, and the one regulators actually care about.

15

Professional usability

Does it fit how fee-earners already work, does it keep a human in the loop by design, and how much training does adoption really take? A tool nobody uses protects nobody — but high usability with a bad evidence trail is a trap, which is why this carries less weight than the three above.

15

Commercial honesty

Published pricing, renewals that behave, data you can export without a fight, and a value case that survives year two. Tools with transparent pricing outscore equally capable rivals that hide the number until a sales call — because opacity is a cost you pay later.

02The four tiers

TierScoreMeaning
A85–100Adopt with controls. We would put client data through it under a written policy, with named caveats. No tool has reached this tier yet.
B72–84Conditional. A good tool with real caveats. Read the write-up before a pilot, and put the caveats in your policy rather than in your optimism.
C58–71Watch. Useful today for non-confidential work. Not for client data until the governance story finishes.
Dbelow 58No client data. Not until the answers change. We will say so plainly and re-score when they do.
why the top grade is empty

An A is not a good review — it is a statement that we would stake our own reputation on the tool's confidentiality and evidence story for a firm like yours. Several tools here are close. Handing out an A for a polished demo is how an index becomes a brochure. We would rather the top tier be empty and be believed.

03The evidence standard

Every score states what it is based on. We use four kinds of evidence, strongest first:

  1. Our own testing. Hands-on use against tasks we specify, including attempts to make the tool misbehave. Rare in this preview build; it will be the backbone of v1.
  2. Vendor documentation and terms. DPAs, subprocessor lists, trust centres, security pages, published pricing. Quotable, and the basis of most entries today.
  3. Answers to our questionnaire. We send vendors the same 24 questions. Where they answer, we say so; where they decline, the score reflects the gap and the entry says "not disclosed".
  4. Firm reports. What firms tell us in confidence about how a tool behaved in production. Never attributed, and never the sole basis for a score — but it is often what prompts a re-test.

Where we could not verify something, the entry says so. "Not published" is a finding, not an absence of one.

04What we will not do

05Corrections, re-scores and how tools get in

Any vendor or firm can challenge a score. Send evidence; if it holds up, we correct the entry and publish the correction with its reasoning — including when the correction moves a tool up. If a vendor declines to engage at all, the entry keeps its documentation-based score and says the vendor did not respond.

Tools enter the index because firms use them, not because vendors ask. Every entry is re-reviewed at least annually, and sooner when the product changes materially.

06The limits, stated plainly

P
Published by Probative Co

An independent compliance and AI governance practice working with UK law firms on AML review, DPIA and AI policy, ISO/IEC 42001 readiness, and private AI deployment. The index exists because the same questions kept arriving from COLPs with nowhere independent to look.

Browse the index Submit a tool or correction