Legal AI Procurement Has a Testing Problem
Legal AI products are becoming more capable, more numerous and more difficult to distinguish.
Procurement methodology needs to catch up. A typical procurement might shortlist several products, assess functionality, security, integration and price, and invite each vendor to demonstrate its technology. The problem is obvious: each vendor is effectively sitting a different exam.
The demonstrations are different. The examples are different. The claims are different. And much of the evidence about performance has been generated by the company selling the product.
For conventional software, this may matter less. A system either integrates with an API, supports a particular workflow or generates a required document. Generative AI is different. Performance is probabilistic and highly dependent on the task, the information provided and the circumstances in which the system is used.
That makes independent comparative testing increasingly important.
Product claims are not benchmarks
There is already evidence that apparently comparable legal AI products can perform quite differently when subjected to the same independent test.
Researchers at Stanford constructed a preregistered dataset of more than 200 legal research questions and used it to test leading proprietary AI legal research tools.
The differences were substantial. Lexis+ AI answered 65% of the researchers’ questions accurately. Westlaw AI-Assisted Research answered 42% accurately. The researchers also found hallucination rates above 17% across the products tested, with Westlaw AI-Assisted Research hallucinating more frequently. The peer-reviewed study concluded that some provider claims about reliability had been overstated.
The results should not be read as a contemporary league table of those products. Products change, and the benchmark tested particular legal research tasks at a particular point in time.
That is precisely the point.
A procurement team needs to know how the products it is considering perform now, on the work for which they are being procured.
Give the products the same exam
A better approach is conceptually simple. Create a representative set of legal matters that shortlisted products have not previously seen. Establish the expected results and scoring criteria in advance. Give each product the same material and the same tasks.
Same files. Same questions. Same scoring.
The result might look something like this:
Product A: 91.4%
Product B: 83.7%
Product C: 72.1%
The aggregate score is only the beginning. Testing can identify differences in factual accuracy, issue identification, extraction, hallucination, identification of missing information, source attribution, consistency, speed and performance as matters become more complex. The result may not even be that Product A is simply “better” than Product B. One product might perform exceptionally well on routine extraction while another is materially better when information is incomplete or contradictory.
That is useful procurement evidence.
Synthetic Test Files putting tools to the test
Synthetic data makes realistic testing possible
Legal technology creates a particular testing problem: the most representative test material is often the material that organisations should be most reluctant to provide to multiple prospective vendors.
Real legal files contain confidential information, personal information and, frequently, highly sensitive information.
Synthetic legal data offers an alternative. A synthetic test corpus can reproduce the characteristics of real legal work without reproducing an actual person’s matter. Files can contain pleadings, correspondence, contracts, witness accounts, financial information, chronologies and evidence. They can deliberately contain missing documents, irrelevant material, conflicting accounts and factual ambiguity.
Crucially, the expected answer can be established when the synthetic matter is created.
This isn’t a particularly radical proposition. UK Government AI procurement guidance expressly contemplates synthetic data as a privacy-preserving technique and recommends testing models under a range of conditions, defining acceptable performance and designing for reproducibility. It also contemplates technology challenges in which vendors compete against each other. The Australian Government’s AI Technical Standard similarly calls for defined test criteria, separation between formal test data and development data, a degree of independence between testers and developers, and testing using real-world and synthetic data.
NIST goes further in articulating the underlying evaluation principle: AI accuracy should be measured using clearly defined, realistic test sets representative of the conditions in which the system is expected to be used.
UK Government Guidelines for AI Procurement
NIST AI Risk Management Framework
Australian Government AI Technical Standard
For justice technology, test the difficult cases
The case for representative testing is particularly strong where technology is being deployed in courts, tribunals and other justice institutions. A system may perform extremely well when presented with complete documents, clearly expressed facts and legally literate users.
That is not always the environment in which justice institutions operate.
A meaningful court technology benchmark should include missing evidence, procedural errors, conflicting accounts, ambiguous facts, irrelevant material and matters in which the appropriate response is to identify uncertainty or escalate the matter to a human.
There is an institutional dimension to this.
Courts exercise public power. Their authority depends not simply upon producing efficient outcomes, but upon institutional legitimacy: people must have reason to trust that processes are fair, consistent and capable of dealing appropriately with the case before them.
Introducing AI into those processes without understanding how it behaves across realistic and difficult cases risks more than a disappointing technology investment. It can affect confidence in the institution using it.
For access-to-justice technology, test the people most likely to be failed
The same principle has an even more immediate application to access-to-justice technology. A legal information or triage product may perform impressively for a digitally confident user who understands their problem, answers questions accurately and has all their documents.
That tells us relatively little about how it will perform for many of the people who most need legal assistance.
Testing should include users represented through synthetic scenarios involving limited legal or digital literacy, incomplete or inconsistent information, language or accessibility barriers, multiple intersecting legal and non-legal problems, misunderstanding of legal processes, vulnerability and circumstances requiring human intervention.
These should not be peripheral edge cases added after a product has been selected.
For an access-to-justice product, performance at the edges may be one of the most important measures of performance. An average accuracy score can conceal the fact that a system performs least reliably for the people for whom an incorrect answer carries the greatest consequences.
JusticeDataStudio.com puts LegalTech and JusticeTech to the test.
Benchmark before you buy
Comparative benchmarking is not a substitute for privacy, security, regulatory, accessibility or broader technology due diligence.
It answers a narrower question:
Given the same legal work, under the same conditions, how do these products actually compare?
As legal AI becomes embedded in professional practice and justice systems, that seems an increasingly basic question for procurement teams to answer before, rather than after, purchasing a product. Justice Data Studio develops synthetic legal test corpora and independent comparative benchmarks for legal, court and access-to-justice technology.
More on comparative benchmarking
Don’t let the vendor write the exam.