Academic benchmark of seven AI legal research products reports citation error rates varying by an order of magnitude, and finds retrieval grounding reduces but does not eliminate fabrication
An academic team published a benchmark measuring citation accuracy in AI legal research tools, testing seven commercial products against a set of two thousand questions drawn from state and federal practice. The paper reports error rates varying by an order of magnitude between products and finds that retrieval-augmented systems reduce but do not eliminate fabricated citations. The authors released the question set and scoring code, and invited vendors to submit corrections before a second round of testing later this year.
Read the long piece on LexRegister
Key points
- Seven commercial products were tested against two thousand questions. Reuters
- Reported error rates vary by an order of magnitude across products. Financial TimesABA Journal
- The question set and scoring code were released publicly. Law360
- Vendors may submit corrections before a second testing round later this year. ReutersLaw360
Sources60 articles across 28 outlets
Artificial Lawyer
LawSites
ABA Journal
Above the Law
Financial Times
Law360
Legaltech News
Reuters
Bloomberg Law
Corporate Counsel
Courthouse News Service
Global Legal Post
JD Supra
Law.com
Legal Dive
Legal Futures
Legal IT Insider
National Law Review
The American Lawyer
The Guardian
The Lawyer
Thomson Reuters Institute
El País
Frankfurter Allgemeine Zeitung
Handelsblatt
International Association of Privacy Professionals Daily Dashboard
Le Monde
The Journal of the American Bar Association Legal Technology Resource Center