Academic benchmark of seven AI legal research products reports citation error rates varying by an order of magnitude, and finds retrieval grounding reduces but does not eliminate fabrication
An academic team published a benchmark measuring citation accuracy in AI legal research tools, testing seven commercial products against a set of two thousand questions drawn from state and federal practice. The paper reports error rates varying by an order of magnitude between products and finds that retrieval-augmented systems reduce but do not eliminate fabricated citations. The authors released the question set and scoring code, and invited vendors to submit corrections before a second round of testing later this year.
Read the long piece on LexRegister