Case Search — Legal Citation Verification (mellea-lrc)
- Role: Graduate Research Assistant, Georgia Tech CSSE — validation stage owner
- Focus: LLM Verification · Retrieval Grounding · Benchmark Design · Error Attribution
- Live Demo: mellea-lrc-visualizer.woodygoodenough.com
- Source Code: gt-csse/mellea-lrc · gt-csse/mellea-lrc-viz
- Benchmark: false-citation-bench on HuggingFace
Project Overview
mellea-lrc is an end-to-end system built at Georgia Tech’s Center for Scientific Software Engineering on IBM’s Mellea framework. It consumes raw legal filings, extracts every citation, and checks each one against CourtListener, the public court archive. I am responsible for the validation stage: deciding, citation by citation, whether the cited record actually supports the claim made from it.
The Validation Stage
Checking happens against an incomplete archive — many genuine filings were never uploaded to CourtListener. The validation stage is designed around that constraint:
- Absence is not falsity. The system never concludes that a citation is fabricated just because a lookup failed. What it established stays separate from what it could not determine.
- The groundedness gate. A substantive verdict is accepted only if the model quotes the cited page and that quote can be located again in the retrieved text. A judgment the model cannot ground is reported as a failure, not as a finding.
- Error attribution. Verdicts use an explicit vocabulary — FILING_DEFECT, RECORD_DEFECT, UNDETERMINED — so a downstream reader knows what kind of evidence sits behind each outcome.
Results
- Validation: 89.1% accuracy over 423 citations, zero false positives
- Citation extraction improved from 88.6% to 94.8% recall at 100% precision
The Benchmark
Alongside the system we published false-citation-bench: 26 real legal documents containing AI fabrications, with 594 annotated citation identifiers and a public protocol so anyone’s solution to the same problem can be evaluated the same way. Extraction and validation are scored independently.