SallyIP

Home / Why citation integrity is not legal correctness

Why citation integrity is not legal correctness

In one frozen 100-question run, every scored answer cited real retrieved sources — and most propositions still failed entailment. This note explains the gap every legal-AI buyer should ask about.

The case study (frozen P0 full, 100Q, gemini-flash-lite-latest, 2026-09-09)

Three distinct properties

  1. Existence: does the cited source exist and was it retrieved for this answer? (Mechanical; SallyIP enforces.)
  2. Fidelity: is the quotation exact? (Mechanical; 95.8% exact here, 4.2% missing — also a FAIL.)
  3. Entailment: does the source actually support the claim? (Hard; needs NLI plus practitioner judgment.)

Most legal-AI marketing collapses all three into "cited, therefore correct". The P0 run is published precisely because it shows the collapse failing — with the failing build blocked from release. Peer-reviewed context sharpens the point: purpose-built legal tools still miss 17–33% (Magesh et al.), so a vendor showing only existence-level metrics is showing the easiest third.

How buyers can use this

Ask any vendor for the three numbers separately on a stated sample: existence rate, exact-quote rate, entailment rate. If only one "accuracy" number exists, the other two are hiding inside it. Then ask for the practitioner-graded legal-correctness score — and treat its absence the way this page treats ours: PENDING, prominently, until humans grade.

All results · methodology