SallyIP

Home / Hallucination, measured — not marketed

Hallucination, measured — not marketed

Peer-reviewed research finds general LLMs hallucinate on most tested legal queries. SallyIP measures its own grounding mechanics on frozen benchmarks and publishes the misses alongside the wins.

What peer review found

SallyIP grounding benchmark 100 (run f8dfe146, gemini-flash-lite-latest, 2026-09-09)

Unverified quotes are not deleted quietly: they are de-quoted with telemetry preserved, so the failure stays measurable release after release instead of vanishing into a rephrased answer.

Mechanical verdicts are regression signals, not lawyer review. No smoke test populates a headline; every metric carries sample size and headline eligibility. Practitioner grading: PENDING (0/25) — no legal-correctness score exists for these answers, and none is shown until humans grade.

Data: latest.json · headline-metrics.csv · why integrity ≠ correctness

Can AI hallucinate legal citations?

Yes — studies show most tested legal queries fail on general models, and purpose-built tools still miss a substantial share. SallyIP counters with mechanical citation and quote checks plus published miss rates, and refuses or qualifies when evidence is missing.

All benchmarks