Home / AI for intellectual property lawyers
SallyIP is a verification-first AI workspace for intellectual property work — patent research, drafting and prosecution, trademark clearance, and IP contract review — where every material conclusion links to retrieved evidence.
General legal AI produces fluent answers. IP practice needs something narrower: answers traceable to the exact passage that supports them, with the misses measured and published. SallyIP retrieves evidence before generating, verifies each quotation word-for-word, strips citations that point nowhere, and qualifies or refuses when evidence is missing.
Patent attorneys and agents, IP boutiques, in-house IP teams, and inventors preparing disclosures — teams that need reviewable output, not confident prose.
SallyIP retrieves evidence first and generates second. Hybrid lexical-plus-vector retrieval with relevance thresholds (zero results is a valid result); exact-quote verification of every cited passage; a citation-integrity guard that removes dangling source labels before display; entailment grading of each proposition-citation pair (entails, partial, context, contradicts, does not support); and an answer mode of VERIFIED, QUALIFIED, or RESEARCH REQUIRED. Citation integrity is not legal correctness: a real citation can still support a wrong conclusion, which is why entailment checks and practitioner review exist. Operating rule: no evidence, no assertion.
| Benchmark (sample, model, date) | Recorded result | Status |
|---|---|---|
| Grounding benchmark 100 grounding-100 / 2026-09-09 · n=95 · gemini-flash-lite-latest · 2026-09-09 | Authority recall: 100%; Zero dangling citations: 100%; Exact-quote verification: 72.5%; Unverified quotes (flagged, never silently passed): 22.5% | VERIFIED INTERNAL |
| Frozen v1.0 P0 full regression (100 questions) v1.0 frozen · n=100 · gemini-flash-lite-latest · 2026-09-09 | Authority recall: 98.0% (PASS at boundary, target ≥98%; 2 execution failures: s102b-08, s112d-02); Citation integrity of scored answers: 100%; Exact-quote verification: 95.8% (target ≥95%: PASS); Missing-quote rate: 4.2% (target <2%: FAIL); Citation entailment: 66.7% (target ≥95%: FAIL); Unsupported-proposition rate: 81.0% (target <2%: FAIL) | BLOCKED |
| Adversarial v1 (live) adversarial-v1 · n=12 · gemini-flash-lite-latest · 2026-09-10 | Accurate: 50.0%; Hallucinated: 0%; Incomplete (declined despite answerable evidence): 50.0% | VERIFIED INTERNAL |
| Stanford-style bench v1 (automated) stanford-bench v1 · n=24 · gemini-flash-lite-latest (runner default; actual value in DB run context, not extracted) · 2026-09-09 | Accurate: 58.3%; Hallucinated: 0%; Incomplete: 41.7% | VERIFIED INTERNAL |
| Ablation: guards vs base answer (simulation) ablation 2026-09-10 · n=8 · none in loop (pure-local simulation, 0 live model calls) · 2026-09-10 | Injected fabrications flagged by full pipeline: 100%; Injected fabrications caught by substance scoring alone: 25% | VERIFIED INTERNAL |
| Retrieval bench v1 (offline) v1-frozen · n=120 · n/a (offline engine test) · 2026-09-10 | Checks passed: 82.5%; Family-resolution (metric under revision): 0% — bench expectation contradicts its own withheld data | PENDING |
| Ablation-25: guards vs substance on 22 new synthetic fabrications ablation-25 / 2026-09-10 · n=28 · none in loop (pure-local simulation, 0 live model calls) · 2026-09-10 | New fabrications flagged by full pipeline: 100%; Combined fabrications flagged (26 fab): 100%; Good answers falsely flagged: 0% | VERIFIED INTERNAL |
Full tables, run references and the blocked-release disclosure: SallyIP benchmarks. Peer-reviewed context: General LLMs hallucinate on 58–88% of legal queries (Dahl et al., Stanford HAI 2024 (peer-reviewed)); Purpose-built legal RAG tools: 17–33% (Lexis+ AI ~17%, Westlaw AI-Assisted Research ~33%, CoCounsel-lineage ~17% with 60%+ refusals) (Magesh et al., Stanford/JELS 2024–2025 (peer-reviewed); tested product versions may differ).
SallyIP is a verification-first AI workspace for intellectual property work — patent research, drafting and prosecution, trademark clearance, and IP contract review. Every material conclusion links to retrieved evidence with exact-quote verification; where evidence is insufficient, it says so.
An approach where retrieval and verification gate every answer: evidence first, exact-quote checks, dangling-citation removal, entailment grading, and VERIFIED / QUALIFIED / RESEARCH REQUIRED modes — with miss rates published, including failures.
Yes. Peer-reviewed studies find general LLMs hallucinate on most legal queries tested, and even purpose-built tools miss a substantial share. SallyIP counters with mechanical citation-integrity and quote checks plus published miss rates — for example 22.5% of quotes unverified (27/120) on the grounding bench, flagged rather than silently passed.
Open SallyIPHow verification works