SallyIP

Home / AI for intellectual property lawyers

AI for intellectual property lawyers

SallyIP is a verification-first AI workspace for intellectual property work — patent research, drafting and prosecution, trademark clearance, and IP contract review — where every material conclusion links to retrieved evidence.

What IP AI means here

General legal AI produces fluent answers. IP practice needs something narrower: answers traceable to the exact passage that supports them, with the misses measured and published. SallyIP retrieves evidence before generating, verifies each quotation word-for-word, strips citations that point nowhere, and qualifies or refuses when evidence is missing.

Scope, stated plainly. Available product capabilities: patent research, drafting and prosecution workflows; trademark clearance and intelligence; IP contract review; litigation evidence chronologies. Guidance and research support (evidence-grounded answers, educational resources, no dedicated production workflow): copyright questions, design-rights questions, trade-secret protection questions. Developing capability: trademark monitoring and portfolio intelligence. Nothing on this page is filing-ready output; practitioner review is required.

Who it is for

Patent attorneys and agents, IP boutiques, in-house IP teams, and inventors preparing disclosures — teams that need reviewable output, not confident prose.

Problems addressed

How verification works

SallyIP retrieves evidence first and generates second. Hybrid lexical-plus-vector retrieval with relevance thresholds (zero results is a valid result); exact-quote verification of every cited passage; a citation-integrity guard that removes dangling source labels before display; entailment grading of each proposition-citation pair (entails, partial, context, contradicts, does not support); and an answer mode of VERIFIED, QUALIFIED, or RESEARCH REQUIRED. Citation integrity is not legal correctness: a real citation can still support a wrong conclusion, which is why entailment checks and practitioner review exist. Operating rule: no evidence, no assertion.

Evidence, with failures shown

Benchmark (sample, model, date)Recorded resultStatus
Grounding benchmark 100
grounding-100 / 2026-09-09 · n=95 · gemini-flash-lite-latest · 2026-09-09
Authority recall: 100%; Zero dangling citations: 100%; Exact-quote verification: 72.5%; Unverified quotes (flagged, never silently passed): 22.5%VERIFIED INTERNAL
Frozen v1.0 P0 full regression (100 questions)
v1.0 frozen · n=100 · gemini-flash-lite-latest · 2026-09-09
Authority recall: 98.0% (PASS at boundary, target ≥98%; 2 execution failures: s102b-08, s112d-02); Citation integrity of scored answers: 100%; Exact-quote verification: 95.8% (target ≥95%: PASS); Missing-quote rate: 4.2% (target <2%: FAIL); Citation entailment: 66.7% (target ≥95%: FAIL); Unsupported-proposition rate: 81.0% (target <2%: FAIL)BLOCKED
Adversarial v1 (live)
adversarial-v1 · n=12 · gemini-flash-lite-latest · 2026-09-10
Accurate: 50.0%; Hallucinated: 0%; Incomplete (declined despite answerable evidence): 50.0%VERIFIED INTERNAL
Stanford-style bench v1 (automated)
stanford-bench v1 · n=24 · gemini-flash-lite-latest (runner default; actual value in DB run context, not extracted) · 2026-09-09
Accurate: 58.3%; Hallucinated: 0%; Incomplete: 41.7%VERIFIED INTERNAL
Ablation: guards vs base answer (simulation)
ablation 2026-09-10 · n=8 · none in loop (pure-local simulation, 0 live model calls) · 2026-09-10
Injected fabrications flagged by full pipeline: 100%; Injected fabrications caught by substance scoring alone: 25%VERIFIED INTERNAL
Retrieval bench v1 (offline)
v1-frozen · n=120 · n/a (offline engine test) · 2026-09-10
Checks passed: 82.5%; Family-resolution (metric under revision): 0% — bench expectation contradicts its own withheld dataPENDING
Ablation-25: guards vs substance on 22 new synthetic fabrications
ablation-25 / 2026-09-10 · n=28 · none in loop (pure-local simulation, 0 live model calls) · 2026-09-10
New fabrications flagged by full pipeline: 100%; Combined fabrications flagged (26 fab): 100%; Good answers falsely flagged: 0%VERIFIED INTERNAL

Full tables, run references and the blocked-release disclosure: SallyIP benchmarks. Peer-reviewed context: General LLMs hallucinate on 58–88% of legal queries (Dahl et al., Stanford HAI 2024 (peer-reviewed)); Purpose-built legal RAG tools: 17–33% (Lexis+ AI ~17%, Westlaw AI-Assisted Research ~33%, CoCounsel-lineage ~17% with 60%+ refusals) (Magesh et al., Stanford/JELS 2024–2025 (peer-reviewed); tested product versions may differ).

What is SallyIP?

SallyIP is a verification-first AI workspace for intellectual property work — patent research, drafting and prosecution, trademark clearance, and IP contract review. Every material conclusion links to retrieved evidence with exact-quote verification; where evidence is insufficient, it says so.

What is verification-first legal AI?

An approach where retrieval and verification gate every answer: evidence first, exact-quote checks, dangling-citation removal, entailment grading, and VERIFIED / QUALIFIED / RESEARCH REQUIRED modes — with miss rates published, including failures.

Can AI hallucinate legal citations?

Yes. Peer-reviewed studies find general LLMs hallucinate on most legal queries tested, and even purpose-built tools miss a substantial share. SallyIP counters with mechanical citation-integrity and quote checks plus published miss rates — for example 22.5% of quotes unverified (27/120) on the grounding bench, flagged rather than silently passed.

Open SallyIPHow verification works

Related

Patent AI · Trademark AI · Benchmarks · Glossary