Home / Tested against outside yardsticks
Beyond internal benches, SallyIP runs stratified subsets of external suites — with sample sizes prominent and full runs honestly marked BLOCKED where compute or budget is missing.
7/23 passed; 14 grounded refusals (no USPTO record access offline — correct fail-closed behavior); 2 drafts incomplete pending judge layers. 0 poison hits across all 23 outputs — fake MPEP sections, fake cases and fabricated cites all rejected or refused. Original scoring unmodified.
15/16 overall — PATENT 4/4, TRADEMARK 4/4, COPYRIGHT 3/4 (one regex brittleness kept unpatched as evidence), OTHER 4/4. Full 10,374-item run BLOCKED (compute + grading budget).
Entailment-only baseline 2/4 → 4/4 with additive ownership-signal layer (12 new tests green). Scope is a 4-item probe; real CUAD phrasing remains a generalization risk, stated openly.
26/26 injected fabrications detected, 0 false positives observed. Base-model substance scoring caught 1/4 on the original set and 0/22 on the expanded set — guards contribute the detection. Simulation with fixed passages; guard-behavior evidence, not a model-quality claim.
Train exemplars and synthetic representatives — headline-ineligible by design. Full 162-task run BLOCKED (paid key + harness + gradebook).
Run references: ABIGAIL/IPBench/CUAD subsets and ablations per benchmarks/external-validation-wave.md (2026-09-10, baseline d4c2870); adversarial-v1 live run 3368daf1; stanford-24 run fe26da44; grounding-100 run f8dfe146; ablation-25 report 2026-09-10.