SallyIP

Home / How we measure (and what we refuse to claim)

How we measure (and what we refuse to claim)

Five separate sections, never merged into one score. Mechanical grader verdicts are regression signals, not lawyer review and not publications.

The five sections

  1. Internal frozen — frozen datasets, verifiable internal runs (grounding-100, P0 full, regression-10q, retrieval, unit, stanford-24).
  2. External — peer-reviewed context and stratified subsets (ABIGAIL, IPBench, CUAD probes, LegalBench probe).
  3. Adversarial — live traps, poison pills, synthetic fabrications (adv-v1, ablation, ablation-25).
  4. Practitioner — human legal-correctness grading. Status: PENDING (0/25). No score shown until it exists.
  5. Blocked / pending — work that cannot run without missing inputs, listed explicitly instead of silently skipped.

Rules

Limitations

Mechanical grading cannot judge legal correctness. Free-tier model errors excluded items from denominators. Per-run Sally commits live in DB run contexts and are UNKNOWN in file reports unless stated. Pack coverage (US 7-passage, India/UK thin) limits recall breadth.

Commit and UNKNOWN discipline

Every file report records the baseline commit it ran against; cells the artifact does not state read UNKNOWN rather than inheriting values from other runs. Two reports may cover the same run (for example the grounding golden baseline appears in both the manifest and the cross-bench rollup) — the scoreboard cites both sources instead of merging them. This discipline is what makes the five sections combinable by readers but never pre-merged by us.

Results · latest.json