Evidence Centered Benchmark Design
Stub
Methodology for designing benchmarks that explicitly tie task evidence to the underlying capability claims. Referenced from CLASP as a more rigorous alternative to LLM-as-a-judge for capability scoring. Needs full definition, methodology, and citation to canonical sources.
See also
- Agentic Vulnerability Discovery and End-to-End Harness Evaluation instantiate this design method on one capability, with differential crash validation and cumulative stage gating as the concrete evidence-to-claim links.
- Measuring Security-Agent Effectiveness is a practitioner instance: security ground truth is irreducibly noisy (double-digit analyst disagreement), so evaluation must score multi-dimensional reasoning rather than assume a transparent-label oracle.