feat(evals): close coverage gaps — 6 authored objective sets + StrategyQA + run batches
Judge-free objective binary sets (balanced ~15/15, ~30 each) for skills that were previously unmeasured or proxy-only, consumable by run-routing-data.js: - reversibility (one-way vs two-way door) - first-principles (fundamental limit vs convention) - theory-of-constraints (is the proposed stage the binding bottleneck?) - second-order (does the action achieve its goal once 2nd-order effects resolve?) - map-territory (does observed behavior contradict the stated docs/spec?) - margin-of-safety (is the provisioning buffer adequate for the stated risk?) Also: point the strategyqa ingest source at ChilleD/StrategyQA (the prior voidful/ ref was unserved) with balanceScan — a real objective multi-hop set for second-order. Add run-confirm.sh (red-team/fermi at power) and run-rerun.sh (the 7 new objective evals). HF datasets gitignored per policy. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
T
Travis Boudreaux committed
0091d63baeef38bb896512b44129fea44e8e24fd
Parent: a93ab8a