feat(m5): powered primary batch — 7 gated skills evaluated, all NO-LIFT
Gate-checked and powered 7 approved skills (red-team, occams-razor, five-whys-plus, systems, archetypes, map-territory, kepner-tregoe) using solver claude-sonnet-4-6, CONC=4, isolation ON (EVAL_RUN=m5-primary). All are post-edit runs against M4-reworked skill versions on frozen calibrated splits. Results (all NO-LIFT, none passed ≥5pp + p<0.05 primary gate): - red-team: +1.4pp p=1.0 (security-adversarial decisive, n=70) - five-whys-plus: +0pp p=0.752 (SWE-bench, n=150) - systems: +1.3pp p=0.724 (SWE-bench, n=150) - occams-razor: +4.7pp p=0.096 (SWE-bench, n=150, directional) - map-territory: +2.7pp p=0.221 (exploratory: surface-mismatch) - kepner-tregoe: +2.7pp p=0.289 (exploratory: surface-mismatch) - archetypes: +0.9pp p=1.0 (systems pairwise, n=117, quarantine candidate) map-territory and kepner-tregoe tagged 'exploratory: surface-mismatch (powered on fault-localization, not native surface)' — their no-lift is NOT an honest kill. archetypes was an in-band quarantine candidate given a fair calibrated chance. Scorecard (JSON+MD) updated with full evidence tuples + provenance + verdict taxonomy. New run-binary-decision.js runner for systems-product-strategy-pairwise family. Assertions fulfilled: VAL-POWERED-001/004/005/006/009/010/012/013, VAL-DATASET-001/009, VAL-SKILLS-008.
T
Travis Boudreaux committed
b11f5538afcca5ffa91c725f83ccbd58b4b5e2ea
Parent: e0bb0d9