feat(evals): judge rubrics + statistical power analysis
reasoning-quality + behavioral-pairwise judge rubrics (consumed by the runners) and power.js (Wilson CI, sign/McNemar, required-N — shows the 108-problem set is underpowered; ~194 needed for a 10pp aggregate). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
T
Travis Boudreaux committed
a055d7d26248a52186dfe88bb2484d506ed99a9c
Parent: df41397