Add Round 5: Skill Intelligence Index (V2) + honest README findings
Round 5 widens the benchmark from 3 arms to 15 conditions (9 public rival skills + 5 prompt/reference arms + no-skill baseline), generates equal-length answers, judges all 15 blind (18 independent-scoring passes), and prices every condition in tokens (Fable 5 $10/$50 per M, tiktoken cl100k proxy). Honest headline: a free one-line self-critique prompt leads judged reasoning quality (8.28); this skill is #3 of 15 (7.84) in a top-cluster near-tie, is #1 of 15 on objective rigor (falsifier presence double the next best), and is the heaviest context-load in the field, so it is dominated on the quality-vs-cost frontier. Published in full because a calibration skill that hid its own benchmark would be a contradiction. Ships: raw answers + judge JSON + rotation keymap + token telemetry, a three-track dashboard (INDEX_EXPANDED.html), three editorial SVG banners, and the as-run harness. README + CHANGELOG updated with the findings and the V2 work list (trim load, keep falsifiers, fold in a self-critique pass). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
J
johna2an committed
78409aa4c098663923450490af37a6578e1c8b50
Parent: 99c3bd5