SIGN IN SIGN UP

Add Round 5: Skill Intelligence Index (V2) + honest README findings

Round 5 widens the benchmark from 3 arms to 15 conditions (9 public rival
skills + 5 prompt/reference arms + no-skill baseline), generates equal-length
answers, judges all 15 blind (18 independent-scoring passes), and prices every
condition in tokens (Fable 5 $10/$50 per M, tiktoken cl100k proxy).

Honest headline: a free one-line self-critique prompt leads judged reasoning
quality (8.28); this skill is #3 of 15 (7.84) in a top-cluster near-tie, is
#1 of 15 on objective rigor (falsifier presence double the next best), and is
the heaviest context-load in the field, so it is dominated on the
quality-vs-cost frontier. Published in full because a calibration skill that
hid its own benchmark would be a contradiction.

Ships: raw answers + judge JSON + rotation keymap + token telemetry, a
three-track dashboard (INDEX_EXPANDED.html), three editorial SVG banners, and
the as-run harness. README + CHANGELOG updated with the findings and the V2
work list (trim load, keep falsifiers, fold in a self-critique pass).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
J
johna2an committed
78409aa4c098663923450490af37a6578e1c8b50
Parent: 99c3bd5