SIGN IN SIGN UP

docs(benchmarks): re-measure impact accuracy at the current pins

The 0.71 average F1 was inherited from a 2026-05-25 capture whose fastapi rows
came from two 1-file commits that commit 841c364 later replaced as
unrepresentative. Re-ran impact_accuracy against the configured test commits.

Average F1 over the same 13 commits is 0.693 (was 0.714) and precision 0.546
(was 0.578). The whole movement is fastapi: 0.834 -> 0.697, now graded on its
21- and 23-file commits instead of the two 1-file ones. Every other repo is
within rounding of its previous value, which is the expected result and a good
check that nothing else shifted. Recall stays 1.000 and stays labelled a
circular graph-derived upper bound.

Co-change mode has now been captured and is NOT usable: every graded commit
returns predicted_files = 0, so its 0.000 F1 measures a broken harness rather
than the predictor. Documented as such; no co-change number is quoted.

Also records the raw 2026-08-02 CSVs alongside the 2026-05-25 ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNv8JqBb46stATZYUinQtn
T
Tirth Kanani committed
8257a5681695930c35558db30b4d36c05ee25450
Parent: 459e762