Online eval v10: F1=0.406 on 66 real-world PRs
Prompt improvements (v7): - Added categories 15-17: code quality, documentation, missing functionality - Expectation-driven review with hypothesis formation - Three-question filter (nitpick, verify, good-catch) - Expanded CSS/styling category with responsive design - Verification-first rule: "did you grep or assume?" Online eval pipeline fix: - Include analyzed PRs in load query (assembled + analyzed) Key finding: bug recall ceiling is 21% (17/80 bug GTs). Model finds real bugs but different ones than humans fix. Style/refactor/docs (31% of GT) are unreachable without diluting bug precision.
B
bhaktatejas922 committed
60274b2d7e4e2654f2057998fba45071d2ea113a
Parent: 1b3face