SIGN IN SIGN UP

Online eval v10: F1=0.406 on 66 real-world PRs

Prompt improvements (v7):
- Added categories 15-17: code quality, documentation, missing functionality
- Expectation-driven review with hypothesis formation
- Three-question filter (nitpick, verify, good-catch)
- Expanded CSS/styling category with responsive design
- Verification-first rule: "did you grep or assume?"

Online eval pipeline fix:
- Include analyzed PRs in load query (assembled + analyzed)

Key finding: bug recall ceiling is 21% (17/80 bug GTs). Model finds
real bugs but different ones than humans fix. Style/refactor/docs
(31% of GT) are unreachable without diluting bug precision.
B
bhaktatejas922 committed
60274b2d7e4e2654f2057998fba45071d2ea113a
Parent: 1b3face