SIGN IN SIGN UP

docs(examples): add Qwen3-32B SWE-bench Verified eval (Phase 3) (#3671)

* docs(examples): add Qwen3-32B SWE-bench Verified eval (Phase 3)

Follow-up to the Qwen3-32B data+recipe (#3594) and the shared eval harness (#3585). Docs only — reuses the model-agnostic ../eval scripts, no new scripts.

- Fill in the qwen3_32b Phase 3 section: goal, scaffold, run/grade commands (NOPARSER=1 + hermes fallback, thinking-on, TP=4xDP=2), Qwen3 serving notes, and the base-vs-SFT resolve table (base 18.2% -> SFT step499 24.4%, +6.2pt).

- Clarify the shared eval README NOPARSER bullet to cover the Qwen3 hermes case.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Abhishree <abhishreetm@gmail.com>

* docs(examples): tidy Qwen3-32B eval wording

- Drop the stale '(eval in a follow-up PR)' note and the gentler-LR trailing clause; minor shared-eval README wording.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Abhishree <abhishreetm@gmail.com>

---------

Signed-off-by: Abhishree <abhishreetm@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
A
Abhishree Thittenamane committed
fc3162d1e1d549982a3c43ecc75f8dd0e03ec444
Parent: 504967b
Committed by GitHub <noreply@github.com> on 8/26/2026, 12:29:02 AM