Default reasoning effort to high (was max) — fixes runaway thinking, 500s, empty content
GLM-5.2 template forces Reasoning Effort: Max unless reasoning_effort==high, so hard/long requests burned the whole token budget thinking (187s, 0 content) -> proxy 500s / timeouts. Ship chat_template.think.jinja that defaults effort to high (client can still override): faster, content reliably produced, reasoning + tool calling intact.
0
0xSero committed
579d78b14112198f57aa141332328a79ab02c5a5
Parent: af3af3d