Make it just work: thinking OFF by default (direct answers), self-verifying launch.sh + chat.sh
- ship chat_template.nothink.jinja; default serves with thinking off so a normal request returns content directly (no empty-content reasoning-token trap) - reasoning parser is conditional: ENABLE_THINKING=1 -> native template + --reasoning-parser glm45 - launch.sh now waits for /health and smoke-tests before printing READY (fails loud otherwise) - add chat.sh one-liner client; tool calling verified in both modes - README rewritten around the one-command flow
0
0xSero committed
8fe3e26c6d03827d8e594532278725f546022dba
Parent: 3aca172