As Hermes' auxiliary model it received requests with no max_tokens and a generation ran to 46k tokens (~70 min) after the client's 60 s timeout, starving the 35B. Cap thinking (reasoning-budget 2048) and total output (n-predict 4096) server-side; preset verified against a throwaway router. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>