Multi-Model Routing After Entering Agent Sessions
多模型路由进入 Agent 会话之后
Multi-model routing saves cost and latency in single-turn Q&A, but falls apart inside multi-turn agent sessions. A real case from vLLM Semantic Router issue #1439: a user said 'looks good, commit it' during a Go refactoring task. The router saw four short words, judged the difficulty as low, and switched to a 0.5B model—which replied with pleasantries and dropped the task. The root cause is the router's narrow view: it can't see prior task state or tool-call progress. Four engineering hurdles make in-session model switching painful: incompatible history formats, Prompt Cache invalidation, non-transferable implicit reasoning tokens, and high glue cost for multimodal artifacts. Three approaches have emerged: Cursor and Claude Code isolate work into subagents with clean contexts; vLLM's SAAR lets the router track session state and lock the model during tool calls; most production agents simply stick to one best fixed model. vLLM's own baseline: a multi-model system must beat the best fixed model on the same budget and latency, or it's not worth the complexity.
Why it matters: An engineering analysis with a concrete failure case, not vague complaining. The vLLM issue #1439 example grounds the argument — useful for anyone building agent inference pipelines. Downside: it's a personal blog, not an official release, and the article body is truncated mid...