Skip to content
r/LocalLLaMA

How small can the orchestration model in an agent be? Separating it from code generation

How small can the orchestration model in an agent be? (separating it from code-gen — that obviously wants a big model)

HomoAgens1 runs a local ReAct orchestration loop on Qwen3.6-35B-A3B, with about 3B active parameters, a 12GB GPU, 30 expert offload, and 40 tokens/s prompt generation; smaller dense models fail first on tool-call discipline, inventing arguments or repeating bad calls, while reasoning is not identified as the first break point.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit experiment rather than a formal release. The VRAM, speed, and failure-mode details put it at the 72 featured threshold.

Read the original ↗Export Markdown