Weekend with Apodex 4B and 35B mini: small search-agent models that don't hallucinate multi-hop answers
Spent the weekend on the Apodex 4b, plus a quick look at the 35b mini
The author ran Apodex 1.0's open models on a single 3090. The 4B-SFT was wired into a ReAct harness with a search tool for multi-hop questions where answers sit three links deep—it hallucinates far less than other 4B-class models. Apodex claims it beats every open 30B-class model on BrowseComp and BrowseComp-ZH; the author's handful of test questions back that up. The 35B mini has only ~3B active parameters per token but the full 35B weights on disk force heavy CPU offload, making it too slow for anything beyond one-off queries. No official gguf exists yet, so the author converted the 0.8B and 2B themselves and kept the 4B in vLLM. The design idea that caught their attention: the context that checks the answer is not the same context that produced it—a pattern a few groups are pushing, now showing up in models small enough for a single card.
Why it matters: A first-person experiment on a single 3090 with concrete BrowseComp comparisons and a specific claim about reduced hallucination. Kept at the lower end of featured because it's a single community post without a formal paper or cross-source confirmation, and the 35B mention is ...