Skip to content
r/LocalLLaMA

Local LLM Benchmark for Backend Generation via Function Calling: GLM vs Qwen vs DeepSeek

Local LLM Benchmark about Backend Generation by Function Calling (GLM vs Qwen vs DeepSeek)

AutoBe posted a controlled backend-generation benchmark and says qwen3.5-35b-a3b matches gpt-5.4 on DB/API design. One shopping-mall run uses 200–300M tokens, costing $1,000–$1,500 per model at GPT 5.5 pricing. The key caveat is n=4 projects and self-scoring harness bias.

Why it matters: HKR-H/K/R all pass, but Reddit sourcing, n=4 projects, and self-eval harness bias keep it at the low featured band. Concrete cost and test constraints carry the score.

Read the original ↗Export Markdown