Trending storyDeveloping
Baseten tests AI-generated inference engine, up to 90% faster
1 report1 sourceupdated 2 hours ago
What happened
AI digest
Baseten engineers used Claude Code and Fable 5 to generate an inference engine called VibeQwen, running Qwen-3.6-35B-A3B at NVFP4 precision on a single B200, according to an October 3 report. Against a tuned vLLM 0.25.1 baseline, VibeQwen's single-stream decode speed rose by up to 90%. The reported results cover only that model, hardware, precision and single-stream decode case; no other scenarios were tested.
Written by AI from the coverage · updated 1 hour ago
Coverage
Follow the reports to see the story from different sides.
Oct 3
- AI HOT · Tips & opinionsBaseten 实测 AI 生成推理引擎:Qwen-3.6-35B-A3B 单流解码比 vLLM 快最多 90%
Baseten 工程师用 Claude Code 和 Fable 5 生成 VibeQwen 推理引擎,在单张 B200 上运行 NVFP4 精度的 Qwen-3.6-35B-A3B,单流解码速度比调优后的 vLLM 0.25.1 最多提升 90%。
Heat over time
Not enough continuous observations to draw a trend yet.