Skip to content
Trending storyDeveloping

Baseten tests AI-generated inference engine, up to 90% faster

1 report1 sourceupdated 2 hours ago

What happened

AI digest

Baseten engineers used Claude Code and Fable 5 to generate an inference engine called VibeQwen, running Qwen-3.6-35B-A3B at NVFP4 precision on a single B200, according to an October 3 report. Against a tuned vLLM 0.25.1 baseline, VibeQwen's single-stream decode speed rose by up to 90%. The reported results cover only that model, hardware, precision and single-stream decode case; no other scenarios were tested.

Written by AI from the coverage · updated 1 hour ago

Coverage

Follow the reports to see the story from different sides.

Oct 3
  1. AI HOT · Tips & opinions
    Baseten 实测 AI 生成推理引擎:Qwen-3.6-35B-A3B 单流解码比 vLLM 快最多 90%

    Baseten 工程师用 Claude Code 和 Fable 5 生成 VibeQwen 推理引擎,在单张 B200 上运行 NVFP4 精度的 Qwen-3.6-35B-A3B,单流解码速度比调优后的 vLLM 0.25.1 最多提升 90%。

Heat over time

Not enough continuous observations to draw a trend yet.