Skip to content
Trending storyDeveloping

Magnitude ships self-optimizing inference engine

1 report1 sourceupdated 2 hours ago

What happened

AI digest

On October 1, 2026, Magnitude released a self-optimizing open-source inference engine for agents. It compiles and tunes kernels on the user's device. Magnitude says open models run up to 2x faster than on llama.cpp, with Metal decoding 92% faster and CUDA 19% faster. The post is on the Hacker News front page; this is still an early release.

Written by AI from the coverage · updated 2 hours ago

Coverage

Follow the reports to see the story from different sides.

Oct 1
  1. Hacker News front page
    Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

    Magnitude 发布开源推理引擎,会在用户设备上编译并调优 kernel,官方称开源模型运行速度最高比 llama.cpp 快 2 倍,其中 Metal 解码快 92%、CUDA 快 19%。

Heat over time

Not enough continuous observations to draw a trend yet.