Skip to content
Hacker News front page

Linum's JiT-DDT trains text-to-image models 3.6× faster at 4× the pixels

Training Text-to-Image Models 3.6× Faster

Linum introduced JiT-DDT, a text-to-image training architecture that merges VAE and DiT into a single pixel-space model. It trains 3.6× faster in GPU-hours than their previous Linum v2 baseline while generating 512×512 images instead of 256×256. The design builds on the JiT paper's 32×32 token compression and adds an encoder-decoder to recover fine details. Code and weights are open under Apache 2.0, labeled as a research artifact, not a full model release.

Read the original ↗Export Markdown