Liquid AI cuts reasoning-model doom loops from 10.2% to 1.4% with Final Token Preference Optimization
Reducing Doom Loops with Final Token Preference Optimization
Liquid AI introduces Antidoom, a method that targets the exact first token of a repetitive loop in small reasoning models. Using Final Token Preference Optimization (FTPO), it trains the model to prefer coherent alternatives at that single position while leaving the rest of the distribution mostly untouched. On an early LFM2.5-2.6B checkpoint, the loop rate on hard math and coding prompts dropped from 10.2% to 1.4%, and eval scores improved as a result. The approach adapts Antislop and uses chosen/rejected single-token pairs, making it cheaper than RL. The post does not disclose training compute cost or latency impact.
Why it matters: Liquid AI proposes a lightweight fix for doom loops in small reasoning models: identify the first token of the loop and use preference optimization to swap it. The idea is clever, but it's only validated on an early 2.6B checkpoint—no cross-model or larger-scale comparisons ar...