How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
Controlling Reasoning Effort in LLMs
Sebastian Raschka explains how to train a single reasoning model to operate at multiple effort levels instead of always running at full throttle. He starts with GPT-5.6's five effort settings, then defines reasoning models as those producing intermediate step-by-step traces. Two levers exist: training-side RLVR and inference-side token budgets. The core recipe mixes reasoning traces of different lengths in the training data and conditions the model on budget tokens like <|low|> or <|high|>. In his experiments, he fine-tunes DeepSeek-R1-Distill-Qwen-32B with DPO on 1,040 preference pairs. On GSM8K, low-effort mode saves 40% tokens while dropping only 1.5% accuracy; high-effort mode spends 2.3× more tokens for a 2.1% gain. Raschka notes the approach is only validated on math benchmarks so far, and generalization to other domains is unknown. He closes with practical implications for cost and latency, plus the prospect of models self-selecting effort based on question difficulty.
Why it matters: Raschka explains how to train reasoning models to switch effort levels on demand. H and K are solid, but the piece is implementation-heavy so R doesn't fully land. Lands at 78 — clears featured but not 85.