RLM treats context as external data, not a prompt dump
Alex Zhang's Recursive Language Model (RLM) keeps long text outside the model window as external data; a root model queries it via code. With GPT-5-mini, RLM lifted OOLONG-Pairs F1 from 0.04% to 58.0% and BrowseComp-Plus accuracy from 0% to 91.3%. But BrowseComp-Plus has known data contamination, OOLONG-Pairs is author-designed, and baselines were tuned by the authors—discount those numbers. RLM only works at depth=1; depth=2 brings 28x latency and 100x token cost. It performs worse on math and science tasks, and Q95 cost can spike 10x above median. The repo has 5,230 stars; an independent reproduction pushed DeepSeek v3.2 on OOLONG from 0% to 42.1%.
Why it matters: Alex Zhang's RLM flips long-context from 'cram into window' to 'query as external data,' hitting 58.0% and 91.3% on two hard benchmarks at depth=1 with GPT-5-mini. The author's honesty about multi-layer recursion failing is a plus. Cap at 78 because it's still a model-specific...