Ello breaks down the architecture of a real-time AI tutor that responds to a 5-year-old in under one second
Building a real-time AI tutor for 5-year-olds
Ello built an AI tutor for kids ages 4-9 that teaches math and reading. Standard agent tool loops were too slow—frontier models take 2–3 seconds for the first token, and kids tuned out. Their fix: the model streams multiple actions in a single response, and an interpreter executes each action as it arrives, so the child waits for only the first ~30 tokens. They also split the tutor into two agents—a real-time converser and an async planner that updates strategy during gaps when the child is thinking or speaking. Both agents share an append-only event log to avoid coordination delays. The post does not disclose end-to-end latency numbers or which models are used.
Why it matters: Ello's engineering post breaks down a real-time AI tutor into two concrete decisions: ditching the standard agent loop and splitting into a chat agent and a teaching agent. It has numbers (first 30 tokens, 1000ms) and a comparison (standard agent takes 2-3 seconds for the firs...