Skip to content
Hacker News front page

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

This paper proposes an architecture that writes runtime data—user-supplied facts or corrections—directly into model weights instead of re-reading them from the prompt. A compact hypernetwork turns live data into a low-rank modulation of a shared base network, and a Bayesian belief over the latent code is updated online so the effective weights evolve across turns. The claimed benefits: frees context window, persists across turns, and may generalize better than in-context learning. The post only provides the abstract and framework; concrete experiments, model scale, and latency figures are not disclosed yet.

Why it matters: Fresh architecture idea: turn live user input into model weights instead of re-reading prompts. The hypernetwork + Bayesian online update gives it real mechanism, not just a concept. But it's a pure paper with no product path and no known lab behind it — capped at the featured...

Read the original ↗Export Markdown