Skip to content
Hacker News front page

A visual tool that walks through GPT-2's Transformer architecture step by step

Transformers Explained Visually

Polo Club built an interactive page that uses GPT-2 (small) as a teaching model, breaking down embedding, attention, MLP, and output probabilities into clickable steps. You can type your own prompt, adjust temperature and top-k/top-p sampling, and watch how Q/K/V matrices and attention weights are computed per token. The model has 124M parameters, 12 blocks, a vocabulary of 50,257 tokens, and an embedding dimension of 768. It doesn't run the latest models, but the architecture principles are shared with GPT, Llama, and Gemini. The post doesn't mention inference latency or hardware requirements—it's purely a teaching demo.

Why it matters: Polo Club's interactive page dissects GPT-2 in detail, with tunable params and sampling steps — genuinely useful for anyone wanting to understand Transformer internals. Score capped here because it's a teaching tool, not industry news, and GPT-2 as a demo model isn't new.

Read the original ↗Export Markdown