Skip to content
Trending storyPast story

A visual tool that walks through GPT-2's Transformer architecture step by step

1 report1 sourceupdated 8 days ago

What happened

Summary

Polo Club 做了一个教学用的交互页面,拿 GPT-2(小号版)当教具,把嵌入、注意力、MLP 和输出概率拆成可点击的步骤。你可以自己输入提示词,调温度和 top-k/top-p 采样,看每个 token 的 Q/K/V 矩阵和注意力权重是怎么算出来的。模型有 1.24 亿参数、12 层、词表 50257 个 token、嵌入维度 768。它跑的...

Coverage

Follow the reports to see the story from different sides.

Sep 22
  1. Hacker News front pagePick
    A visual tool that walks through GPT-2's Transformer architecture step by step

    Polo Club built an interactive page that uses GPT-2 (small) as a teaching model, breaking down embedding, attention, MLP, and output probabilities into clickable steps. You can type your own prompt, adjust temperature and top-k/top-p sampling, and watch how Q/K/V matrices and attention weights are computed per token. The model has 124M parameters, 12 blocks, a vocabulary of 50,257 tokens, and an embedding dimension of 768. It doesn't run the latest models, but the architecture principles are shared with GPT, Llama, and Gemini. The post doesn't mention inference latency or hardware requirements—it's purely a teaching demo.