3Blue1Brown, created and run by Grant Sanderson, is one of the most respected math and AI education channels on YouTube, known for turning abstract technical concepts into clear visual intuition. This particular video started as a commission for the Computer History Museum, built for an exhibit on the history of chatbots, and Sanderson adapted it for his own audience after noticing that his longer, more technical videos on transformers were sometimes too dense for viewers without a machine learning background.
The video's central idea is the movie-script analogy: a large language model is fundamentally a very powerful next-word predictor, one that has learned to assign probabilities to every possible next word rather than committing to a single guess. That's why the same prompt can produce a different response each time, since the model is sampling from a distribution rather than following a fixed rule. From there, the video briefly touches on how models like GPT are trained on enormous volumes of text, and how that pretraining process is what gives them the raw ability to generate coherent, human-like language before any fine-tuning happens.
What makes this explanation valuable isn't just its simplicity, it's how honestly it scopes itself. The video is explicitly framed as a light introduction, with links to Sanderson's deeper technical series on transformers and attention for viewers who want to go further. That layered approach, a fast on-ramp followed by an optional deep end, is a useful model in itself for how to introduce a technical audience to a complex topic.
Who is this for:
Developers, tech-savvy readers, and anyone who wants to understand what's actually happening behind ChatGPT-style tools.