Intro to LLMs
Andrej Karpathy's "Intro to Large Language Models" came out in 2023, and I still go back to it. The details have aged, but the story hasn't. It gives you the full picture of what an LLM is and how the ideas connect, from "a model is just two files" to pretraining, finetuning, and RLHF.
Every time I rewatched it, the same questions came up: what actually happens inside the network? How does text become numbers? What is attention doing? So I turned the first part of the talk into one interactive page, and I added the pieces I was missing: tokenization, embeddings, positional encoding, and self-attention, each with something you can click, drag, or play with.
If you're new to LLMs and want the big picture before the details, this is a good place to start.
You can also experiment with a larger width at: https://nora-alshareef.github.io/two-files-one-brain/




Comments