Deep Learning from the Inside Out
A ten-part series that takes modern language models apart, the transformer, BERT, GPT-2, T5, and LLaMA, down to their weight matrices, tensor shapes, and FLOP counts. Every claim is one you can check with arithmetic. Written from my UC Berkeley CS199 study and hand-illustrated throughout.
Read the series →
