Free through lesson 3

PyTorch & Deep Learning for Beginners

From tensors and automatic differentiation to neural network basics, techniques for stable training, CNNs and transfer learning, Attention and Transformers, faster inference, and shipping your model. Across 50 lessons, you'll get to the point where you can build a model's internals yourself, train it, and ship it as ONNX and an API. Every piece of code and output was actually run on PyTorch 2.14.0.

Curriculum

The 50 lessons are split into 7 chapters. We recommend going in order from chapter 1, but feel free to jump to whatever interests you. Note: The lessons use torch, so they can't run in the browser. Create a virtual environment on your machine, run pip install torch torchvision, and try things there. You don't need a GPU. If CUDA is available you can use cuda, and on an Apple Silicon Mac you can use mps.

Chapter 1 — Tensors and Automatic Differentiation (lessons 1–8)

You'll get a handle on the two foundations of PyTorch: tensors and automatic differentiation. At the end, you'll solve linear regression with plain gradient descent, without an optimizer.

Chapter 2 — Neural Network Basics (lessons 9–16)

You'll run fully connected layers, activation functions, loss functions, nn.Module, the training loop, and DataLoader. At the end, you'll write one complete binary classifier, from standardization all the way to a confusion matrix.

Chapter 3 — Stabilizing Training (lessons 17–24)

You'll compare learning rates, optimizers, schedulers, initialization, regularization, Dropout, and normalization layers, all on the same data, and check the numbers. At the end, you'll split off validation data and implement early stopping.

Chapter 4 — Images (lessons 25–32)

From how convolution and pooling work to residual connections, data augmentation, and transfer learning. At the end, you'll fine-tune a pretrained resnet18 and measure how it compares to training from scratch.

Chapter 5 — Sequences and Attention (lessons 33–42)

You'll measure the limits of RNNs, write Attention yourself, and assemble a Transformer block. At the end, you'll train a character-level language model on your own Transformer and have it generate sentences.

33

Sequence Data Shapes and Padding

Three dimensions: (batch, time, features). When to use masks vs. pack to align sequences of different lengths.

🔒 Basic
34

Embeddings with nn.Embedding

Same result as one-hot times a matrix, but faster since it just looks up a row. Including what padding_idx means.

🔒 Basic
35

How RNNs Work and Their Limits

The same weights are reused at every time step. Show, in orders of magnitude, how gradients vanish the further back in time they go.

🔒 Basic
36

LSTM and GRU

Add gates so the model learns what to remember and what to forget. Measure the difference against an RNN under the same conditions.

🔒 Basic
37

Writing Attention by Hand

Just softmax the Q·K dot products and mix V. Confirm with numbers why we divide by sqrt(d).

🔒 Basic
38

Multi-Head Attention

Split the dimensions and run attention in parallel. Changing the number of heads doesn't change the parameter count.

🔒 Basic
39

Positional Encoding

Attention alone can't tell word order. The sinusoidal formula vs. learned position embeddings.

🔒 Basic
40

Assembling a Transformer Block

Attention → FFN, each with a residual connection and LayerNorm. Pre-LN vs. Post-LN changes how gradients flow by orders of magnitude.

🔒 Basic
41

Predicting Characters with Your Own Transformer

In 600K parameters and 35 seconds, teach it the grammar of a corpus generated from templates.

🔒 Basic
42

Using the Built-In Transformer

nn.TransformerEncoderLayer has the same internals as yours. Just watch out for the mask types and the norm_first default.

🔒 Basic

Chapter 6 — Real-World Practice (lessons 43–47)

Experiment logging, reproducibility, faster inference, saving and loading, TorchScript and torch.compile. You'll gather the tools you need to reproduce the same results again and to make models fast enough for production.

Chapter 7 — Finishing Up (lessons 48–50)

You'll export to ONNX, wrap the model in an inference API with FastAPI, and get it into a shape you can ship. Finally, you'll fold the tools from all 50 lessons into a single script and sort out where to go next.

Once you've finished all 50 lessons, learn data preprocessing and classic methods in Python & Machine Learning for Beginners, and how to call large pretrained models and build apps on them in Intro to AI App Development. To put an inference API into production, head to Intro to Shipping and Running a Service. A membership unlocks every course.