02.05 · UNIT 02 · PyTorch in One Hour: models and training · Lesson
Multilayer neural networks
A multilayer perceptron stacks linear layers with nonlinear activations.
PLAIN-LANGUAGE INTRODUCTION
What is this?
A multilayer perceptron stacks linear layers with nonlinear activations.
One simple example
A network with 50 inputs, hidden layers of 30 and 20 units, and 3 outputs has 2213 trainable parameters.
What goes in?
A batch with 50 features per example.
What comes out?
Three raw scores, called logits, for each example.
Why does it matter?
nn.Module registers every parameter so the optimizer and the checkpoint can find them.
What is it not?
The model returns logits. It does not apply softmax for you.
WORK THROUGH THE IDEA
See the idea in more detail
- Subclass
torch.nn.Module. Put the layers in__init__and the data flow inforward. - Each
nn.Linearholds a weight matrix and a bias vector. The three layers hold1530,620, and63trainable values. torch.nn.Sequentialruns its layers in order. TheReLUlayers hold no parameters.- Use
torch.no_grad()for inference and applysoftmaxyourself for probabilities. - Common mistake: applying softmax before a loss that expects raw logits.