Tensor by Tensor

02.05 · UNIT 02 · PyTorch in One Hour: models and training · Lesson

Multilayer neural networks

A multilayer perceptron stacks linear layers with nonlinear activations.

PLAIN-LANGUAGE INTRODUCTION

What is this?

A multilayer perceptron stacks linear layers with nonlinear activations.

One simple example

A network with 50 inputs, hidden layers of 30 and 20 units, and 3 outputs has 2213 trainable parameters.

What goes in?

A batch with 50 features per example.

What comes out?

Three raw scores, called logits, for each example.

Why does it matter?

nn.Module registers every parameter so the optimizer and the checkpoint can find them.

What is it not?

The model returns logits. It does not apply softmax for you.

WORK THROUGH THE IDEA

See the idea in more detail

  1. Subclass torch.nn.Module. Put the layers in __init__ and the data flow in forward.
  2. Each nn.Linear holds a weight matrix and a bias vector. The three layers hold 1530, 620, and 63 trainable values.
  3. torch.nn.Sequential runs its layers in order. The ReLU layers hold no parameters.
  4. Use torch.no_grad() for inference and apply softmax yourself for probabilities.
  5. Common mistake: applying softmax before a loss that expects raw logits.
Open the detailed notes ↗