Tensor by Tensor

02.09 · UNIT 03 · PyTorch in One Hour: saving and scale · Lesson

Training on GPUs

A GPU runs the same tensor operations faster. Move the model and the data to one device.

PLAIN-LANGUAGE INTRODUCTION

What is this?

A GPU runs the same tensor operations faster. Move the model and the data to one device.

One simple example

After .to('cuda'), adding [1,2,3] and [4,5,6] gives tensor([5.,7.,9.], device='cuda:0').

What goes in?

A model, data batches, and a device name such as cuda or mps.

What comes out?

The same results, computed on the chosen device.

Why does it matter?

Large models and datasets train much faster on a GPU.

What is it not?

A small dataset gives no speed-up because the transfer cost dominates.

WORK THROUGH THE IDEA

See the idea in more detail

  1. A tensor lives on one device. All tensors in one operation must use the same device.
  2. Move tensors and models with .to(). Single-GPU training adds a device line, a model move, and a data move.
  3. The portable form torch.device('cuda' if torch.cuda.is_available() else 'cpu') runs with or without a GPU.
  4. DistributedDataParallel (DDP) starts one process per GPU, gives each a unique minibatch, and synchronizes the gradients.
  5. Common mistake: DDP needs several processes, so run it as a script with torchrun, not in a notebook.
Open the detailed notes ↗