02.09 · UNIT 03 · PyTorch in One Hour: saving and scale · Lesson
Training on GPUs
A GPU runs the same tensor operations faster. Move the model and the data to one device.
PLAIN-LANGUAGE INTRODUCTION
What is this?
A GPU runs the same tensor operations faster. Move the model and the data to one device.
One simple example
After .to('cuda'), adding [1,2,3] and [4,5,6] gives tensor([5.,7.,9.], device='cuda:0').
What goes in?
A model, data batches, and a device name such as cuda or mps.
What comes out?
The same results, computed on the chosen device.
Why does it matter?
Large models and datasets train much faster on a GPU.
What is it not?
A small dataset gives no speed-up because the transfer cost dominates.
WORK THROUGH THE IDEA
See the idea in more detail
- A tensor lives on one device. All tensors in one operation must use the same device.
- Move tensors and models with
.to(). Single-GPU training adds a device line, a model move, and a data move. - The portable form
torch.device('cuda' if torch.cuda.is_available() else 'cpu')runs with or without a GPU. DistributedDataParallel(DDP) starts one process per GPU, gives each a unique minibatch, and synchronizes the gradients.- Common mistake: DDP needs several processes, so run it as a script with
torchrun, not in a notebook.