Lesson 9
Training on GPUs
Introduction
A GPU trains a network faster than a CPU. This lesson shows GPU compute, single-GPU training, and multi-GPU training with DDP.
- Learning goal
Move a model and its data to a GPU, run a training loop on one device, and start a distributed training script with one process per GPU.
- Before you start
Training loops, model modules, data loaders, and the loss and optimizer steps from earlier lessons.
Lesson plan
- Name the device of a tensor and move tensors to a GPU.
- Add three lines to the training loop for single-GPU training.
- Split a dataset across GPUs and synchronize gradients with DDP.
9.1 PyTorch computations on GPU devices
A device is where data lives and where operations run. The CPU and the GPU are devices. A tensor lives on one device. Its operations run on the same device.
First, check that the runtime supports GPU compute.
print(torch.cuda.is_available()) # True
Add two tensors. The computation runs on the CPU by default.
tensor_1 = torch.tensor([1., 2., 3.])
tensor_2 = torch.tensor([4., 5., 6.])
print(tensor_1 + tensor_2)
# Expected: tensor([5., 7., 9.])
Move the tensors to the GPU with the .to() method. Run the addition there.
tensor_1 = tensor_1.to("cuda")
tensor_2 = tensor_2.to("cuda")
print(tensor_1 + tensor_2)
# Expected: tensor([5., 7., 9.], device='cuda:0')
The result shows device='cuda:0'. The tensors live on the first GPU. With several GPUs, choose a device with .to("cuda:0"), .to("cuda:1"), and so on.
All tensors in one operation must use the same device. One CPU tensor and one GPU tensor fail with a RuntimeError.
tensor_1 = torch.tensor([1., 2., 3.]).to("cuda")
tensor_2 = torch.tensor([4., 5., 6.]) # still on the CPU
print(tensor_1 + tensor_2)
# Expected:
# RuntimeError: Expected all tensors to be on the same device,
# but found at least two devices, cuda:0 and cpu!
Move both tensors to one GPU. PyTorch does the rest.
9.2 Single-GPU training
Change three lines of the training loop. The new lines define a device, move the model, and move the data.
torch.manual_seed(123)
model = NeuralNetwork(num_inputs=2, num_outputs=2)
# New: define a device that defaults to a GPU.
device = torch.device("cuda")
# New: transfer the model to the GPU.
model.to(device)
optimizer = torch.optim.SGD(model.parameters(), lr=0.5)
num_epochs = 3
for epoch in range(num_epochs):
model.train()
for batch_idx, (features, labels) in enumerate(train_loader):
# New: transfer the data to the GPU.
features, labels = features.to(device), labels.to(device)
logits = model(features)
loss = F.cross_entropy(logits, labels)
optimizer.zero_grad()
loss.backward()
optimizer.step()
model.eval()
The output is the same as on the CPU. This result is a good check.
# Expected:
# Epoch: 001/003 | Batch 000/002 | Train/Val Loss: 0.75
# Epoch: 001/003 | Batch 001/002 | Train/Val Loss: 0.65
# Epoch: 002/003 | Batch 000/002 | Train/Val Loss: 0.44
# Epoch: 002/003 | Batch 001/002 | Train/Val Loss: 0.13
# Epoch: 003/003 | Batch 000/002 | Train/Val Loss: 0.03
# Epoch: 003/003 | Batch 001/002 | Train/Val Loss: 0.00
You can write .to("cuda") instead of device = torch.device("cuda"). A portable form runs on a CPU if no GPU is present. Use this form in shared code.
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
This toy dataset gives no speed-up. The transfer from the CPU to the GPU costs time. Large models and large datasets gain a lot.
Apple Silicon. Use this line to use the M-series chip.
device = torch.device("mps" if torch.backends.mps.is_available() else "cpu")
Common pitfalls
- Move the model and the data to the same device. A mismatch causes a RuntimeError.
- Keep the device in one variable. Do not write
"cuda"in several lines. - A small dataset gives no speed-up, because the transfer cost is larger than the compute.
- A Mac uses the device name
"mps", not"cuda".
9.3 Training with multiple GPUs
Distributed training divides the training across several GPUs and machines.
One GPU can be too slow, and model development needs many training rounds. More devices cut the training time.
The basic case is PyTorch DistributedDataParallel, or DDP. DDP splits the input data across the devices. Each device processes its subset at the same time.
How DDP works:
- PyTorch starts one process on each GPU.
- Each process holds a copy of the model.
- Each copy gets a different, non-overlapping minibatch. The
DistributedSamplermakes the non-overlapping batches. - Each copy sees different samples. The copies give different logits and different gradients.
- PyTorch averages and synchronizes the gradients. Then the copies do not diverge.


The benefit of DDP is more speed for the dataset. Two GPUs can process an epoch in about half the time of one GPU. The time decreases more with the number of GPUs. Eight GPUs can process an epoch about eight times faster.
Common pitfalls
- Do not use DDP in a notebook. DDP starts multiple processes, and each process needs its own Python interpreter.
- Run the DDP code as a script with
torchrun, not in Jupyter.
Import the extra classes and functions for distributed training.
import platform
from torch.utils.data.distributed import DistributedSampler
from torch.nn.parallel import DistributedDataParallel as DDP
from torch.distributed import init_process_group, destroy_process_group
- DistributedSampler divides the dataset among the processes.
- init_process_group starts the distributed mode. Call it at the start of the training script.
- destroy_process_group stops the distributed mode. Call it at the end, to free the resources.
The helper ddp_setup starts the process group. It sets the master address and port, unless torchrun already provides them. It uses the NCCL backend. Then it sets the device for the process.
def ddp_setup(rank, world_size):
# Only set MASTER_ADDR and MASTER_PORT if torchrun does not define them.
if "MASTER_ADDR" not in os.environ:
os.environ["MASTER_ADDR"] = "localhost"
if "MASTER_PORT" not in os.environ:
os.environ["MASTER_PORT"] = "12345"
if platform.system() == "Windows":
os.environ["USE_LIBUV"] = "0"
init_process_group(backend="gloo", rank=rank, world_size=world_size)
else:
init_process_group(backend="nccl", rank=rank, world_size=world_size)
torch.cuda.set_device(rank)
The main function wraps the model with DDP. DDP synchronizes the gradient updates across all GPUs.
def main(rank, world_size, num_epochs):
ddp_setup(rank, world_size)
train_loader, test_loader = prepare_dataset()
model = NeuralNetwork(num_inputs=2, num_outputs=2)
model.to(rank)
optimizer = torch.optim.SGD(model.parameters(), lr=0.5)
model = DDP(model, device_ids=[rank]) # wrap the model with DDP
for epoch in range(num_epochs):
train_loader.sampler.set_epoch(epoch) # reshuffle
model.train()
for features, labels in train_loader:
features, labels = features.to(rank), labels.to(rank)
logits = model(features)
loss = F.cross_entropy(logits, labels)
optimizer.zero_grad()
loss.backward()
optimizer.step()
print(f"[GPU{rank}] Epoch: {epoch+1:03d}/{num_epochs:03d}"
f" | Batchsize {labels.shape[0]:03d}"
f" | Train/Val Loss: {loss:.2f}")
model.eval()
train_acc = compute_accuracy(model, train_loader, device=rank)
print(f"[GPU{rank}] Training accuracy", train_acc)
test_acc = compute_accuracy(model, test_loader, device=rank)
print(f"[GPU{rank}] Test accuracy", test_acc)
destroy_process_group()
This script shows DDP training for the NeuralNetwork model. Open the box for the full script.
Full DDP script (DDP-script-torchrun.py)
import torch
import torch.nn.functional as F
from torch.utils.data import Dataset, DataLoader
# NEW imports:
import os
import platform
from torch.utils.data.distributed import DistributedSampler
from torch.nn.parallel import DistributedDataParallel as DDP
from torch.distributed import init_process_group, destroy_process_group
# NEW: initialize a distributed process group (1 process / GPU)
def ddp_setup(rank, world_size):
# Only set MASTER_ADDR and MASTER_PORT if not defined by torchrun
if "MASTER_ADDR" not in os.environ:
os.environ["MASTER_ADDR"] = "localhost"
if "MASTER_PORT" not in os.environ:
os.environ["MASTER_PORT"] = "12345"
if platform.system() == "Windows":
os.environ["USE_LIBUV"] = "0"
init_process_group(backend="gloo", rank=rank, world_size=world_size)
else:
init_process_group(backend="nccl", rank=rank, world_size=world_size)
torch.cuda.set_device(rank)
class ToyDataset(Dataset):
def __init__(self, X, y):
self.features = X
self.labels = y
def __getitem__(self, index):
return self.features[index], self.labels[index]
def __len__(self):
return self.labels.shape[0]
class NeuralNetwork(torch.nn.Module):
def __init__(self, num_inputs, num_outputs):
super().__init__()
self.layers = torch.nn.Sequential(
torch.nn.Linear(num_inputs, 30),
torch.nn.ReLU(),
torch.nn.Linear(30, 20),
torch.nn.ReLU(),
torch.nn.Linear(20, num_outputs),
)
def forward(self, x):
return self.layers(x)
def prepare_dataset():
X_train = torch.tensor([
[-1.2, 3.1], [-0.9, 2.9], [-0.5, 2.6],
[ 2.3, -1.1], [ 2.7, -1.5]
])
y_train = torch.tensor([0, 0, 0, 1, 1])
X_test = torch.tensor([[-0.8, 2.8], [2.6, -1.6]])
y_test = torch.tensor([0, 1])
# Uncomment to enlarge the dataset for up to 8 GPUs:
# factor = 4
# X_train = torch.cat([X_train + torch.randn_like(X_train) * 0.1 for _ in range(factor)])
# y_train = y_train.repeat(factor)
train_ds = ToyDataset(X_train, y_train)
test_ds = ToyDataset(X_test, y_test)
train_loader = DataLoader(
dataset=train_ds, batch_size=2, shuffle=False,
pin_memory=True, drop_last=True,
sampler=DistributedSampler(train_ds) # NEW
)
test_loader = DataLoader(dataset=test_ds, batch_size=2, shuffle=False)
return train_loader, test_loader
def main(rank, world_size, num_epochs):
ddp_setup(rank, world_size)
train_loader, test_loader = prepare_dataset()
model = NeuralNetwork(num_inputs=2, num_outputs=2)
model.to(rank)
optimizer = torch.optim.SGD(model.parameters(), lr=0.5)
model = DDP(model, device_ids=[rank]) # NEW: wrap model with DDP
for epoch in range(num_epochs):
train_loader.sampler.set_epoch(epoch) # NEW: reshuffle
model.train()
for features, labels in train_loader:
features, labels = features.to(rank), labels.to(rank)
logits = model(features)
loss = F.cross_entropy(logits, labels)
optimizer.zero_grad()
loss.backward()
optimizer.step()
print(f"[GPU{rank}] Epoch: {epoch+1:03d}/{num_epochs:03d}"
f" | Batchsize {labels.shape[0]:03d}"
f" | Train/Val Loss: {loss:.2f}")
model.eval()
train_acc = compute_accuracy(model, train_loader, device=rank)
print(f"[GPU{rank}] Training accuracy", train_acc)
test_acc = compute_accuracy(model, test_loader, device=rank)
print(f"[GPU{rank}] Test accuracy", test_acc)
destroy_process_group() # NEW
def compute_accuracy(model, dataloader, device):
model = model.eval()
correct = 0.0
total_examples = 0
for idx, (features, labels) in enumerate(dataloader):
features, labels = features.to(device), labels.to(device)
with torch.no_grad():
logits = model(features)
predictions = torch.argmax(logits, dim=1)
correct += torch.sum(labels == predictions)
total_examples += len(labels)
return (correct / total_examples).item()
if __name__ == "__main__":
if "WORLD_SIZE" in os.environ:
world_size = int(os.environ["WORLD_SIZE"])
else:
world_size = 1
if "LOCAL_RANK" in os.environ:
rank = int(os.environ["LOCAL_RANK"])
elif "RANK" in os.environ:
rank = int(os.environ["RANK"])
else:
rank = 0
if rank == 0: # print only on rank 0
print("PyTorch version:", torch.__version__)
print("CUDA available:", torch.cuda.is_available())
print("Number of GPUs available:", torch.cuda.device_count())
torch.manual_seed(123)
num_epochs = 3
main(rank, world_size, num_epochs)
How the script works:
- The
__main__block runs when the file runs as a standalone script. It prints the number of GPUs. It sets the random seed. - Use
torchrunto start the script. It starts one process per GPU. It gives each process a rank and other data through environment variables. - The
__main__block reads these variables withos.environ. It passes them tomain(). main()callsddp_setup, loads the data, makes the model, and runs the loop.- Move the model and the data to the correct GPU with
.to(rank). The rank is the GPU index for the process. - Wrap the model with DDP. DDP synchronizes the gradient updates across all GPUs.
- The training loader uses
sampler=DistributedSampler(train_ds). ddp_setupsets the master address and port, unlesstorchrunalready provides them. It starts the process group with the NCCL backend. It sets the device for the process.
Run the script with torchrun. It comes with PyTorch.
torchrun --nproc_per_node=2 DDP-script-torchrun.py
# To run on all available GPUs:
torchrun --nproc_per_node=$(nvidia-smi -L | wc -l) DDP-script-torchrun.py
The tiny dataset limits the GPU count. Uncomment the lines with factor = 4 in the script. The larger dataset works on up to 8 GPUs.
Output on one GPU:
# Expected:
# PyTorch version: 2.0.1+cu117
# CUDA available: True
# Number of GPUs available: 1
# [GPU0] Epoch: 001/003 | Batchsize 002 | Train/Val Loss: 0.62
# [GPU0] Epoch: 001/003 | Batchsize 002 | Train/Val Loss: 0.32
# [GPU0] Epoch: 002/003 | Batchsize 002 | Train/Val Loss: 0.11
# [GPU0] Epoch: 002/003 | Batchsize 002 | Train/Val Loss: 0.07
# [GPU0] Epoch: 003/003 | Batchsize 002 | Train/Val Loss: 0.02
# [GPU0] Epoch: 003/003 | Batchsize 002 | Train/Val Loss: 0.03
# [GPU0] Training accuracy 1.0
# [GPU0] Test accuracy 1.0
Output on two GPUs:
# Expected:
# Number of GPUs available: 2
# [GPU1] Epoch: 001/003 | Batchsize 002 | Train/Val Loss: 0.60
# [GPU0] Epoch: 001/003 | Batchsize 002 | Train/Val Loss: 0.59
# [GPU0] Epoch: 002/003 | Batchsize 002 | Train/Val Loss: 0.16
# [GPU1] Epoch: 002/003 | Batchsize 002 | Train/Val Loss: 0.17
# [GPU0] Epoch: 003/003 | Batchsize 002 | Train/Val Loss: 0.05
# [GPU1] Epoch: 003/003 | Batchsize 002 | Train/Val Loss: 0.05
GPU0 processes some batches. GPU1 processes the other batches. The accuracy lines repeat, because each process prints the result. Guard the print with if rank == 0 to stop the repeated lines.
if rank == 0: # only print in the first process
print("Test accuracy: ", accuracy)
Try it
You have one CPU tensor and one CUDA tensor. Describe the fastest fix and the error without the fix. Then write the portable device line for a script that must run with or without a GPU.
Reveal the worked answer
# Move both tensors to one device.
tensor_1 = tensor_1.to("cuda")
tensor_2 = tensor_2.to("cuda")
print(tensor_1 + tensor_2)
# Expected: tensor([5., 7., 9.], device='cuda:0')
# Without the move:
# RuntimeError: Expected all tensors to be on the same device,
# but found at least two devices, cuda:0 and cpu!
# Portable device line:
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
Move both tensors to the same device before the operation. A mixed-device operation gives a RuntimeError. The portable line chooses the GPU when one is present, and the CPU when one is absent.
Further reading
Books:
- Machine Learning with PyTorch and Scikit-Learn (2022) by Sebastian Raschka, Hayden Liu, and Vahid Mirjalili.
- Deep Learning with PyTorch (2021) by Eli Stevens, Luca Antiga, and Thomas Viehmann.
Articles:
- Introduction to Calculus by Sebastian Raschka. PDF.
- Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning (2018) by Sebastian Raschka. arXiv.
- Finetuning Large Language Models On A Single GPU Using Gradient Accumulation by Sebastian Raschka. Article.
Large models
DDP trains a model that fits on one GPU. If one model does not fit on one GPU, use Fully Sharded Data Parallel (FSDP). FSDP distributes large layers across several GPUs.
Recap
A tensor lives on one device, and its operations run on that device. Move tensors and models with .to(). Keep every tensor in one operation on the same device.
Single-GPU training needs three new lines: define the device, move the model, and move the data. Use the portable device form in shared code.
DDP starts one process per GPU, gives each process a model copy and a unique minibatch, and synchronizes the gradients. Run DDP as a script with torchrun, not in a notebook.
Reference: DistributedDataParallel documentation.