Lesson 9

Training on GPUs

Introduction

A GPU trains a network faster than a CPU. This lesson shows GPU compute, single-GPU training, and multi-GPU training with DDP.

Learning goal

Move a model and its data to a GPU, run a training loop on one device, and start a distributed training script with one process per GPU.

Before you start

Training loops, model modules, data loaders, and the loss and optimizer steps from earlier lessons.

Lesson plan

  1. Name the device of a tensor and move tensors to a GPU.
  2. Add three lines to the training loop for single-GPU training.
  3. Split a dataset across GPUs and synchronize gradients with DDP.

9.1 PyTorch computations on GPU devices

A device is where data lives and where operations run. The CPU and the GPU are devices. A tensor lives on one device. Its operations run on the same device.

First, check that the runtime supports GPU compute.

print(torch.cuda.is_available())   # True

Add two tensors. The computation runs on the CPU by default.

tensor_1 = torch.tensor([1., 2., 3.])
tensor_2 = torch.tensor([4., 5., 6.])

print(tensor_1 + tensor_2)
# Expected: tensor([5., 7., 9.])

Move the tensors to the GPU with the .to() method. Run the addition there.

tensor_1 = tensor_1.to("cuda")
tensor_2 = tensor_2.to("cuda")

print(tensor_1 + tensor_2)
# Expected: tensor([5., 7., 9.], device='cuda:0')

The result shows device='cuda:0'. The tensors live on the first GPU. With several GPUs, choose a device with .to("cuda:0"), .to("cuda:1"), and so on.

All tensors in one operation must use the same device. One CPU tensor and one GPU tensor fail with a RuntimeError.

tensor_1 = torch.tensor([1., 2., 3.]).to("cuda")
tensor_2 = torch.tensor([4., 5., 6.])          # still on the CPU

print(tensor_1 + tensor_2)
# Expected:
# RuntimeError: Expected all tensors to be on the same device,
# but found at least two devices, cuda:0 and cpu!

Move both tensors to one GPU. PyTorch does the rest.

9.2 Single-GPU training

Change three lines of the training loop. The new lines define a device, move the model, and move the data.

torch.manual_seed(123)
model = NeuralNetwork(num_inputs=2, num_outputs=2)

# New: define a device that defaults to a GPU.
device = torch.device("cuda")
# New: transfer the model to the GPU.
model.to(device)

optimizer = torch.optim.SGD(model.parameters(), lr=0.5)
num_epochs = 3

for epoch in range(num_epochs):
    model.train()
    for batch_idx, (features, labels) in enumerate(train_loader):
        # New: transfer the data to the GPU.
        features, labels = features.to(device), labels.to(device)

        logits = model(features)
        loss = F.cross_entropy(logits, labels)

        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

    model.eval()

The output is the same as on the CPU. This result is a good check.

# Expected:
# Epoch: 001/003 | Batch 000/002 | Train/Val Loss: 0.75
# Epoch: 001/003 | Batch 001/002 | Train/Val Loss: 0.65
# Epoch: 002/003 | Batch 000/002 | Train/Val Loss: 0.44
# Epoch: 002/003 | Batch 001/002 | Train/Val Loss: 0.13
# Epoch: 003/003 | Batch 000/002 | Train/Val Loss: 0.03
# Epoch: 003/003 | Batch 001/002 | Train/Val Loss: 0.00

You can write .to("cuda") instead of device = torch.device("cuda"). A portable form runs on a CPU if no GPU is present. Use this form in shared code.

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

This toy dataset gives no speed-up. The transfer from the CPU to the GPU costs time. Large models and large datasets gain a lot.

Apple Silicon. Use this line to use the M-series chip.

device = torch.device("mps" if torch.backends.mps.is_available() else "cpu")

Common pitfalls

  • Move the model and the data to the same device. A mismatch causes a RuntimeError.
  • Keep the device in one variable. Do not write "cuda" in several lines.
  • A small dataset gives no speed-up, because the transfer cost is larger than the compute.
  • A Mac uses the device name "mps", not "cuda".

9.3 Training with multiple GPUs

Distributed training divides the training across several GPUs and machines.

One GPU can be too slow, and model development needs many training rounds. More devices cut the training time.

The basic case is PyTorch DistributedDataParallel, or DDP. DDP splits the input data across the devices. Each device processes its subset at the same time.

How DDP works:

  1. PyTorch starts one process on each GPU.
  2. Each process holds a copy of the model.
  3. Each copy gets a different, non-overlapping minibatch. The DistributedSampler makes the non-overlapping batches.
  4. Each copy sees different samples. The copies give different logits and different gradients.
  5. PyTorch averages and synchronizes the gradients. Then the copies do not diverge.
DDP copies the model onto each GPU and splits the data into unique minibatches.
Figure 12: Sebastian Raschka, "PyTorch in One Hour" (source).
Forward and backward pass run on each GPU, then the gradients are synchronized.
Figure 13: Sebastian Raschka, "PyTorch in One Hour" (source).

The benefit of DDP is more speed for the dataset. Two GPUs can process an epoch in about half the time of one GPU. The time decreases more with the number of GPUs. Eight GPUs can process an epoch about eight times faster.

Common pitfalls

  • Do not use DDP in a notebook. DDP starts multiple processes, and each process needs its own Python interpreter.
  • Run the DDP code as a script with torchrun, not in Jupyter.

Import the extra classes and functions for distributed training.

import platform
from torch.utils.data.distributed import DistributedSampler
from torch.nn.parallel import DistributedDataParallel as DDP
from torch.distributed import init_process_group, destroy_process_group

The helper ddp_setup starts the process group. It sets the master address and port, unless torchrun already provides them. It uses the NCCL backend. Then it sets the device for the process.

def ddp_setup(rank, world_size):
    # Only set MASTER_ADDR and MASTER_PORT if torchrun does not define them.
    if "MASTER_ADDR" not in os.environ:
        os.environ["MASTER_ADDR"] = "localhost"
    if "MASTER_PORT" not in os.environ:
        os.environ["MASTER_PORT"] = "12345"

    if platform.system() == "Windows":
        os.environ["USE_LIBUV"] = "0"
        init_process_group(backend="gloo", rank=rank, world_size=world_size)
    else:
        init_process_group(backend="nccl", rank=rank, world_size=world_size)

    torch.cuda.set_device(rank)

The main function wraps the model with DDP. DDP synchronizes the gradient updates across all GPUs.

def main(rank, world_size, num_epochs):
    ddp_setup(rank, world_size)

    train_loader, test_loader = prepare_dataset()
    model = NeuralNetwork(num_inputs=2, num_outputs=2)
    model.to(rank)
    optimizer = torch.optim.SGD(model.parameters(), lr=0.5)

    model = DDP(model, device_ids=[rank])   # wrap the model with DDP

    for epoch in range(num_epochs):
        train_loader.sampler.set_epoch(epoch)   # reshuffle
        model.train()
        for features, labels in train_loader:
            features, labels = features.to(rank), labels.to(rank)
            logits = model(features)
            loss = F.cross_entropy(logits, labels)
            optimizer.zero_grad()
            loss.backward()
            optimizer.step()
            print(f"[GPU{rank}] Epoch: {epoch+1:03d}/{num_epochs:03d}"
                  f" | Batchsize {labels.shape[0]:03d}"
                  f" | Train/Val Loss: {loss:.2f}")

    model.eval()

    train_acc = compute_accuracy(model, train_loader, device=rank)
    print(f"[GPU{rank}] Training accuracy", train_acc)
    test_acc = compute_accuracy(model, test_loader, device=rank)
    print(f"[GPU{rank}] Test accuracy", test_acc)

    destroy_process_group()

This script shows DDP training for the NeuralNetwork model. Open the box for the full script.

Full DDP script (DDP-script-torchrun.py)
import torch
import torch.nn.functional as F
from torch.utils.data import Dataset, DataLoader

# NEW imports:
import os
import platform
from torch.utils.data.distributed import DistributedSampler
from torch.nn.parallel import DistributedDataParallel as DDP
from torch.distributed import init_process_group, destroy_process_group


# NEW: initialize a distributed process group (1 process / GPU)
def ddp_setup(rank, world_size):
    # Only set MASTER_ADDR and MASTER_PORT if not defined by torchrun
    if "MASTER_ADDR" not in os.environ:
        os.environ["MASTER_ADDR"] = "localhost"
    if "MASTER_PORT" not in os.environ:
        os.environ["MASTER_PORT"] = "12345"

    if platform.system() == "Windows":
        os.environ["USE_LIBUV"] = "0"
        init_process_group(backend="gloo", rank=rank, world_size=world_size)
    else:
        init_process_group(backend="nccl", rank=rank, world_size=world_size)

    torch.cuda.set_device(rank)


class ToyDataset(Dataset):
    def __init__(self, X, y):
        self.features = X
        self.labels = y

    def __getitem__(self, index):
        return self.features[index], self.labels[index]

    def __len__(self):
        return self.labels.shape[0]


class NeuralNetwork(torch.nn.Module):
    def __init__(self, num_inputs, num_outputs):
        super().__init__()
        self.layers = torch.nn.Sequential(
            torch.nn.Linear(num_inputs, 30),
            torch.nn.ReLU(),
            torch.nn.Linear(30, 20),
            torch.nn.ReLU(),
            torch.nn.Linear(20, num_outputs),
        )

    def forward(self, x):
        return self.layers(x)


def prepare_dataset():
    X_train = torch.tensor([
        [-1.2,  3.1], [-0.9,  2.9], [-0.5,  2.6],
        [ 2.3, -1.1], [ 2.7, -1.5]
    ])
    y_train = torch.tensor([0, 0, 0, 1, 1])
    X_test = torch.tensor([[-0.8, 2.8], [2.6, -1.6]])
    y_test = torch.tensor([0, 1])

    # Uncomment to enlarge the dataset for up to 8 GPUs:
    # factor = 4
    # X_train = torch.cat([X_train + torch.randn_like(X_train) * 0.1 for _ in range(factor)])
    # y_train = y_train.repeat(factor)

    train_ds = ToyDataset(X_train, y_train)
    test_ds  = ToyDataset(X_test, y_test)

    train_loader = DataLoader(
        dataset=train_ds, batch_size=2, shuffle=False,
        pin_memory=True, drop_last=True,
        sampler=DistributedSampler(train_ds)   # NEW
    )
    test_loader = DataLoader(dataset=test_ds, batch_size=2, shuffle=False)
    return train_loader, test_loader


def main(rank, world_size, num_epochs):
    ddp_setup(rank, world_size)

    train_loader, test_loader = prepare_dataset()
    model = NeuralNetwork(num_inputs=2, num_outputs=2)
    model.to(rank)
    optimizer = torch.optim.SGD(model.parameters(), lr=0.5)

    model = DDP(model, device_ids=[rank])   # NEW: wrap model with DDP

    for epoch in range(num_epochs):
        train_loader.sampler.set_epoch(epoch)   # NEW: reshuffle
        model.train()
        for features, labels in train_loader:
            features, labels = features.to(rank), labels.to(rank)
            logits = model(features)
            loss = F.cross_entropy(logits, labels)
            optimizer.zero_grad()
            loss.backward()
            optimizer.step()
            print(f"[GPU{rank}] Epoch: {epoch+1:03d}/{num_epochs:03d}"
                  f" | Batchsize {labels.shape[0]:03d}"
                  f" | Train/Val Loss: {loss:.2f}")

    model.eval()

    train_acc = compute_accuracy(model, train_loader, device=rank)
    print(f"[GPU{rank}] Training accuracy", train_acc)
    test_acc = compute_accuracy(model, test_loader, device=rank)
    print(f"[GPU{rank}] Test accuracy", test_acc)

    destroy_process_group()   # NEW


def compute_accuracy(model, dataloader, device):
    model = model.eval()
    correct = 0.0
    total_examples = 0
    for idx, (features, labels) in enumerate(dataloader):
        features, labels = features.to(device), labels.to(device)
        with torch.no_grad():
            logits = model(features)
        predictions = torch.argmax(logits, dim=1)
        correct += torch.sum(labels == predictions)
        total_examples += len(labels)
    return (correct / total_examples).item()


if __name__ == "__main__":
    if "WORLD_SIZE" in os.environ:
        world_size = int(os.environ["WORLD_SIZE"])
    else:
        world_size = 1

    if "LOCAL_RANK" in os.environ:
        rank = int(os.environ["LOCAL_RANK"])
    elif "RANK" in os.environ:
        rank = int(os.environ["RANK"])
    else:
        rank = 0

    if rank == 0:   # print only on rank 0
        print("PyTorch version:", torch.__version__)
        print("CUDA available:", torch.cuda.is_available())
        print("Number of GPUs available:", torch.cuda.device_count())

    torch.manual_seed(123)
    num_epochs = 3
    main(rank, world_size, num_epochs)

How the script works:

Run the script with torchrun. It comes with PyTorch.

torchrun --nproc_per_node=2 DDP-script-torchrun.py

# To run on all available GPUs:
torchrun --nproc_per_node=$(nvidia-smi -L | wc -l) DDP-script-torchrun.py

The tiny dataset limits the GPU count. Uncomment the lines with factor = 4 in the script. The larger dataset works on up to 8 GPUs.

Output on one GPU:

# Expected:
# PyTorch version: 2.0.1+cu117
# CUDA available: True
# Number of GPUs available: 1
# [GPU0] Epoch: 001/003 | Batchsize 002 | Train/Val Loss: 0.62
# [GPU0] Epoch: 001/003 | Batchsize 002 | Train/Val Loss: 0.32
# [GPU0] Epoch: 002/003 | Batchsize 002 | Train/Val Loss: 0.11
# [GPU0] Epoch: 002/003 | Batchsize 002 | Train/Val Loss: 0.07
# [GPU0] Epoch: 003/003 | Batchsize 002 | Train/Val Loss: 0.02
# [GPU0] Epoch: 003/003 | Batchsize 002 | Train/Val Loss: 0.03
# [GPU0] Training accuracy 1.0
# [GPU0] Test accuracy 1.0

Output on two GPUs:

# Expected:
# Number of GPUs available: 2
# [GPU1] Epoch: 001/003 | Batchsize 002 | Train/Val Loss: 0.60
# [GPU0] Epoch: 001/003 | Batchsize 002 | Train/Val Loss: 0.59
# [GPU0] Epoch: 002/003 | Batchsize 002 | Train/Val Loss: 0.16
# [GPU1] Epoch: 002/003 | Batchsize 002 | Train/Val Loss: 0.17
# [GPU0] Epoch: 003/003 | Batchsize 002 | Train/Val Loss: 0.05
# [GPU1] Epoch: 003/003 | Batchsize 002 | Train/Val Loss: 0.05

GPU0 processes some batches. GPU1 processes the other batches. The accuracy lines repeat, because each process prints the result. Guard the print with if rank == 0 to stop the repeated lines.

if rank == 0:  # only print in the first process
    print("Test accuracy: ", accuracy)

Try it

You have one CPU tensor and one CUDA tensor. Describe the fastest fix and the error without the fix. Then write the portable device line for a script that must run with or without a GPU.

Reveal the worked answer
# Move both tensors to one device.
tensor_1 = tensor_1.to("cuda")
tensor_2 = tensor_2.to("cuda")
print(tensor_1 + tensor_2)
# Expected: tensor([5., 7., 9.], device='cuda:0')

# Without the move:
# RuntimeError: Expected all tensors to be on the same device,
# but found at least two devices, cuda:0 and cpu!

# Portable device line:
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

Move both tensors to the same device before the operation. A mixed-device operation gives a RuntimeError. The portable line chooses the GPU when one is present, and the CPU when one is absent.

Further reading

Books:

Articles:

Large models

DDP trains a model that fits on one GPU. If one model does not fit on one GPU, use Fully Sharded Data Parallel (FSDP). FSDP distributes large layers across several GPUs.

Recap

A tensor lives on one device, and its operations run on that device. Move tensors and models with .to(). Keep every tensor in one operation on the same device.

Single-GPU training needs three new lines: define the device, move the model, and move the data. Use the portable device form in shared code.

DDP starts one process per GPU, gives each process a model copy and a unique minibatch, and synchronizes the gradients. Run DDP as a script with torchrun, not in a notebook.

Reference: DistributedDataParallel documentation.