Tensor by Tensor

02.06 · UNIT 02 · PyTorch in One Hour: models and training · Lesson

Data loaders

A Dataset defines one record. A DataLoader shuffles records and groups them into batches.

PLAIN-LANGUAGE INTRODUCTION

What is this?

A Dataset defines one record. A DataLoader shuffles records and groups them into batches.

One simple example

A loader with batch_size=2 over 5 examples gives batches of 2, 2, and 1.

What goes in?

Feature and label tensors, plus a batch size and a shuffle choice.

What comes out?

One batch of features and labels per loop step.

Why does it matter?

Separating record lookup from batching keeps both jobs simple and reusable.

What is it not?

drop_last=True on a test loader removes real examples from evaluation.

WORK THROUGH THE IDEA

See the idea in more detail

  1. A custom Dataset needs __init__, __getitem__, and __len__.
  2. __getitem__ returns one record. __len__ returns the number of records.
  3. shuffle=True mixes the order before each epoch. One complete pass is one epoch.
  4. drop_last=True drops a short final batch. num_workers greater than 0 loads data in parallel processes.
  5. Common mistake: labels must start at 0 and use torch.long. Keep features as floats.
Open the detailed notes ↗