02.06 · UNIT 02 · PyTorch in One Hour: models and training · Lesson
Data loaders
A Dataset defines one record. A DataLoader shuffles records and groups them into batches.
PLAIN-LANGUAGE INTRODUCTION
What is this?
A Dataset defines one record. A DataLoader shuffles records and groups them into batches.
One simple example
A loader with batch_size=2 over 5 examples gives batches of 2, 2, and 1.
What goes in?
Feature and label tensors, plus a batch size and a shuffle choice.
What comes out?
One batch of features and labels per loop step.
Why does it matter?
Separating record lookup from batching keeps both jobs simple and reusable.
What is it not?
drop_last=True on a test loader removes real examples from evaluation.
WORK THROUGH THE IDEA
See the idea in more detail
- A custom
Datasetneeds__init__,__getitem__, and__len__. __getitem__returns one record.__len__returns the number of records.shuffle=Truemixes the order before each epoch. One complete pass is one epoch.drop_last=Truedrops a short final batch.num_workersgreater than 0 loads data in parallel processes.- Common mistake: labels must start at 0 and use
torch.long. Keep features as floats.