JK. V. Jobin / Learning notes
Deep learning · Sample lecture 01 / 07

Deep learning / 01

When does a model stop learning?

Generalization, overfitting, and the small interventions that help neural networks travel beyond their training data.

K. V. JobinLecture notes · 8 min
EPOCH 01EPOCH 30 LOSSGENERALIZATION GAP
TrainingValidation

01 / The failure mode

A low training loss can lie.

A model can memorize the examples it has seen and still fail on new ones. Watch both curves: validation loss is the signal that tells us when memorization starts to win.

Δthe generalization gap
Lval − LtrainMeasure the difference on held-out data.
TRAININGEPOCHBEST CHECKPOINT

02 / The toolkit

Four ways to constrain a model.

01 / WEIGHTS

Weight decay

Penalize large weights to favor smoother functions.

AdamW
02 / ACTIVATIONS

Dropout

Prevent hidden units from relying on one another.

nn.Dropout
03 / SCALE

Batch norm

Keep intermediate activations on a stable scale.

nn.BatchNorm
04 / TIME

Early stopping

Keep the checkpoint with the best validation loss.

patience = 5

03 / One objective, two goals

Regularize the function, not the score.

Weight decay asks for a model that fits the data without using unnecessarily large parameters.

Ltotalwhat training minimizes
=
Ldatafit the observed examples
+
λ ||w||²prefer smaller weights

Increase λ to constrain the model more. Too much regularization can make it underfit; tune against validation data.

04 / A repeatable workflow

Make validation part of the loop.

01 / SPLIT

Hold out data

Keep validation examples away from gradient updates.

02 / TRAIN

Fit the model

Update weights on training batches only.

03 / MONITOR

Track the trend

Evaluate validation loss after each epoch.

04 / RESTORE

Keep the best

Stop after patience runs out; restore best weights.

best_epoch = arg minepoch LvalChoose by a signal the optimizer did not train on.

05 / Evaluation discipline

The test set is a one-shot exam.

Use validation results to choose settings. Touch the test set once, after decisions are final.

TRAIN

Fit parameters.

VALIDATE

Choose the recipe.

TEST

Report once.

06 / Takeaways

Learn the signal. Leave the noise.

01

Training performance alone does not measure generalization.

02

Regularizers constrain different parts of the learning process.

03

Validation loss guides stopping and model selection.

04

Keep the test set out of the tuning loop.

Next: try it in code

Compare a plain MNIST classifier with weight decay, dropout, batch normalization, and early stopping in the accompanying tutorial.

Presenter mode

Up next