Weight decay
Penalize large weights to favor smoother functions.
AdamWDeep learning / 01
Generalization, overfitting, and the small interventions that help neural networks travel beyond their training data.
01 / The failure mode
A model can memorize the examples it has seen and still fail on new ones. Watch both curves: validation loss is the signal that tells us when memorization starts to win.
02 / The toolkit
Penalize large weights to favor smoother functions.
AdamWPrevent hidden units from relying on one another.
nn.DropoutKeep intermediate activations on a stable scale.
nn.BatchNormKeep the checkpoint with the best validation loss.
patience = 503 / One objective, two goals
Weight decay asks for a model that fits the data without using unnecessarily large parameters.
Increase λ to constrain the model more. Too much regularization can make it underfit; tune against validation data.
04 / A repeatable workflow
Keep validation examples away from gradient updates.
Update weights on training batches only.
Evaluate validation loss after each epoch.
Stop after patience runs out; restore best weights.
05 / Evaluation discipline
Use validation results to choose settings. Touch the test set once, after decisions are final.
Fit parameters.
Choose the recipe.
Report once.
06 / Takeaways
Training performance alone does not measure generalization.
Regularizers constrain different parts of the learning process.
Validation loss guides stopping and model selection.
Keep the test set out of the tuning loop.
Compare a plain MNIST classifier with weight decay, dropout, batch normalization, and early stopping in the accompanying tutorial.