All lessons in FoundationsTất cả bài trong chương Foundations20
Gradient Descent: How Models Turn Error Into an Update
Understand loss surfaces, gradients, learning rates and the optimization loop that trains most neural models.
On this pageMục lục bài viết8
Training a model means finding parameters that make its predictions useful. Gradient descent is the basic mechanism that turns how wrong the model is into how its parameters should change.
Start with an objective
A model produces predictions. A loss function converts the difference between prediction and target into a number the optimizer can minimize.
parameters
↓
model → prediction
↓
loss
↓
gradient
↓
parameter update
The loss does not need to perfectly represent the real product goal, but it needs to produce a training signal aligned enough with that goal.
The gradient points uphill
For a parameter, the derivative tells us how the loss changes when that parameter changes slightly. Across many parameters, the gradient collects those local slopes into a vector.
The gradient points toward increasing loss, so gradient descent moves in the opposite direction:
parameter = parameter - learning_rate * gradient
That single pattern scales from a line with two parameters to neural networks with billions.
Learning rate controls step size
A very small learning rate makes progress slow. A very large one can overshoot useful regions or make training unstable.
small step: · · · · · → minimum
large step: ↗ ↘ may bounce around
Modern optimizers adapt or reshape updates, but the learning-rate intuition remains important.
The loss surface is high-dimensional
Visual diagrams show a ball rolling down a two-dimensional bowl. Real neural networks have enormous parameter spaces with saddles, flat regions and complicated curvature.
The picture is still useful as an intuition, but do not take it literally. Training succeeds because local gradient information, architecture, initialization, normalization, optimizer design and data all interact.
Batches make the gradient noisy
Computing a gradient over the full dataset can be expensive. Mini-batch training estimates the gradient from a subset of examples.
This introduces noise, but the noise is not necessarily harmful. It can make optimization computationally practical and sometimes helps explore the parameter space.
Backpropagation is efficient bookkeeping
A neural network is a composition of functions. Backpropagation applies the chain rule efficiently so each parameter receives its contribution to the final loss.
You rarely calculate these derivatives manually in production. Automatic differentiation does that work. But knowing what is being propagated prevents the training loop from feeling magical.
Optimization and generalization are different questions
A model can achieve very low training loss and still perform poorly on unseen data.
Optimization asks:
Can we fit the training objective?
Generalization asks:
Did we learn a pattern that survives outside the training set?
Do not confuse a successful optimizer with a successful model.
Keep this loop in your head
Whenever a model is trained, identify:
- the parameters;
- the forward computation;
- the loss;
- the gradient path;
- the optimizer update;
- the data used for that update.
Most training systems are larger versions of this loop with better engineering around it.