All lessons in Machine LearningTất cả bài trong chương Machine Learning19

What Does It Mean for a Machine to Learn?

A clean model of supervised learning: data, hypothesis, objective, optimization and generalization.

August 19, 2026
On this pageMục lục bài viết7

Machine learning is often introduced as "systems that learn from data." That is true but too vague to guide engineering decisions.

A more useful decomposition is:

data
  ↓
model family
  ↓
objective
  ↓
optimization
  ↓
selected parameters
  ↓
generalization to unseen data

The model family defines what can be learned

Before training, you choose a family of possible functions: linear models, decision trees, neural networks and so on.

Training does not search every imaginable program. It searches within the structure you provide.

This is an inductive bias. A linear model prefers linear relationships. A convolutional network encodes assumptions about local spatial structure. Architecture is therefore part of the learning problem.

Data provides evidence

Training examples tell the model which patterns should be useful. But datasets also contain noise, shortcuts and historical bias.

The model does not automatically know which correlation you intended it to learn.

If a background pattern predicts the label more easily than the actual object, a vision model may use the background. If a timestamp leaks the target, a tabular model may exploit it.

The objective says what "better" means

Training requires an objective that compares behavior to a desired target.

For regression, mean squared error may be reasonable. For classification, cross-entropy is common. Ranking and generative tasks need other objectives.

The objective is not the product requirement. It is a differentiable or optimizable proxy for it.

Optimization chooses parameters

The optimizer adjusts parameters to reduce the training objective. This is the part people often call "learning," but optimization alone is not enough.

A model that memorizes training examples can optimize the loss very well and still be useless.

Generalization is the real test

The valuable behavior is performance on data the model did not directly optimize against.

That is why train/validation/test separation matters. The test set acts as a proxy for future unseen cases.

A useful model discovers structure that transfers beyond the exact training examples.

Learning is a system, not a model class

When performance is poor, the fix may not be "use a bigger model." The problem may be:

  • the target is badly defined;
  • the input lacks necessary information;
  • labels are inconsistent;
  • train and production distributions differ;
  • evaluation measures the wrong behavior;
  • the model family cannot express the relationship.

The model is one component in a learning system.

A reusable debugging frame

When a machine-learning project struggles, inspect five layers:

  1. Task - what behavior is actually required?
  2. Data - what evidence is available and how was it sampled?
  3. Model - what patterns can the architecture represent?
  4. Objective - what behavior does training reward?
  5. Evaluation - how do we know the result generalizes?

That frame is more durable than memorizing a list of algorithms.