All lessons in FoundationsTất cả bài trong chương Foundations20

Statistics for AI: What Can You Conclude From Data?

A practical introduction to sampling, estimation, variance, confidence and why model evaluation is a statistical problem.

August 19, 2026
On this pageMục lục bài viết7

Probability asks what outcomes are possible under a model. Statistics often works in the opposite direction: given observed data, what can we reasonably infer about the underlying process?

That distinction becomes extremely practical when evaluating AI systems.

A dataset is a sample, not the world

Suppose an evaluation set contains 1,000 queries. Even if the test set is carefully designed, it is still a sample from a much larger space of possible user behavior.

The measured accuracy is therefore an estimate, not an eternal property of the model.

A change from 84.1% to 84.7% may reflect a real improvement, noise, or a shift concentrated in one slice. Statistics helps you avoid treating every decimal movement as meaningful.

Mean alone is not enough

Averages compress information. For latency, an average can look healthy while a minority of requests are painfully slow. For model quality, overall accuracy can look good while one critical intent fails repeatedly.

Useful summaries include:

  • mean and median;
  • quantiles such as p90 or p95;
  • variance or standard deviation;
  • distributions by slice;
  • counts, not only percentages.

Sampling changes conclusions

If your evaluation set over-represents easy cases, the model appears stronger than it is in production. If it over-represents rare failures, the opposite happens.

This is why dataset construction is part of evaluation design.

Ask:

What population do I care about?
        ↓
How was this sample collected?
        ↓
Which behaviors are over/under-represented?

Correlation is not causation

Two variables can move together without one causing the other. A model may learn a proxy feature that correlates with the label in training data but breaks when the environment changes.

This is one reason shortcuts and dataset leakage are dangerous: the model can score well by exploiting a correlation that does not represent the intended task.

Confidence is about repeated sampling

A confidence interval is not simply "the model is probably inside this range." It describes the behavior of an estimation procedure under repeated samples.

The deeper lesson for engineering is simpler: measurements have uncertainty.

For small slices, one or two examples can move the metric dramatically. Always inspect sample size before comparing percentages.

Statistical thinking for model evaluation

When a new model beats the old one, ask:

  1. On how many cases?
  2. Which cases changed?
  3. Are improvements concentrated in one slice?
  4. Did any high-value behavior regress?
  5. Is the comparison paired on the same examples?
  6. Could annotation noise explain the difference?

That mindset prevents evaluation from becoming scoreboard watching.

The practical goal

Statistics for AI is less about memorizing tests and more about learning to distrust conclusions that ignore sampling, uncertainty and distribution shift.

A metric is evidence. It is not the whole explanation.