All lessons in FoundationsTất cả bài trong chương Foundations20

Probability for AI: Modeling Uncertainty Instead of Hiding It

The probability concepts that matter most for machine learning: random variables, conditional probability, expectation and distributions.

August 19, 2026
On this pageMục lục bài viết7

AI systems constantly operate under uncertainty. A classifier does not know the true class before observing data. A language model does not know the next token. A detector does not know whether a region contains an object.

Probability gives us a language for representing that uncertainty instead of pretending every prediction is certain.

Random variables describe uncertain outcomes

A random variable assigns numerical values to possible outcomes. For a binary classifier, a simple variable could be:

Y = 1  if the email is spam
Y = 0  otherwise

Before observing all relevant evidence, Y is uncertain. A model may estimate:

P(Y = 1 | x) = 0.87

The vertical bar means given. The model is not claiming that Y equals 0.87; it is estimating a probability conditioned on input x.

Conditional probability is everywhere

Most supervised learning can be viewed as learning something about a conditional distribution:

input x  →  model  →  distribution over y

For classification, we often want P(y | x). In generative modeling, we may model how likely data is under some conditions. In an autoregressive language model, the next token is predicted conditioned on previous tokens.

The key idea is that probability changes when information changes.

Expectation is a weighted average over possibilities

The expected value summarizes the average outcome under a probability distribution.

If a system has outcomes with different probabilities and costs, expected value provides a principled way to reason about the trade-off.

This becomes important in losses, calibration and decision theory. A probability estimate can be technically accurate while the downstream decision is still poor if the business cost of false positives and false negatives is asymmetric.

A distribution contains more than a mean

Two datasets can have the same average and behave very differently.

A: tightly concentrated near the mean
B: widely spread with extreme values

Variance describes spread. Distribution shape can reveal skew, heavy tails or multiple modes.

In AI engineering, this matters because aggregate averages often hide the slice where the system fails.

Bayes' rule is about updating beliefs

Bayes' rule formalizes how prior beliefs change after seeing evidence. The practical intuition matters more than memorizing the formula:

prior belief
    + evidence
        ↓
updated belief

This pattern appears in probabilistic modeling, diagnosis, filtering and any problem where evidence arrives over time.

Probability is not confidence theater

A model output of 0.95 is useful only if that number has a meaningful relationship to real-world frequency. If predictions assigned 95% confidence are correct only 70% of the time, the model is poorly calibrated.

So distinguish:

  • score: an arbitrary model output used for ranking;
  • probability: a number intended to represent uncertainty;
  • calibration: whether predicted probabilities match observed outcomes.

What to carry forward

You do not need to solve every probability exercise before building models. You should be comfortable with:

  • conditional probability;
  • random variables;
  • expectation and variance;
  • common distributions;
  • independence versus dependence;
  • the idea of updating beliefs with evidence.

These ideas reappear in classification, loss functions, Bayesian methods, language modeling and evaluation.