Brian Wong.
← Writing

September 5, 2026

MSE and MAE

Why MSE Predicts the Mean and MAE Predicts the Median

Mean Absolute Error (MAE) and Mean Squared Error (MSE) are two of the most common loss functions for regression.

A common explanation is that MSE is more sensitive to outliers because it squares errors, while MAE is more robust because it uses absolute errors. That is true, but it does not fully explain how these losses affect a model’s predictions.

The deeper point is this:

  • Minimizing MSE leads to predicting a conditional mean.
  • Minimizing MAE leads to predicting a conditional median.

To see why, first consider the simplest case: a model that must make the same constant prediction for every observation.

Mean Squared Error (MSE)

The Mean Squared Error is

MSE⁡=1n∑i=1n(yi−y^i)2,\operatorname{MSE} = \frac{1}{n} \sum_{i=1}^{n} (y_i-\hat y_i)^2,

where:

  • yiy_i is the actual target value.
  • y^i\hat y_i is the predicted value.
  • nn is the number of observations.

Because errors are squared, a large error is penalized much more heavily than a small error. For example, an error of 10 contributes 102=10010^2=100, while an error of 1 contributes only 12=11^2=1.

This quadratic penalty is why MSE is sensitive to outliers.

Why MSE predicts the mean

Suppose the model must output one constant value aa for every target:

y^i=a.\hat y_i=a.

The MSE objective becomes

L(a)=1n∑i=1n(yi−a)2.L(a) = \frac{1}{n} \sum_{i=1}^{n} (y_i-a)^2.

To find the best value of aa, differentiate with respect to aa:

dLda=1n∑i=1n2(a−yi).\frac{dL}{da} = \frac{1}{n} \sum_{i=1}^{n} 2(a-y_i).

Set the derivative equal to zero:

1n∑i=1n2(a−yi)=0.\frac{1}{n} \sum_{i=1}^{n} 2(a-y_i)=0.

Removing the constant factor gives

∑i=1n(a−yi)=0.\sum_{i=1}^{n}(a-y_i)=0.

Therefore,

na−∑i=1nyi=0,na-\sum_{i=1}^{n}y_i=0,

so

a=1n∑i=1nyi.a = \frac{1}{n} \sum_{i=1}^{n}y_i.

Thus, the constant prediction that minimizes MSE is the sample mean.

Because the squared-error loss is convex, this solution is the global minimum.

Mean Absolute Error (MAE)

The Mean Absolute Error is

MAE⁡=1n∑i=1n∣yi−y^i∣.\operatorname{MAE} = \frac{1}{n} \sum_{i=1}^{n} |y_i-\hat y_i|.

Unlike MSE, MAE penalizes errors linearly. Doubling an error doubles its contribution to the loss:

∣10∣=10,∣1∣=1.|10|=10, \qquad |1|=1.

This makes MAE less influenced by extreme values than MSE. An outlier still matters, but its effect does not grow quadratically.

Why MAE predicts the median

Again, suppose the model must output a single constant aa:

L(a)=1n∑i=1n∣yi−a∣.L(a) = \frac{1}{n} \sum_{i=1}^{n} |y_i-a|.

Sort the target values:

y(1)≤y(2)≤⋯≤y(n).y_{(1)} \le y_{(2)} \le \cdots \le y_{(n)}.

For a value of aa that is not exactly equal to one of the targets, the derivative is

dLda=#{yi<a}−#{yi>a}n.\frac{dL}{da} = \frac{ \#\{y_i<a\} - \#\{y_i>a\} }{n}.

This expression has a simple interpretation:

  • Every target below aa contributes a slope of +1+1: increasing aa increases its error.
  • Every target above aa contributes a slope of −1-1: increasing aa decreases its error.

If more observations lie above aa than below it, the derivative is negative. Increasing aa reduces the total absolute error.

If more observations lie below aa than above it, the derivative is positive. Increasing aa increases the total absolute error.

Therefore, the loss is minimized when the number of observations below aa and above aa is balanced. Equivalently,

#{yi≤a}≥n2and#{yi≥a}≥n2.\#\{y_i\le a\}\ge \frac{n}{2} \quad\text{and}\quad \#\{y_i\ge a\}\ge \frac{n}{2}.

That is precisely the definition of a median.

Since absolute-error loss is convex, any such value is a global minimizer.

Example: an outlier changes MSE much more

Consider the targets

[1,2,4,100].[1,2,4,100].

The mean is

1+2+4+1004=26.75.\frac{1+2+4+100}{4}=26.75.

The median is not unique because there are four values. Every value in the interval

[2,4][2,4]

minimizes MAE.

Constant prediction aaSum of absolute errors
110+1+3+99=1030+1+3+99=103
221+0+2+98=1011+0+2+98=101
332+1+1+97=1012+1+1+97=101
443+2+0+96=1013+2+0+96=101
10010099+98+96+0=29399+98+96+0=293
Mean =26.75=26.7525.75+24.75+22.75+73.25=146.525.75+24.75+22.75+73.25=146.5

The MAE-optimal constant prediction is any value from 22 to 44. The value 100100 contributes a large absolute error regardless of whether the prediction is 2, 3, or 4, so it does not pull the optimum toward itself.

By contrast, under MSE, the error from 100100 is squared. Predicting too far below 100 becomes extremely expensive, so the MSE-optimal prediction shifts upward toward the mean, 26.7526.75.

The more general result

In a real regression problem, the model does not usually produce one global constant. It produces a prediction based on input features X=xX=x.

The population-optimal prediction under squared loss is

fMSE(x)=E[Y∣X=x].f_{\text{MSE}}(x) = \mathbb{E}[Y\mid X=x].

In words: a sufficiently flexible model trained to minimize expected MSE estimates the conditional mean of YY given X=xX=x.

Under absolute loss, the population-optimal prediction is a conditional median:

fMAE(x)∈Median⁡(Y∣X=x).f_{\text{MAE}}(x) \in \operatorname{Median}(Y\mid X=x).

In words: a model trained to minimize expected MAE estimates the conditional median of YY given X=xX=x.

This distinction matters. The statement “MSE predicts the mean” does not mean that every MSE regression model outputs the overall average target value. Rather, it means that for each feature pattern xx, the ideal prediction is the average outcome among observations with similar features.

Practical takeaway

Choose the loss according to the decision you want the model to make:

  • Use MSE when large mistakes should receive disproportionately large penalties, or when predicting the conditional average is the right goal.
  • Use MAE when a typical or central outcome is more useful and you want less sensitivity to unusually large errors.
  • Use a different loss if overprediction and underprediction have different real-world costs. For example, quantile loss can target upper or lower conditional quantiles rather than the mean or median.

The difference between MSE and MAE is therefore not only about outlier robustness. Each loss function encodes a different definition of the “best” prediction.