Mean Absolute Error (MAE) and Mean Squared Error (MSE) are two of the most common loss functions for regression.
A common explanation is that MSE is more sensitive to outliers because it squares errors, while MAE is more robust because it uses absolute errors. That is true, but it does not fully explain how these losses affect a model’s predictions.
The deeper point is this:
- Minimizing MSE leads to predicting a conditional mean.
- Minimizing MAE leads to predicting a conditional median.
To see why, first consider the simplest case: a model that must make the same constant prediction for every observation.
Mean Squared Error (MSE)
The Mean Squared Error is
where:
- is the actual target value.
- is the predicted value.
- is the number of observations.
Because errors are squared, a large error is penalized much more heavily than a small error. For example, an error of 10 contributes , while an error of 1 contributes only .
This quadratic penalty is why MSE is sensitive to outliers.
Why MSE predicts the mean
Suppose the model must output one constant value for every target:
The MSE objective becomes
To find the best value of , differentiate with respect to :
Set the derivative equal to zero:
Removing the constant factor gives
Therefore,
so
Thus, the constant prediction that minimizes MSE is the sample mean.
Because the squared-error loss is convex, this solution is the global minimum.
Mean Absolute Error (MAE)
The Mean Absolute Error is
Unlike MSE, MAE penalizes errors linearly. Doubling an error doubles its contribution to the loss:
This makes MAE less influenced by extreme values than MSE. An outlier still matters, but its effect does not grow quadratically.
Why MAE predicts the median
Again, suppose the model must output a single constant :
Sort the target values:
For a value of that is not exactly equal to one of the targets, the derivative is
This expression has a simple interpretation:
- Every target below contributes a slope of : increasing increases its error.
- Every target above contributes a slope of : increasing decreases its error.
If more observations lie above than below it, the derivative is negative. Increasing reduces the total absolute error.
If more observations lie below than above it, the derivative is positive. Increasing increases the total absolute error.
Therefore, the loss is minimized when the number of observations below and above is balanced. Equivalently,
That is precisely the definition of a median.
Since absolute-error loss is convex, any such value is a global minimizer.
Example: an outlier changes MSE much more
Consider the targets
The mean is
The median is not unique because there are four values. Every value in the interval
minimizes MAE.
| Constant prediction | Sum of absolute errors |
|---|---|
| Mean |
The MAE-optimal constant prediction is any value from to . The value contributes a large absolute error regardless of whether the prediction is 2, 3, or 4, so it does not pull the optimum toward itself.
By contrast, under MSE, the error from is squared. Predicting too far below 100 becomes extremely expensive, so the MSE-optimal prediction shifts upward toward the mean, .
The more general result
In a real regression problem, the model does not usually produce one global constant. It produces a prediction based on input features .
The population-optimal prediction under squared loss is
In words: a sufficiently flexible model trained to minimize expected MSE estimates the conditional mean of given .
Under absolute loss, the population-optimal prediction is a conditional median:
In words: a model trained to minimize expected MAE estimates the conditional median of given .
This distinction matters. The statement “MSE predicts the mean” does not mean that every MSE regression model outputs the overall average target value. Rather, it means that for each feature pattern , the ideal prediction is the average outcome among observations with similar features.
Practical takeaway
Choose the loss according to the decision you want the model to make:
- Use MSE when large mistakes should receive disproportionately large penalties, or when predicting the conditional average is the right goal.
- Use MAE when a typical or central outcome is more useful and you want less sensitivity to unusually large errors.
- Use a different loss if overprediction and underprediction have different real-world costs. For example, quantile loss can target upper or lower conditional quantiles rather than the mean or median.
The difference between MSE and MAE is therefore not only about outlier robustness. Each loss function encodes a different definition of the “best” prediction.