#1: Mean Absolute Error (MAE)
📖 Definition
Mean Absolute Error (MAE) measures the average magnitude of errors in predictions, without considering their direction. It represents the average distance between predicted values and actual values.
🧮 Mathematical Equation & Components
$$\text{MAE} = \frac{1}{N} \sum_{i=1}^{N} \left| y_i - \hat{y}_i \right|$$
- N = number of data points
- yi = actual value
- ŷi = predicted value
🧰 Typical Use Cases
- Regression tasks where interpretability and robustness to outliers are valued.
- Situations where equal penalty is assigned to all deviations.
- Applications like retail forecasting, cost estimation, energy usage prediction.
⭐ Strength Points
- Easy to interpret – in the same unit as the target.
- Robust to outliers compared to MSE or RMSE.
- Good when large errors should not be exaggerated.
⚠️ Weakness Points
- Doesn’t indicate error direction – over vs under prediction.
- Less smooth for gradient-based optimizers than MSE.
- Underpenalizes large deviations if they matter to stakeholders.
🎯 Best Practice Recommendation
- Use MAE when outliers are not critical and you need clear interpretability.
- Combine with RMSE or MSE to assess penalization of large errors.
- Ideal for communicating average error in applied scenarios.
#2: Mean Squared Error (MSE)
📖 Definition
Mean Squared Error (MSE) measures the average of the squares of the errors—the squared difference between predicted and actual values. It emphasizes larger errors by squaring them, making it particularly sensitive to significant prediction mistakes.
🧮 Mathematical Equation & Components
$$\text{MSE} = \frac{1}{N} \sum_{i=1}^{N} \left( y_i - \hat{y}_i \right)^2$$
- N = number of observations
- yi = actual value
- ŷi = predicted value
🧰 Typical Use Cases
- Used as the default loss function in many regression models.
- Appropriate when **large errors must be penalized more heavily**.
- Applications in forecasting, risk modeling, and finance.
⭐ Strength Points
- Differentiable and convex — ideal for optimization algorithms.
- Penalizes large deviations, so well-suited for high-risk environments.
- Standard benchmark in model competitions and research papers.
⚠️ Weakness Points
- Highly sensitive to outliers — one large error can skew the metric.
- Not directly interpretable — units are squared.
- Can mislead when data has high variance or noise.
🎯 Best Practice Recommendation
- Use when outliers should be heavily penalized.
- Pair with
MAEto assess robustness and interpretability. - Use
RMSEto make results easier to interpret in target units.
#3: Root Mean Squared Error (RMSE)
📖 Definition
Root Mean Squared Error (RMSE) is the square root of Mean Squared Error (MSE). It preserves MSE’s sensitivity to larger errors while restoring the error unit to match the original target variable—making it more interpretable.
🧮 Mathematical Equation & Components
$$\text{RMSE} = \sqrt{\frac{1}{N} \sum_{i=1}^{N} \left( y_i - \hat{y}_i \right)^2}$$
- N = number of predictions
- yi = actual value
- ŷi = predicted value
🧰 Typical Use Cases
- Forecasting models where interpretability in the original scale matters
- Regression tasks with concern for large error penalties
- Common in energy models, physics, and time series evaluation
⭐ Strength Points
- Results are in the same units as the target — highly interpretable
- Retains MSE's sensitivity to large deviations
- Standard metric for benchmarking and competitions
⚠️ Weakness Points
- Still sensitive to outliers
- May
if not visualized properly - Not robust for noisy or skewed distributions
🎯 Best Practice Recommendation
- Use RMSE when interpretability and **error magnitude** matter
- Compare with
MAEfor deeper insight:- If
RMSE ≫ MAE, your model may be struggling with outliers
- If
- Use normalized RMSE (e.g.,
RMSE / mean(target)) to compare across tasks
#4: R-squared (Coefficient of Determination)
📖 Definition
R-squared (R²) measures the proportion of variance in the target variable that can be explained by the input features. It reflects how well your model fits the data.
🧮 Mathematical Equation & Components
$$R^2 = 1 - \frac{\sum (y_i - \hat{y}_i)^2}{\sum (y_i - \bar{y})^2}$$
- yi = actual value
- ŷi = predicted value
- ȳ = mean of actual values
- Numerator = Residual Sum of Squares (RSS)
- Denominator = Total Sum of Squares (TSS)
🧰 Typical Use Cases
- Evaluating model fit in linear regression
- Exploratory data analysis to gauge explainable variance
- Benchmarking models against simple baselines
⭐ Strength Points
- Intuitive: expresses result as a percentage of variance explained
- Normalized scale (0 to 1) — with possible negatives for poor fits
- Useful for comparing model performance on the same dataset
⚠️ Weakness Points
- Can be misleading for non-linear models or overfitted models
- Does not penalize complexity — can overestimate poor fits
- Doesn’t describe the magnitude of error
🎯 Best Practice Recommendation
- Use R² as a quick snapshot of model fit
- Prefer
Adjusted R²when comparing models with different numbers of features - Combine with
MAEorRMSEfor error scale analysis
#5: Adjusted R-squared
📖 Definition
Adjusted R-squared modifies the regular R² score by penalizing the inclusion of unnecessary predictors. It improves upon R² by factoring in model complexity, helping avoid overfitting.
🧮 Mathematical Equation & Components
$$\text{Adjusted } R^2 = 1 - \left( \frac{(1 - R^2)(n - 1)}{n - p - 1} \right)$$
- R² = standard coefficient of determination
- n = number of observations
- p = number of independent variables (predictors)
🧰 Typical Use Cases
- Model comparison with varying numbers of features
- Linear regression diagnostics to avoid overfitting
- Feature selection and pruning
⭐ Strength Points
- Penalizes models for adding non-informative variables
- More realistic than R² when many predictors are involved
- Favors parsimonious, generalizable models
⚠️ Weakness Points
- Still assumes a linear model structure
- Less meaningful for tree-based or nonparametric models
- Not intuitive for non-technical audiences
🎯 Best Practice Recommendation
- Use when comparing multiple linear models with different numbers of features
- Helpful in feature engineering for identifying relevant predictors
- Always evaluate alongside
MAE,RMSE, orcross-validation metrics
#6: Mean Absolute Percentage Error (MAPE)
📖 Definition
MAPE measures prediction error as a percentage of the actual value. It averages the absolute percent difference between predicted and actual values across all predictions— making it ideal for understanding relative error.
🧮 Mathematical Equation & Components
$$\text{MAPE} = \frac{100\%}{N} \sum_{i=1}^{N} \left| \frac{y_i - \hat{y}_i}{y_i} \right|$$
- $y_i$ = actual value
- $\hat{y}_i$ = predicted value
- $N$ = number of observations
🧰 Typical Use Cases
- Business forecasting (e.g., sales, revenue)
- Time series analysis with large scale differences
- Model results reporting for non-technical audiences
⭐ Strength Points
- Easy to interpret as percentage error
- Helps with scale-independent evaluation
- Enables model comparison across domains
⚠️ Weakness Points
- Undefined or unstable when actual value ($y_i$) is zero or near-zero
- Can be biased toward under-predictions
- Penalizes large over-predictions more than large under-predictions
🎯 Best Practice Recommendation
- Use when relative accuracy is more meaningful than absolute
- Avoid when data includes zeros or near-zeros—use
sMAPEinstead - Always report with
MAEorRMSEto contextualize absolute performance
#7: Symmetric Mean Absolute Percentage Error (sMAPE)
📖 Definition
sMAPE is an adjusted version of MAPE that provides a balanced view of relative error, treating over- and under-predictions symmetrically. It's especially useful when dealing with values near zero or when MAPE’s asymmetry introduces bias.
🧮 Mathematical Equation & Components
$$\text{sMAPE} = \frac{100\%}{N} \sum_{i=1}^{N} \frac{|\hat{y}_i - y_i|}{(|y_i| + |\hat{y}_i|)/2}$$
- $y_i$ = actual value
- $\hat{y}_i$ = predicted value
- $N$ = number of data points
🧰 Typical Use Cases
- Forecasting in sectors like retail, energy, traffic
- Performance metrics for **zero-inclusive data**
- Dashboard-ready metrics for **executive reporting**
⭐ Strength Points
- Better than MAPE at handling **zero or near-zero values**
- **Balanced treatment** of over- vs. under-estimation
- Great for **comparative analysis across multiple series**
⚠️ Weakness Points
- Still unstable if both prediction and actual ≈ 0
- Can be **less intuitive** to explain than plain MAPE
- Not natively available in some libraries like
scikit-learn
🎯 Best Practice Recommendation
- Prefer sMAPE over MAPE when your data contains zeros or small values
- Use it in forecasting tasks where **balanced relative error** is important
- Always cross-check with
MAEorRMSEfor absolute error magnitude
#8: Huber Loss
📖 Definition
Huber Loss is a hybrid metric that blends the best of Mean Squared Error (MSE) and
Mean Absolute Error (MAE). It behaves like MSE for small errors (ensuring smooth optimization)
and like MAE for large errors (reducing sensitivity to outliers).
🧮 Mathematical Equation & Components
$$ L_\delta(y, \hat{y}) = \begin{cases} \frac{1}{2}(y - \hat{y})^2 & \text{if } |y - \hat{y}| \leq \delta \\ \delta \cdot (|y - \hat{y}| - \frac{1}{2} \delta) & \text{otherwise} \end{cases} $$
- $y$ = actual value
- $\hat{y}$ = predicted value
- $\delta$ = threshold that determines the switch between MSE and MAE behavior
🧰 Typical Use Cases
- When data contains **outliers or noisy observations**
- **Robust regression** settings and **gradient boosting frameworks**
- Used in **deep learning loss functions** and **industrial predictive systems**
⭐ Strength Points
- **Balances robustness and smooth optimization**
- Reduces the impact of outliers without ignoring them entirely
- Widely supported in libraries like
XGBoost,TensorFlow,PyTorch
⚠️ Weakness Points
- Requires selection of **$\delta$**, which affects model sensitivity
- Slightly more **complex to compute** and tune compared to MAE or MSE
- Performance depends on **good hyperparameter tuning**
🎯 Best Practice Recommendation
- Use Huber Loss when your data has **moderate outliers** or noisy labels
- Perform **cross-validation** to tune the
δparameter effectively - Combine with MAE/MSE evaluations for comprehensive analysis
#9: Explained Variance Score
📖 Definition
The Explained Variance Score quantifies how much of the variance in the target variable is
captured by the model. It’s closely related to R² but focuses more on variance reduction
rather than error relative to mean.
🧮 Mathematical Equation & Components
$$ \text{Explained Variance} = 1 - \frac{\text{Var}(y - \hat{y})}{\text{Var}(y)} $$
- $y$: true values
- $\hat{y}$: predicted values
- Var($y - \hat{y}$): variance of prediction errors (residuals)
- Var($y$): variance of actual target values
🧰 Typical Use Cases
- Evaluating how well a model captures target variability
- Continuous regression tasks where variance matters more than absolute accuracy
- Time series, energy forecasting, or environmental modeling
⭐ Strength Points
- Highlights the model’s ability to reduce unexplained variation
- Less impacted by constant shift/bias than R²
- Helps distinguish models with similar RMSE but different variance coverage
⚠️ Weakness Points
- Can be negative if the model increases variance over baseline
- May be misleading if model has high bias (systematic error)
- Not as intuitive as RMSE or R² for non-technical audiences
🎯 Best Practice Recommendation
- Use in tandem with
R²andRMSEfor a more complete regression assessment - Especially valuable when variance containment is the modeling priority
- Avoid comparing across datasets with drastically different variance scales
#9: Explained Variance Score
📖 Definition
The Explained Variance Score quantifies how much of the variance in the target variable is
captured by the model. It’s closely related to R² but focuses more on variance reduction
rather than error relative to mean.
🧮 Mathematical Equation & Components
$$ \text{Explained Variance} = 1 - \frac{\text{Var}(y - \hat{y})}{\text{Var}(y)} $$
- $y$: true values
- $\hat{y}$: predicted values
- Var($y - \hat{y}$): variance of prediction errors (residuals)
- Var($y$): variance of actual target values
🧰 Typical Use Cases
- Evaluating how well a model captures target variability
- Continuous regression tasks where variance matters more than absolute accuracy
- Time series, energy forecasting, or environmental modeling
⭐ Strength Points
- Highlights the model’s ability to reduce unexplained variation
- Less impacted by constant shift/bias than R²
- Helps distinguish models with similar RMSE but different variance coverage
⚠️ Weakness Points
- Can be negative if the model increases variance over baseline
- May be misleading if model has high bias (systematic error)
- Not as intuitive as RMSE or R² for non-technical audiences
🎯 Best Practice Recommendation
- Use in tandem with
R²andRMSEfor a more complete regression assessment - Especially valuable when variance containment is the modeling priority
- Avoid comparing across datasets with drastically different variance scales
#10: Mean Bias Deviation (MBD)
📖 Definition
Mean Bias Deviation (MBD) measures the average directional error of your model’s predictions. Unlike MAE or RMSE, which focus on magnitude, MBD reveals whether your model consistently over- or under-predicts.
🧮 Mathematical Equation & Components
$$ \text{MBD} = \frac{1}{N} \sum_{i=1}^{N} ( \hat{y}_i - y_i ) $$
- $y_i$: actual value
- $\hat{y}_i$: predicted value
- $N$: total number of samples
🧰 Typical Use Cases
- Detecting systematic prediction bias in regression models
- Evaluating and tuning model calibration
- Critical in energy forecasting, finance, climate models
⭐ Strength Points
- Identifies whether model is biased toward over- or under-predicting
- Easy to interpret and useful for model diagnostics
- Helpful for bias correction workflows
⚠️ Weakness Points
- Positive and negative errors cancel out, possibly hiding large errors
- Should not be used in isolation—combine with
MAEorRMSE - Sensitive to outliers
🎯 Best Practice Recommendation
- Use MBD to detect
not captured by other metrics - Pair with
MAE,RMSE, orR²to complete the picture - Visualize MBD over time or by feature to identify model drift
📊 Comparative Atlas of Regression Metrics
🧭 1. Key Trade-offs Between Metrics
| Metric | Measures | Sensitive to Outliers | Easy to Interpret | Directional Bias | Penalizes Large Errors |
|---|---|---|---|---|---|
MAE |
Avg. magnitude of error | ❌ Low | ✅ Yes | ❌ No | ❌ No |
MSE |
Avg. squared error | ✅ High | ⚠️ Somewhat | ❌ No | ✅ Yes |
RMSE |
Root of squared error | ✅ High | ✅ Yes | ❌ No | ✅ Yes |
R² |
Variance explained | ⚠️ Depends | ✅ Yes | ❌ No | ❌ No |
Adjusted R² |
Adjusted variance | ⚠️ Depends | ⚠️ Somewhat | ❌ No | ❌ No |
MAPE |
% error | ✅ High (near 0) | ✅ Yes | ❌ No | ❌ No |
sMAPE |
Symmetric % error | ⚠️ Medium | ✅ Yes | ❌ No | ❌ No |
Huber |
Combined MAE/MSE | ⚠️ Medium | ⚠️ Model-specific | ❌ No | ✅ Yes (conditionally) |
Explained Variance |
Variance reduced | ❌ Low | ⚠️ Somewhat | ❌ No | ❌ No |
MBD |
Avg. signed bias | ⚠️ Somewhat | ⚠️ Model-specific | ✅ Yes | ❌ No |
🎯 2. Use Case Selection Table
| Use Case | Recommended Metric(s) |
|---|---|
| General regression | RMSE, MAE, R² |
| Outlier-resistant models | MAE, Huber, sMAPE |
| Penalize large deviations | MSE, RMSE |
| Compare models of different sizes | Adjusted R² |
| Forecast error as % | MAPE, sMAPE |
| Explainability to stakeholders | MAE, RMSE, MAPE |
| Calibration check (bias) | MBD |
| Variance-focused analysis | Explained Variance, R² |
| Robust loss function for training | Huber Loss |
🧠 3. Visual Mapping Guide (What to Plot)
| Metric | Suggested Plot | What It Shows |
|---|---|---|
MAE, RMSE |
Residual plots, error histograms | Spread and magnitude of errors |
MAPE / sMAPE |
Line plot of percentage errors | Relative forecasting performance |
R² / Adjusted R² |
Model fit plots (predicted vs actual) | Fit and dispersion pattern |
MBD |
Residual mean line, bias plots | Directional drift |
Explained Variance |
Target vs prediction variance bar chart | Model's explanatory power |
🧪 4. Statistical Behavior & Optimization Suitability
| Metric | Convex | Differentiable | Used in Loss Functions |
|---|---|---|---|
MAE |
✅ | ❌ (non-smooth) | ⚠️ Sometimes |
MSE |
✅ | ✅ | ✅ Standard |
RMSE |
✅ | ✅ | ✅ (rarely directly) |
Huber |
✅ | ✅ | ✅ Robust loss |
📦 5. Ensemble Use & Reporting
Best practice: report multiple metrics together to capture different dimensions of model performance:
- MAE + RMSE → General error + sensitivity to large deviations
- MAPE or sMAPE → Percentage-based errors, ideal for stakeholder-facing reports
- R² + Adjusted R² → Model fit + complexity-aware comparison
- MBD + Explained Variance → Bias + variance interpretation
- Huber Loss → Use during training for robustness to outliers
📋 Regression Metrics Selection Checklist
MAE– easy to explainRMSE– same units, emphasizes large errorsMAPE/sMAPE– percent error
MAE– robustHuber Loss– adaptivesMAPE– if percentage error needed- No outliers? → Use
MSE,RMSE, orMAPE
MSE– quadratic penaltyRMSE– amplifies large errorsHuber Loss– combines linear and quadratic
MBD– Mean Bias Deviation- Check
residual plots - Consider
Explained Variance Score
MAPE/sMAPE– for relative errorMAEandMBD– for interpretability and drift detection
Adjusted R²- Also track
R²andExplained Variance
MSE– standard lossHuber Loss– robust training loss
MAPE/sMAPE– scale-invariant- Normalize
MAE/RMSEby mean target value if needed