#1: Mean Absolute Error (MAE)

📖 Definition

Mean Absolute Error (MAE) measures the average magnitude of errors in predictions, without considering their direction. It represents the average distance between predicted values and actual values.

🧮 Mathematical Equation & Components

$$\text{MAE} = \frac{1}{N} \sum_{i=1}^{N} \left| y_i - \hat{y}_i \right|$$

  • N = number of data points
  • yi = actual value
  • ŷi = predicted value

🧰 Typical Use Cases

  • Regression tasks where interpretability and robustness to outliers are valued.
  • Situations where equal penalty is assigned to all deviations.
  • Applications like retail forecasting, cost estimation, energy usage prediction.

⭐ Strength Points

  • Easy to interpret – in the same unit as the target.
  • Robust to outliers compared to MSE or RMSE.
  • Good when large errors should not be exaggerated.

⚠️ Weakness Points

  • Doesn’t indicate error direction – over vs under prediction.
  • Less smooth for gradient-based optimizers than MSE.
  • Underpenalizes large deviations if they matter to stakeholders.

🎯 Best Practice Recommendation

  • Use MAE when outliers are not critical and you need clear interpretability.
  • Combine with RMSE or MSE to assess penalization of large errors.
  • Ideal for communicating average error in applied scenarios.

#2: Mean Squared Error (MSE)

📖 Definition

Mean Squared Error (MSE) measures the average of the squares of the errors—the squared difference between predicted and actual values. It emphasizes larger errors by squaring them, making it particularly sensitive to significant prediction mistakes.

🧮 Mathematical Equation & Components

$$\text{MSE} = \frac{1}{N} \sum_{i=1}^{N} \left( y_i - \hat{y}_i \right)^2$$

  • N = number of observations
  • yi = actual value
  • ŷi = predicted value

🧰 Typical Use Cases

  • Used as the default loss function in many regression models.
  • Appropriate when **large errors must be penalized more heavily**.
  • Applications in forecasting, risk modeling, and finance.

⭐ Strength Points

  • Differentiable and convex — ideal for optimization algorithms.
  • Penalizes large deviations, so well-suited for high-risk environments.
  • Standard benchmark in model competitions and research papers.

⚠️ Weakness Points

  • Highly sensitive to outliers — one large error can skew the metric.
  • Not directly interpretable — units are squared.
  • Can mislead when data has high variance or noise.

🎯 Best Practice Recommendation

  • Use when outliers should be heavily penalized.
  • Pair with MAE to assess robustness and interpretability.
  • Use RMSE to make results easier to interpret in target units.

#3: Root Mean Squared Error (RMSE)

📖 Definition

Root Mean Squared Error (RMSE) is the square root of Mean Squared Error (MSE). It preserves MSE’s sensitivity to larger errors while restoring the error unit to match the original target variable—making it more interpretable.

🧮 Mathematical Equation & Components

$$\text{RMSE} = \sqrt{\frac{1}{N} \sum_{i=1}^{N} \left( y_i - \hat{y}_i \right)^2}$$

  • N = number of predictions
  • yi = actual value
  • ŷi = predicted value

🧰 Typical Use Cases

  • Forecasting models where interpretability in the original scale matters
  • Regression tasks with concern for large error penalties
  • Common in energy models, physics, and time series evaluation

⭐ Strength Points

  • Results are in the same units as the target — highly interpretable
  • Retains MSE's sensitivity to large deviations
  • Standard metric for benchmarking and competitions

⚠️ Weakness Points

  • Still sensitive to outliers
  • May if not visualized properly
  • Not robust for noisy or skewed distributions

🎯 Best Practice Recommendation

  • Use RMSE when interpretability and **error magnitude** matter
  • Compare with MAE for deeper insight:
    • If RMSE ≫ MAE, your model may be struggling with outliers
  • Use normalized RMSE (e.g., RMSE / mean(target)) to compare across tasks

#4: R-squared (Coefficient of Determination)

📖 Definition

R-squared (R²) measures the proportion of variance in the target variable that can be explained by the input features. It reflects how well your model fits the data.

🧮 Mathematical Equation & Components

$$R^2 = 1 - \frac{\sum (y_i - \hat{y}_i)^2}{\sum (y_i - \bar{y})^2}$$

  • yi = actual value
  • ŷi = predicted value
  • ȳ = mean of actual values
  • Numerator = Residual Sum of Squares (RSS)
  • Denominator = Total Sum of Squares (TSS)

🧰 Typical Use Cases

  • Evaluating model fit in linear regression
  • Exploratory data analysis to gauge explainable variance
  • Benchmarking models against simple baselines

⭐ Strength Points

  • Intuitive: expresses result as a percentage of variance explained
  • Normalized scale (0 to 1) — with possible negatives for poor fits
  • Useful for comparing model performance on the same dataset

⚠️ Weakness Points

  • Can be misleading for non-linear models or overfitted models
  • Does not penalize complexity — can overestimate poor fits
  • Doesn’t describe the magnitude of error

🎯 Best Practice Recommendation

  • Use R² as a quick snapshot of model fit
  • Prefer Adjusted R² when comparing models with different numbers of features
  • Combine with MAE or RMSE for error scale analysis

#5: Adjusted R-squared

📖 Definition

Adjusted R-squared modifies the regular R² score by penalizing the inclusion of unnecessary predictors. It improves upon R² by factoring in model complexity, helping avoid overfitting.

🧮 Mathematical Equation & Components

$$\text{Adjusted } R^2 = 1 - \left( \frac{(1 - R^2)(n - 1)}{n - p - 1} \right)$$

  • R² = standard coefficient of determination
  • n = number of observations
  • p = number of independent variables (predictors)

🧰 Typical Use Cases

  • Model comparison with varying numbers of features
  • Linear regression diagnostics to avoid overfitting
  • Feature selection and pruning

⭐ Strength Points

  • Penalizes models for adding non-informative variables
  • More realistic than R² when many predictors are involved
  • Favors parsimonious, generalizable models

⚠️ Weakness Points

  • Still assumes a linear model structure
  • Less meaningful for tree-based or nonparametric models
  • Not intuitive for non-technical audiences

🎯 Best Practice Recommendation

  • Use when comparing multiple linear models with different numbers of features
  • Helpful in feature engineering for identifying relevant predictors
  • Always evaluate alongside MAE, RMSE, or cross-validation metrics

#6: Mean Absolute Percentage Error (MAPE)

📖 Definition

MAPE measures prediction error as a percentage of the actual value. It averages the absolute percent difference between predicted and actual values across all predictions— making it ideal for understanding relative error.

🧮 Mathematical Equation & Components

$$\text{MAPE} = \frac{100\%}{N} \sum_{i=1}^{N} \left| \frac{y_i - \hat{y}_i}{y_i} \right|$$

  • $y_i$ = actual value
  • $\hat{y}_i$ = predicted value
  • $N$ = number of observations

🧰 Typical Use Cases

  • Business forecasting (e.g., sales, revenue)
  • Time series analysis with large scale differences
  • Model results reporting for non-technical audiences

⭐ Strength Points

  • Easy to interpret as percentage error
  • Helps with scale-independent evaluation
  • Enables model comparison across domains

⚠️ Weakness Points

  • Undefined or unstable when actual value ($y_i$) is zero or near-zero
  • Can be biased toward under-predictions
  • Penalizes large over-predictions more than large under-predictions

🎯 Best Practice Recommendation

  • Use when relative accuracy is more meaningful than absolute
  • Avoid when data includes zeros or near-zeros—use sMAPE instead
  • Always report with MAE or RMSE to contextualize absolute performance

#7: Symmetric Mean Absolute Percentage Error (sMAPE)

📖 Definition

sMAPE is an adjusted version of MAPE that provides a balanced view of relative error, treating over- and under-predictions symmetrically. It's especially useful when dealing with values near zero or when MAPE’s asymmetry introduces bias.

🧮 Mathematical Equation & Components

$$\text{sMAPE} = \frac{100\%}{N} \sum_{i=1}^{N} \frac{|\hat{y}_i - y_i|}{(|y_i| + |\hat{y}_i|)/2}$$

  • $y_i$ = actual value
  • $\hat{y}_i$ = predicted value
  • $N$ = number of data points

🧰 Typical Use Cases

  • Forecasting in sectors like retail, energy, traffic
  • Performance metrics for **zero-inclusive data**
  • Dashboard-ready metrics for **executive reporting**

⭐ Strength Points

  • Better than MAPE at handling **zero or near-zero values**
  • **Balanced treatment** of over- vs. under-estimation
  • Great for **comparative analysis across multiple series**

⚠️ Weakness Points

  • Still unstable if both prediction and actual ≈ 0
  • Can be **less intuitive** to explain than plain MAPE
  • Not natively available in some libraries like scikit-learn

🎯 Best Practice Recommendation

  • Prefer sMAPE over MAPE when your data contains zeros or small values
  • Use it in forecasting tasks where **balanced relative error** is important
  • Always cross-check with MAE or RMSE for absolute error magnitude

#8: Huber Loss

📖 Definition

Huber Loss is a hybrid metric that blends the best of Mean Squared Error (MSE) and Mean Absolute Error (MAE). It behaves like MSE for small errors (ensuring smooth optimization) and like MAE for large errors (reducing sensitivity to outliers).

🧮 Mathematical Equation & Components

$$ L_\delta(y, \hat{y}) = \begin{cases} \frac{1}{2}(y - \hat{y})^2 & \text{if } |y - \hat{y}| \leq \delta \\ \delta \cdot (|y - \hat{y}| - \frac{1}{2} \delta) & \text{otherwise} \end{cases} $$

  • $y$ = actual value
  • $\hat{y}$ = predicted value
  • $\delta$ = threshold that determines the switch between MSE and MAE behavior

🧰 Typical Use Cases

  • When data contains **outliers or noisy observations**
  • **Robust regression** settings and **gradient boosting frameworks**
  • Used in **deep learning loss functions** and **industrial predictive systems**

⭐ Strength Points

  • **Balances robustness and smooth optimization**
  • Reduces the impact of outliers without ignoring them entirely
  • Widely supported in libraries like XGBoost, TensorFlow, PyTorch

⚠️ Weakness Points

  • Requires selection of **$\delta$**, which affects model sensitivity
  • Slightly more **complex to compute** and tune compared to MAE or MSE
  • Performance depends on **good hyperparameter tuning**

🎯 Best Practice Recommendation

  • Use Huber Loss when your data has **moderate outliers** or noisy labels
  • Perform **cross-validation** to tune the δ parameter effectively
  • Combine with MAE/MSE evaluations for comprehensive analysis

#9: Explained Variance Score

📖 Definition

The Explained Variance Score quantifies how much of the variance in the target variable is captured by the model. It’s closely related to R² but focuses more on variance reduction rather than error relative to mean.

🧮 Mathematical Equation & Components

$$ \text{Explained Variance} = 1 - \frac{\text{Var}(y - \hat{y})}{\text{Var}(y)} $$

  • $y$: true values
  • $\hat{y}$: predicted values
  • Var($y - \hat{y}$): variance of prediction errors (residuals)
  • Var($y$): variance of actual target values

🧰 Typical Use Cases

  • Evaluating how well a model captures target variability
  • Continuous regression tasks where variance matters more than absolute accuracy
  • Time series, energy forecasting, or environmental modeling

⭐ Strength Points

  • Highlights the model’s ability to reduce unexplained variation
  • Less impacted by constant shift/bias than R²
  • Helps distinguish models with similar RMSE but different variance coverage

⚠️ Weakness Points

  • Can be negative if the model increases variance over baseline
  • May be misleading if model has high bias (systematic error)
  • Not as intuitive as RMSE or R² for non-technical audiences

🎯 Best Practice Recommendation

  • Use in tandem with R² and RMSE for a more complete regression assessment
  • Especially valuable when variance containment is the modeling priority
  • Avoid comparing across datasets with drastically different variance scales

#9: Explained Variance Score

📖 Definition

The Explained Variance Score quantifies how much of the variance in the target variable is captured by the model. It’s closely related to R² but focuses more on variance reduction rather than error relative to mean.

🧮 Mathematical Equation & Components

$$ \text{Explained Variance} = 1 - \frac{\text{Var}(y - \hat{y})}{\text{Var}(y)} $$

  • $y$: true values
  • $\hat{y}$: predicted values
  • Var($y - \hat{y}$): variance of prediction errors (residuals)
  • Var($y$): variance of actual target values

🧰 Typical Use Cases

  • Evaluating how well a model captures target variability
  • Continuous regression tasks where variance matters more than absolute accuracy
  • Time series, energy forecasting, or environmental modeling

⭐ Strength Points

  • Highlights the model’s ability to reduce unexplained variation
  • Less impacted by constant shift/bias than R²
  • Helps distinguish models with similar RMSE but different variance coverage

⚠️ Weakness Points

  • Can be negative if the model increases variance over baseline
  • May be misleading if model has high bias (systematic error)
  • Not as intuitive as RMSE or R² for non-technical audiences

🎯 Best Practice Recommendation

  • Use in tandem with R² and RMSE for a more complete regression assessment
  • Especially valuable when variance containment is the modeling priority
  • Avoid comparing across datasets with drastically different variance scales

#10: Mean Bias Deviation (MBD)

📖 Definition

Mean Bias Deviation (MBD) measures the average directional error of your model’s predictions. Unlike MAE or RMSE, which focus on magnitude, MBD reveals whether your model consistently over- or under-predicts.

🧮 Mathematical Equation & Components

$$ \text{MBD} = \frac{1}{N} \sum_{i=1}^{N} ( \hat{y}_i - y_i ) $$

  • $y_i$: actual value
  • $\hat{y}_i$: predicted value
  • $N$: total number of samples

🧰 Typical Use Cases

  • Detecting systematic prediction bias in regression models
  • Evaluating and tuning model calibration
  • Critical in energy forecasting, finance, climate models

⭐ Strength Points

  • Identifies whether model is biased toward over- or under-predicting
  • Easy to interpret and useful for model diagnostics
  • Helpful for bias correction workflows

⚠️ Weakness Points

  • Positive and negative errors cancel out, possibly hiding large errors
  • Should not be used in isolation—combine with MAE or RMSE
  • Sensitive to outliers

🎯 Best Practice Recommendation

  • Use MBD to detect not captured by other metrics
  • Pair with MAE, RMSE, or R² to complete the picture
  • Visualize MBD over time or by feature to identify model drift

📊 Comparative Atlas of Regression Metrics

🧭 1. Key Trade-offs Between Metrics

Metric Measures Sensitive to Outliers Easy to Interpret Directional Bias Penalizes Large Errors
MAE Avg. magnitude of error ❌ Low ✅ Yes ❌ No ❌ No
MSE Avg. squared error ✅ High ⚠️ Somewhat ❌ No ✅ Yes
RMSE Root of squared error ✅ High ✅ Yes ❌ No ✅ Yes
R² Variance explained ⚠️ Depends ✅ Yes ❌ No ❌ No
Adjusted R² Adjusted variance ⚠️ Depends ⚠️ Somewhat ❌ No ❌ No
MAPE % error ✅ High (near 0) ✅ Yes ❌ No ❌ No
sMAPE Symmetric % error ⚠️ Medium ✅ Yes ❌ No ❌ No
Huber Combined MAE/MSE ⚠️ Medium ⚠️ Model-specific ❌ No ✅ Yes (conditionally)
Explained Variance Variance reduced ❌ Low ⚠️ Somewhat ❌ No ❌ No
MBD Avg. signed bias ⚠️ Somewhat ⚠️ Model-specific ✅ Yes ❌ No

🎯 2. Use Case Selection Table

Use Case Recommended Metric(s)
General regression RMSE, MAE, R²
Outlier-resistant models MAE, Huber, sMAPE
Penalize large deviations MSE, RMSE
Compare models of different sizes Adjusted R²
Forecast error as % MAPE, sMAPE
Explainability to stakeholders MAE, RMSE, MAPE
Calibration check (bias) MBD
Variance-focused analysis Explained Variance, R²
Robust loss function for training Huber Loss

🧠 3. Visual Mapping Guide (What to Plot)

Metric Suggested Plot What It Shows
MAE, RMSE Residual plots, error histograms Spread and magnitude of errors
MAPE / sMAPE Line plot of percentage errors Relative forecasting performance
R² / Adjusted R² Model fit plots (predicted vs actual) Fit and dispersion pattern
MBD Residual mean line, bias plots Directional drift
Explained Variance Target vs prediction variance bar chart Model's explanatory power

🧪 4. Statistical Behavior & Optimization Suitability

Metric Convex Differentiable Used in Loss Functions
MAE ✅ ❌ (non-smooth) ⚠️ Sometimes
MSE ✅ ✅ ✅ Standard
RMSE ✅ ✅ ✅ (rarely directly)
Huber ✅ ✅ ✅ Robust loss

📦 5. Ensemble Use & Reporting

Best practice: report multiple metrics together to capture different dimensions of model performance:

  • MAE + RMSE → General error + sensitivity to large deviations
  • MAPE or sMAPE → Percentage-based errors, ideal for stakeholder-facing reports
  • R² + Adjusted R² → Model fit + complexity-aware comparison
  • MBD + Explained Variance → Bias + variance interpretation
  • Huber Loss → Use during training for robustness to outliers

📋 Regression Metrics Selection Checklist

  • MAE – easy to explain
  • RMSE – same units, emphasizes large errors
  • MAPE / sMAPE – percent error
  • MAE – robust
  • Huber Loss – adaptive
  • sMAPE – if percentage error needed
  • No outliers? → Use MSE, RMSE, or MAPE
  • MSE – quadratic penalty
  • RMSE – amplifies large errors
  • Huber Loss – combines linear and quadratic
  • MBD – Mean Bias Deviation
  • Check residual plots
  • Consider Explained Variance Score
  • MAPE / sMAPE – for relative error
  • MAE and MBD – for interpretability and drift detection
  • Adjusted R²
  • Also track R² and Explained Variance
  • MSE – standard loss
  • Huber Loss – robust training loss
  • MAPE / sMAPE – scale-invariant
  • Normalize MAE / RMSE by mean target value if needed