📏 Mean Absolute Error (MAE)

“Every mistake counts — just how much, not how big.”

📍 Definition

MAE computes the average absolute difference between the predicted and actual values:

$$ \text{MAE} = \frac{1}{n} \sum_{i=1}^{n} \left| y_i - \hat{y}_i \right| $$

  • \( y_i \): actual value
  • \( \hat{y}_i \): predicted value
  • \( n \): number of samples
  • Also known as L1 loss

📘 Derivative (Subgradient)

MAE is not differentiable at zero, but subgradients are used:

$$ \frac{\partial \text{MAE}}{\partial \hat{y}_i} = \begin{cases} \frac{1}{n} & \text{if } \hat{y}_i > y_i \\\\ -\frac{1}{n} & \text{if } \hat{y}_i < y_i \\\\ \text{undefined} & \text{if } \hat{y}_i = y_i \end{cases} $$

Auto-diff frameworks handle this gracefully with subgradients.

🧠 Intuition

  • Treats all errors linearly — no exaggeration of large ones
  • More robust to outliers than MSE
  • Slower to converge (gradient is constant)

🧮 Use Cases

ApplicationWhy MAE?
Robust regressionLess affected by noisy labels
Fair error measurementEqual penalty for all deviations
Sensor dataWhere noise spikes should not dominate
ForecastingTo avoid squaring large prediction gaps

🔄 Behavior

  • Zero loss only when prediction = truth
  • Constant slope — no acceleration from large errors
  • Not smooth at error = 0

⚠️ Trade-offs

ProCon
Robust to outliersSlower convergence than MSE
Simple and interpretableGradient doesn’t scale with error
No squaringMay lead to multiple minima (non-convex surface)

🔬 Visualization Tip

  • Plot the absolute error bars per sample
  • Overlay MAE and MSE loss surfaces — MAE is more angular

📦 PyTorch Example


import torch.nn as nn
loss_fn = nn.L1Loss()
loss = loss_fn(predictions, targets)
  

📊 MSE vs MAE Summary

Feature MSE MAE
PenaltySquared errorAbsolute error
Outlier sensitivityHighLow
GradientProportionalConstant
ConvergenceFasterSlower
Loss surfaceSmooth, convexAngular, piecewise

📚 Related Topics


🎯 Mean Squared Error (MSE)

“When your model guesses wrong, MSE tells it how wrong — and by how much.”

📍 Definition

MSE measures the average of the squared differences between predictions and actual targets:

$$ \text{MSE} = \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2 $$

  • \( y_i \): true value
  • \( \hat{y}_i \): predicted value
  • \( n \): number of samples
  • Also known as L2 Loss or squared loss

📘 Derivative (for backpropagation)

$$ \frac{\partial \text{MSE}}{\partial \hat{y}_i} = \frac{2}{n} (\hat{y}_i - y_i) $$

This gradient is used during gradient descent to update model parameters.

🧠 Intuition

  • Squaring punishes larger errors more
  • Smooth and differentiable — ideal for optimization
  • Sensitive to outliers due to the squaring

🧮 Use Cases

ApplicationWhy MSE?
Linear regressionCanonical loss for fitting lines
Neural nets for regressionSmooth gradient helps convergence
AutoencodersCommon loss for reconstructing images/signals
Time-series predictionPenalizes long-term deviation

🔄 Behavior

  • Zero loss only when prediction = truth exactly
  • Always non-negative
  • Squared error grows rapidly with bigger mistakes

⚠️ Limitations

IssueEffect
Sensitive to outliersBig errors dominate total loss
No robustness to noiseCan distort training
Not ideal for classificationRequires continuous targets

🔬 Visual Tip

  • Plot true vs predicted and mark squared error lines
  • Show loss surface of a simple linear model under MSE — convex and smooth

📦 PyTorch Example


import torch.nn as nn
loss_fn = nn.MSELoss()
loss = loss_fn(predictions, targets)
  

📊 Comparison Snapshot

Loss Function Use Case Robust to Outliers? Gradient Smooth?
MSERegression❌✅
MAERegression✅❌ (nondifferentiable at 0)
Cross-EntropyClassification—✅

📚 Related Topics


📏 R²: Coefficient of Determination

“How much of the variation in the data did your model explain?”

📍 Definition

R² compares the total variance in the target to the variance left unexplained by the model:

$$ R^2 = 1 - \frac{\sum_{i=1}^{n} (y_i - \hat{y}_i)^2}{\sum_{i=1}^{n} (y_i - \bar{y})^2} $$

  • \( y_i \): true value
  • \( \hat{y}_i \): predicted value
  • \( \bar{y} \): mean of actual values
  • Numerator: residual sum of squares (unexplained)
  • Denominator: total sum of squares (total variance)

🧠 Intuition

  • R² = 1 → perfect fit
  • R² = 0 → model explains no more than the mean
  • R² < 0 → model worse than a horizontal line at the mean
  • Measures proportion of variance explained

🧮 Use Cases

ApplicationWhy R²?
Linear regressionStandard interpretability metric
Model comparisonBenchmark across models
Feature importanceTracks how well inputs explain target
Model diagnosticsCombine with RMSE or MAE for full view

🔄 Behavior

  • High R² doesn’t mean “good” → check residuals too
  • R² improves with more features → be cautious of overfitting
  • Can be misleading in non-linear models

⚠️ Considerations

ProCon
Intuitive, easy to explainInflated by irrelevant features
Standard for regressionCan be negative or misleading
Unitless & relativeNot a loss function (not optimized)

🔬 Visualization Tip

  • Plot predicted vs true values with R² annotated
  • Also plot line of best fit vs line at \( \bar{y} \)

📦 PyTorch (as Eval Metric)


def r_squared(y_true, y_pred):
    ss_res = torch.sum((y_true - y_pred) ** 2)
    ss_tot = torch.sum((y_true - torch.mean(y_true)) ** 2)
    return 1 - ss_res / ss_tot
  

📊 Metric Snapshot

Metric Measures Ideal Value Notes
MSESquared error0Penalizes large errors
MAEAbsolute error0Robust to outliers
R²Explained variance1Proportion explained

📚 Related Topics


📏 Root Mean Squared Error (RMSE)

“Like MSE, but more human — in the same units as your data.”

📍 Definition

RMSE is the square root of MSE:

$$ \text{RMSE} = \sqrt{ \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2 } $$

  • \( y_i \): actual value
  • \( \hat{y}_i \): predicted value
  • \( n \): number of samples
  • It measures the standard deviation of prediction errors

🧠 Intuition

  • Penalizes large errors (like MSE)
  • Interpretable: in same units as original data
  • More intuitive to compare to actual values

📘 Derivative

While RMSE itself is not commonly differentiated directly, it's used for evaluation. During training, we typically use MSE for optimization and compute RMSE for monitoring:

$$ \frac{d(\text{RMSE})}{d\hat{y}_i} = \frac{1}{2 \cdot \text{RMSE} \cdot n} \cdot 2 (\hat{y}_i - y_i) = \frac{\hat{y}_i - y_i}{\text{RMSE} \cdot n} $$

But usually: optimize MSE, report RMSE

🧮 Use Cases

ApplicationWhy RMSE?
Model evaluationEasy-to-understand error metric
ForecastingCompare predicted vs real-world values
Regression competitionsRMSE is often the leaderboard metric
Sensor systemsMeasure average deviation in native units

🔄 Behavior

  • Always non-negative
  • Squaring magnifies larger errors
  • Square root reverses the unit distortion of MSE

⚠️ Trade-offs

ProCon
Readable in real unitsStill sensitive to outliers
Easy to explainNot differentiable at 0 (but not an issue in practice)
Smooth, global error measureCan mask distribution of individual errors

🔬 Visualization Tip

  • Plot actual vs predicted, then show RMSE as average vertical distance
  • Overlay RMSE on regression predictions for intuition

📦 PyTorch Example (as Eval Metric)


import torch
rmse = torch.sqrt(torch.mean((predictions - targets) ** 2))
  

📊 MSE vs RMSE

Metric Units Penalizes Big Errors Easy to Interpret
MSESquared units✅✅✅❌
RMSEOriginal units✅✅✅✅✅✅

📚 Related Topics


🔧 Huber Loss

“When you want to be smooth and resilient — go Huber.”

📍 Definition

Huber Loss behaves like MSE near zero and like MAE for large errors. It’s defined as:

$$ \mathcal{L}_\delta(a) = \begin{cases} \frac{1}{2} a^2 & \text{for } |a| \leq \delta \\\\ \delta \left( |a| - \frac{1}{2} \delta \right) & \text{for } |a| > \delta \end{cases} $$

  • \( a = y_i - \hat{y}_i \): the residual
  • \( \delta \): threshold that controls the transition point

📘 Derivative

$$ \frac{d\mathcal{L}_\delta}{da} = \begin{cases} a & \text{for } |a| \leq \delta \\\\ \delta \cdot \text{sign}(a) & \text{for } |a| > \delta \end{cases} $$

Smooth transition from quadratic to linear gradient

🧠 Intuition

  • Like MSE for small errors (precise, smooth)
  • Like MAE for big errors (robust to outliers)
  • The \( \delta \) hyperparameter controls sensitivity to outliers

🧮 Use Cases

ApplicationWhy Huber?
Robust regressionBetter than MSE when data has outliers
Reinforcement learningPopular in TD-error loss
Sensor readingsSmooth yet tolerant to occasional spikes
Medical/finance predictionsWhere precision and robustness both matter

🔄 Behavior

  • Smooth everywhere — unlike MAE
  • Loss curve is quadratic in center, linear in tails
  • Controlled with \( \delta \) — can be tuned per problem

⚠️ Notes

  • Choosing \( \delta \) too small ≈ MAE
  • Choosing \( \delta \) too large ≈ MSE
  • Adds a hyperparameter, but balances bias/variance

🔬 Visualization Tip

  • Plot MSE, MAE, and Huber on same graph:
    • MSE = sharp curve
    • MAE = V-shape
    • Huber = curve + line hybrid

📦 PyTorch Example


import torch.nn as nn
loss_fn = nn.HuberLoss(delta=1.0)
loss = loss_fn(predictions, targets)
  

📊 Summary Snapshot

Feature MSE MAE Huber
Outlier robustness❌✅✅
Smoothness✅❌✅
HyperparameterNoneNone\( \delta \)
Use casePrecise errorNoisy dataBalanced needs

📚 Related Topics


⚖️ Hinge Loss

“Don’t just classify — separate with a confident margin.”

📍 Definition

Hinge Loss is used for “maximum-margin” classification, particularly with linear classifiers like SVMs.

For binary classification with labels \( y \in \{-1, 1\} \):

$$ \text{Hinge}(y, \hat{y}) = \max(0, 1 - y \cdot \hat{y}) $$

  • \( \hat{y} \): raw output score (not a probability)
  • \( y \): true label in \{-1, 1\}
  • Loss is zero when the prediction is confidently correct (i.e., margin ≥ 1)

📘 Derivative (Subgradient)

$$ \frac{d}{d\hat{y}} \text{Hinge}(y, \hat{y}) = \begin{cases} 0 & \text{if } y \cdot \hat{y} \geq 1 \\\\ - y & \text{otherwise} \end{cases} $$

Only penalizes when prediction is on the wrong side or too close to the decision boundary.

🧠 Intuition

  • Enforces a margin of separation between classes
  • Focuses on confidence as well as correctness
  • Non-zero loss even for correct predictions if they’re too close to the margin

🧮 Use Cases

ApplicationWhy Hinge?
Support Vector MachinesCanonical loss for SVM training
Binary classifiers (with raw logits)Encourages confident separation
Linear margin classifiersEnforces geometric separation
Deep learning variantsCan improve robustness (e.g. max-margin loss)

🔄 Behavior

  • Zero loss for “correct and confident”
  • Linear penalty as predictions drift across the boundary
  • No penalty for margin > 1

⚠️ Considerations

ProCon
Encourages large marginOnly works with raw scores
Sparse gradientsNot probabilistic
Robust to noisy near-margin pointsCan’t use soft labels

🔬 Visualize

  • Plot \( 1 - y \cdot \hat{y} \) vs. hinge loss
  • Highlight region where margin < 1 triggers penalties

📦 PyTorch Implementation


import torch

def hinge_loss(outputs, targets):
    return torch.mean(torch.clamp(1 - outputs * targets, min=0))
  

📊 Comparison Snapshot

Loss Type Probabilistic? Margin-Aware Use Case
MSERegression❌❌Continuous targets
BCEClassification✅❌Probabilities
HingeClassification❌✅Raw output, margin classifiers

📚 Related Topics


📉 Binary Cross-Entropy (Log Loss)

“When your model predicts probability, BCE tells how close it came to the truth.”

📍 Definition

For binary labels \( y \in \{0,1\} \) and predicted probabilities \( \hat{y} \in (0,1) \):

$$ \text{BCE}(y, \hat{y}) = -[y \log(\hat{y}) + (1 - y) \log(1 - \hat{y})] $$

  • \( y \): actual label
  • \( \hat{y} \): model's predicted probability of class 1
  • Loss is averaged over all samples in a batch

📘 Derivative (for backprop)

$$ \frac{d}{d\hat{y}} \text{BCE} = \frac{\hat{y} - y}{\hat{y}(1 - \hat{y})} $$

Auto-diff frameworks apply this stably using log-sum-exp tricks

🧠 Intuition

  • If model says 90% for a true 1 → small loss
  • If model says 10% for a true 1 → high loss
  • Encourages confident, correct predictions
  • Punishes overconfidence on wrong labels

🧮 Use Cases

ApplicationWhy BCE?
Binary classificationStandard loss for 0/1 tasks
Logistic regressionBCE used with sigmoid outputs
Anomaly detectionScore normal vs anomaly
Binary multi-labelApply BCE to each label independently

🔄 Behavior

  • Loss → 0 as prediction → correct label
  • Loss → ∞ as prediction → opposite of label
  • Asymmetric: wrong + confident = harsh penalty

⚠️ Considerations

ProCon
Probabilistic interpretationSensitive to extreme predictions
Works with sigmoidNeeds output in (0,1)
Good for uncertainty modelingCan be unstable without clipping

🔬 Visualization

  • Plot BCE vs predicted probability for \( y = 0 \) and \( y = 1 \)
  • Show steep curves around wrong predictions

📦 PyTorch Example


import torch.nn as nn

# For sigmoid outputs in (0, 1)
loss_fn = nn.BCELoss()

# If using raw logits (faster + numerically stable)
loss_fn = nn.BCEWithLogitsLoss()
  

📊 BCE Summary

Feature Value
Range\[0, ∞)
Best value0 (perfect prediction)
Use caseBinary classification
Pairs withSigmoid activation
Alternate namesLog loss

📚 Related Topics


🎯 Categorical Cross-Entropy (CCE)

“When your model picks from multiple classes, CCE rewards the sharp and correct.”

📍 Definition

For a true one-hot label vector \( \mathbf{y} \in \{0,1\}^K \) and predicted probability vector \( \hat{\mathbf{y}} \in (0,1)^K \):

$$ \text{CCE}(\mathbf{y}, \hat{\mathbf{y}}) = - \sum_{i=1}^{K} y_i \log(\hat{y}_i) $$

Only the log of the predicted probability for the correct class contributes to the loss.

📘 Derivative

$$ \frac{\partial \text{CCE}}{\partial \hat{y}_j} = \begin{cases} - \frac{1}{\hat{y}_j} & \text{if } j = \text{true class} \\\\ 0 & \text{otherwise} \end{cases} $$

Backprop is usually applied through softmax + log as one stable function.

🧠 Intuition

  • Encourages high probability for the correct class
  • Strongly penalizes wrong + confident predictions
  • Best used with softmax in the output layer

🧮 Use Cases

ApplicationWhy CCE?
Multiclass classificationOne correct label among many
Image recognition10-class tasks like CIFAR-10
NLP (language modeling)Vocabulary prediction via softmax
Speech, text, visionCore loss for supervised learning

🔄 Behavior

  • Loss is low when the correct class has high predicted probability
  • Loss grows sharply as correct probability shrinks
  • Logarithmic curve penalizes overconfidence in wrong answers

⚠️ Considerations

ProCon
Probabilistic and interpretableNeeds careful numerical stability
Well-understood in theoryNeeds one-hot or class index labels
Compatible with softmaxNot for multilabel tasks (use BCE)

🔬 Visualization Tip

  • Plot softmax output vs. loss for different class positions
  • Show how loss decreases as the correct class probability increases

📦 PyTorch Example


import torch.nn as nn

# CCE expects raw logits (not softmaxed), and class indices as targets
loss_fn = nn.CrossEntropyLoss()
loss = loss_fn(logits, labels)  # labels: LongTensor of class indices
  

📊 CCE Summary

Feature Value
Loss range\[0, ∞)
Target formatClass index or one-hot
Output formatSoftmax probabilities
Common pairsSoftmax + CCE
Multiclass support✅
Multilabel support❌ (use BCE instead)

📚 Related Topics


🧮 Sparse Categorical Cross-Entropy

“When labels are indices, not vectors — sparse CCE steps in.”

📍 Definition

Sparse CCE shares the same formula as standard CCE:

$$ \text{CCE}(y, \hat{\mathbf{y}}) = - \log(\hat{y}_{\text{true class}}) $$

  • \( y \in \{0, 1, ..., K-1\} \): class index (not one-hot)
  • \( \hat{\mathbf{y}} \in (0,1)^K \): softmax probabilities
  • Picks the log probability of the correct class directly

📘 Derivative

$$ \frac{\partial \text{Loss}}{\partial \hat{y}_i} = \begin{cases} - \frac{1}{\hat{y}_i} & \text{if } i = y \\\\ 0 & \text{otherwise} \end{cases} $$

Same as CCE — targets the correct class index only.

🧠 Intuition

  • Skips one-hot encoding → saves memory
  • Numerically identical to standard CCE
  • Ideal for large output spaces (e.g., vocabularies)

🧮 Use Cases

ApplicationWhy Sparse CCE?
Multiclass classificationEfficient label format
NLP (language modeling)Large vocab class index
Image classificationAvoids one-hot expansion
Transformer decodingIndex-based softmax loss

🔄 Behavior

  • Same gradient and output as categorical cross-entropy
  • Input: logits or softmaxed probabilities
  • Target: class indices like tensor([3, 0, 1])

⚠️ Considerations

ProCon
No need for one-hot labelsCan't handle multilabel targets
More memory efficientStill needs softmax/logits
Same as CCE under the hoodRequires integer class labels

🔬 Visualization Tip

  • Show how label index directly indexes into the softmax vector
  • Compare memory cost of one-hot vs sparse encoding

📦 PyTorch Example


import torch.nn as nn

loss_fn = nn.CrossEntropyLoss()  # logits in, class index out
loss = loss_fn(logits, class_indices)  # class_indices: LongTensor of shape [batch]
  

📊 Summary: CCE vs Sparse CCE

Feature CCE Sparse CCE
TargetOne-hot vectorClass index
MemoryHigherLower
Ease of useMore prep neededSimpler label input
Best forSmall to mid class countLarge-class problems

📚 Related Topics

📐 KL Divergence

“How much extra info is needed to use Q instead of the true P?”

📍 Definition

Given two probability distributions \( P \) (true) and \( Q \) (approximate), KL divergence is:

$$ D_{KL}(P \parallel Q) = \sum_{i} P(i) \log \left( \frac{P(i)}{Q(i)} \right) $$

For continuous distributions:

$$ D_{KL}(P \parallel Q) = \int p(x) \log \left( \frac{p(x)}{q(x)} \right) dx $$

  • \( P \): ground-truth distribution
  • \( Q \): approximate or predicted distribution
  • Not symmetric: \( D_{KL}(P \parallel Q) \ne D_{KL}(Q \parallel P) \)

🧠 Intuition

  • KL tells how much information is lost when using \( Q \) to approximate \( P \)
  • If \( Q \approx P \), then KL ≈ 0
  • Higher KL → worse approximation

📘 Gradient

KL divergence is differentiable with respect to \( Q \), making it usable in neural nets for training distributional outputs (e.g. VAEs, language models).

🧮 Use Cases

ApplicationWhy KL Divergence?
Variational Autoencoders (VAEs)Regularizes latent space against a prior
Language modelingMeasures divergence from expected token distribution
Reinforcement LearningBounds policy updates (TRPO, PPO)
Bayesian deep learningCompares posterior and prior distributions

🔄 Behavior

  • KL \( \geq 0 \)
  • KL = 0 ⇔ \( P = Q \)
  • KL → ∞ if \( Q(i) = 0 \) and \( P(i) > 0 \) (division by zero risk)

⚠️ Considerations

ProCon
Captures full distribution mismatchNot symmetric
Differentiable (good for training)Sensitive to small \( Q \)-values
Core to VAEs, RL, NLPRequires normalized distributions

🔬 Visualization Tip

  • Overlay distributions \( P \) and \( Q \), highlight where mismatch contributes most
  • Visualize KL as "cost" of substituting \( Q \) for \( P \)

📦 PyTorch Example


import torch.nn.functional as F

# Q: log probabilities (logits passed through log_softmax)
# P: true probabilities or one-hot vectors
kl = F.kl_div(Q.log(), P, reduction='batchmean')
  

📊 KL Summary

Feature KL Divergence
TypeAsymmetric distance
OutputScalar \( \geq 0 \)
Used withProbabilities / Softmax
Common inVAEs, RL, NLP
Key roleRegularization / Distribution alignment

📚 Related Topics

🧾 Negative Log-Likelihood (NLL)

“Log the chance of the truth — then flip the sign.”

📍 Definition

For classification using log-probabilities \( \log(\hat{p}) \) and true class label \( y \):

$$ \text{NLL}(\log(\hat{p}), y) = -\log(\hat{p}_y) $$

  • \( \hat{p}_y \): predicted probability for true class \( y \)
  • Input must be log-probabilities → typically from log_softmax

🧠 Intuition

  • Picks the log-probability for the correct class
  • Negates it → confident predictions get low loss
  • Mathematically equivalent to cross-entropy, when paired with softmax

🧮 Use Cases

ApplicationWhy NLL?
Multiclass classificationUsed with precomputed log_softmax
Language modelingPredicting next word/token from vocab
Custom probabilistic modelsExplicit work in log-probability space
Modular pipelinesSeparation of softmax and loss computation

🔄 Behavior

  • Loss → 0 as predicted log-prob of true class → 0 (i.e., prob → 1)
  • Loss grows quickly as confidence in true class drops
  • Stable in log-space, great for numerical safety

⚠️ Considerations

ProCon
Efficient in log-spaceRequires log-probs as input
Pairs with log_softmaxNot for raw logits directly
Numerically stableCan be redundant vs CrossEntropyLoss

🔬 Visualization Tip

  • Plot log-probabilities vs NLL for the correct class
  • Highlight steep rise in loss as log-prob decreases

📦 PyTorch Example


import torch.nn as nn

loss_fn = nn.NLLLoss()
log_probs = torch.log_softmax(logits, dim=1)
loss = loss_fn(log_probs, targets)
  

📊 NLL vs Cross-Entropy

FeatureNLLCrossEntropy
InputLog-probabilitiesRaw logits
Requires log_softmax?✅❌ (included)
Numerical stability✅✅
FlexibilityModular, manualSimplified, automatic

📚 Related Topics