📏 Mean Absolute Error (MAE)
“Every mistake counts — just how much, not how big.”
📍 Definition
MAE computes the average absolute difference between the predicted and actual values:
$$ \text{MAE} = \frac{1}{n} \sum_{i=1}^{n} \left| y_i - \hat{y}_i \right| $$
- \( y_i \): actual value
- \( \hat{y}_i \): predicted value
- \( n \): number of samples
- Also known as L1 loss
📘 Derivative (Subgradient)
MAE is not differentiable at zero, but subgradients are used:
$$ \frac{\partial \text{MAE}}{\partial \hat{y}_i} = \begin{cases} \frac{1}{n} & \text{if } \hat{y}_i > y_i \\\\ -\frac{1}{n} & \text{if } \hat{y}_i < y_i \\\\ \text{undefined} & \text{if } \hat{y}_i = y_i \end{cases} $$
Auto-diff frameworks handle this gracefully with subgradients.
🧠 Intuition
- Treats all errors linearly — no exaggeration of large ones
- More robust to outliers than MSE
- Slower to converge (gradient is constant)
🧮 Use Cases
| Application | Why MAE? |
|---|---|
| Robust regression | Less affected by noisy labels |
| Fair error measurement | Equal penalty for all deviations |
| Sensor data | Where noise spikes should not dominate |
| Forecasting | To avoid squaring large prediction gaps |
🔄 Behavior
- Zero loss only when prediction = truth
- Constant slope — no acceleration from large errors
- Not smooth at error = 0
⚠️ Trade-offs
| Pro | Con |
|---|---|
| Robust to outliers | Slower convergence than MSE |
| Simple and interpretable | Gradient doesn’t scale with error |
| No squaring | May lead to multiple minima (non-convex surface) |
🔬 Visualization Tip
- Plot the absolute error bars per sample
- Overlay MAE and MSE loss surfaces — MAE is more angular
📦 PyTorch Example
import torch.nn as nn
loss_fn = nn.L1Loss()
loss = loss_fn(predictions, targets)
📊 MSE vs MAE Summary
| Feature | MSE | MAE |
|---|---|---|
| Penalty | Squared error | Absolute error |
| Outlier sensitivity | High | Low |
| Gradient | Proportional | Constant |
| Convergence | Faster | Slower |
| Loss surface | Smooth, convex | Angular, piecewise |
📚 Related Topics
🎯 Mean Squared Error (MSE)
“When your model guesses wrong, MSE tells it how wrong — and by how much.”
📍 Definition
MSE measures the average of the squared differences between predictions and actual targets:
$$ \text{MSE} = \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2 $$
- \( y_i \): true value
- \( \hat{y}_i \): predicted value
- \( n \): number of samples
- Also known as L2 Loss or squared loss
📘 Derivative (for backpropagation)
$$ \frac{\partial \text{MSE}}{\partial \hat{y}_i} = \frac{2}{n} (\hat{y}_i - y_i) $$
This gradient is used during gradient descent to update model parameters.
🧠 Intuition
- Squaring punishes larger errors more
- Smooth and differentiable — ideal for optimization
- Sensitive to outliers due to the squaring
🧮 Use Cases
| Application | Why MSE? |
|---|---|
| Linear regression | Canonical loss for fitting lines |
| Neural nets for regression | Smooth gradient helps convergence |
| Autoencoders | Common loss for reconstructing images/signals |
| Time-series prediction | Penalizes long-term deviation |
🔄 Behavior
- Zero loss only when prediction = truth exactly
- Always non-negative
- Squared error grows rapidly with bigger mistakes
⚠️ Limitations
| Issue | Effect |
|---|---|
| Sensitive to outliers | Big errors dominate total loss |
| No robustness to noise | Can distort training |
| Not ideal for classification | Requires continuous targets |
🔬 Visual Tip
- Plot true vs predicted and mark squared error lines
- Show loss surface of a simple linear model under MSE — convex and smooth
📦 PyTorch Example
import torch.nn as nn
loss_fn = nn.MSELoss()
loss = loss_fn(predictions, targets)
📊 Comparison Snapshot
| Loss Function | Use Case | Robust to Outliers? | Gradient Smooth? |
|---|---|---|---|
| MSE | Regression | ❌ | ✅ |
| MAE | Regression | ✅ | ❌ (nondifferentiable at 0) |
| Cross-Entropy | Classification | — | ✅ |
📚 Related Topics
📏 R²: Coefficient of Determination
“How much of the variation in the data did your model explain?”
📍 Definition
R² compares the total variance in the target to the variance left unexplained by the model:
$$ R^2 = 1 - \frac{\sum_{i=1}^{n} (y_i - \hat{y}_i)^2}{\sum_{i=1}^{n} (y_i - \bar{y})^2} $$
- \( y_i \): true value
- \( \hat{y}_i \): predicted value
- \( \bar{y} \): mean of actual values
- Numerator: residual sum of squares (unexplained)
- Denominator: total sum of squares (total variance)
🧠 Intuition
- R² = 1 → perfect fit
- R² = 0 → model explains no more than the mean
- R² < 0 → model worse than a horizontal line at the mean
- Measures proportion of variance explained
🧮 Use Cases
| Application | Why R²? |
|---|---|
| Linear regression | Standard interpretability metric |
| Model comparison | Benchmark across models |
| Feature importance | Tracks how well inputs explain target |
| Model diagnostics | Combine with RMSE or MAE for full view |
🔄 Behavior
- High R² doesn’t mean “good” → check residuals too
- R² improves with more features → be cautious of overfitting
- Can be misleading in non-linear models
⚠️ Considerations
| Pro | Con |
|---|---|
| Intuitive, easy to explain | Inflated by irrelevant features |
| Standard for regression | Can be negative or misleading |
| Unitless & relative | Not a loss function (not optimized) |
🔬 Visualization Tip
- Plot predicted vs true values with R² annotated
- Also plot line of best fit vs line at \( \bar{y} \)
📦 PyTorch (as Eval Metric)
def r_squared(y_true, y_pred):
ss_res = torch.sum((y_true - y_pred) ** 2)
ss_tot = torch.sum((y_true - torch.mean(y_true)) ** 2)
return 1 - ss_res / ss_tot
📊 Metric Snapshot
| Metric | Measures | Ideal Value | Notes |
|---|---|---|---|
| MSE | Squared error | 0 | Penalizes large errors |
| MAE | Absolute error | 0 | Robust to outliers |
| R² | Explained variance | 1 | Proportion explained |
📚 Related Topics
📏 Root Mean Squared Error (RMSE)
“Like MSE, but more human — in the same units as your data.”
📍 Definition
RMSE is the square root of MSE:
$$ \text{RMSE} = \sqrt{ \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2 } $$
- \( y_i \): actual value
- \( \hat{y}_i \): predicted value
- \( n \): number of samples
- It measures the standard deviation of prediction errors
🧠 Intuition
- Penalizes large errors (like MSE)
- Interpretable: in same units as original data
- More intuitive to compare to actual values
📘 Derivative
While RMSE itself is not commonly differentiated directly, it's used for evaluation. During training, we typically use MSE for optimization and compute RMSE for monitoring:
$$ \frac{d(\text{RMSE})}{d\hat{y}_i} = \frac{1}{2 \cdot \text{RMSE} \cdot n} \cdot 2 (\hat{y}_i - y_i) = \frac{\hat{y}_i - y_i}{\text{RMSE} \cdot n} $$
But usually: optimize MSE, report RMSE
🧮 Use Cases
| Application | Why RMSE? |
|---|---|
| Model evaluation | Easy-to-understand error metric |
| Forecasting | Compare predicted vs real-world values |
| Regression competitions | RMSE is often the leaderboard metric |
| Sensor systems | Measure average deviation in native units |
🔄 Behavior
- Always non-negative
- Squaring magnifies larger errors
- Square root reverses the unit distortion of MSE
⚠️ Trade-offs
| Pro | Con |
|---|---|
| Readable in real units | Still sensitive to outliers |
| Easy to explain | Not differentiable at 0 (but not an issue in practice) |
| Smooth, global error measure | Can mask distribution of individual errors |
🔬 Visualization Tip
- Plot actual vs predicted, then show RMSE as average vertical distance
- Overlay RMSE on regression predictions for intuition
📦 PyTorch Example (as Eval Metric)
import torch
rmse = torch.sqrt(torch.mean((predictions - targets) ** 2))
📊 MSE vs RMSE
| Metric | Units | Penalizes Big Errors | Easy to Interpret |
|---|---|---|---|
| MSE | Squared units | ✅✅✅ | ❌ |
| RMSE | Original units | ✅✅✅ | ✅✅✅ |
📚 Related Topics
🔧 Huber Loss
“When you want to be smooth and resilient — go Huber.”
📍 Definition
Huber Loss behaves like MSE near zero and like MAE for large errors. It’s defined as:
$$ \mathcal{L}_\delta(a) = \begin{cases} \frac{1}{2} a^2 & \text{for } |a| \leq \delta \\\\ \delta \left( |a| - \frac{1}{2} \delta \right) & \text{for } |a| > \delta \end{cases} $$
- \( a = y_i - \hat{y}_i \): the residual
- \( \delta \): threshold that controls the transition point
📘 Derivative
$$ \frac{d\mathcal{L}_\delta}{da} = \begin{cases} a & \text{for } |a| \leq \delta \\\\ \delta \cdot \text{sign}(a) & \text{for } |a| > \delta \end{cases} $$
Smooth transition from quadratic to linear gradient
🧠 Intuition
- Like MSE for small errors (precise, smooth)
- Like MAE for big errors (robust to outliers)
- The \( \delta \) hyperparameter controls sensitivity to outliers
🧮 Use Cases
| Application | Why Huber? |
|---|---|
| Robust regression | Better than MSE when data has outliers |
| Reinforcement learning | Popular in TD-error loss |
| Sensor readings | Smooth yet tolerant to occasional spikes |
| Medical/finance predictions | Where precision and robustness both matter |
🔄 Behavior
- Smooth everywhere — unlike MAE
- Loss curve is quadratic in center, linear in tails
- Controlled with \( \delta \) — can be tuned per problem
⚠️ Notes
- Choosing \( \delta \) too small ≈ MAE
- Choosing \( \delta \) too large ≈ MSE
- Adds a hyperparameter, but balances bias/variance
🔬 Visualization Tip
- Plot MSE, MAE, and Huber on same graph:
- MSE = sharp curve
- MAE = V-shape
- Huber = curve + line hybrid
📦 PyTorch Example
import torch.nn as nn
loss_fn = nn.HuberLoss(delta=1.0)
loss = loss_fn(predictions, targets)
📊 Summary Snapshot
| Feature | MSE | MAE | Huber |
|---|---|---|---|
| Outlier robustness | ❌ | ✅ | ✅ |
| Smoothness | ✅ | ❌ | ✅ |
| Hyperparameter | None | None | \( \delta \) |
| Use case | Precise error | Noisy data | Balanced needs |
📚 Related Topics
- Mean Squared Error (MSE)
- Mean Absolute Error (MAE)
- Root Mean Squared Error (RMSE)
- Robust Regression Techniques
⚖️ Hinge Loss
“Don’t just classify — separate with a confident margin.”
📍 Definition
Hinge Loss is used for “maximum-margin” classification, particularly with linear classifiers like SVMs.
For binary classification with labels \( y \in \{-1, 1\} \):
$$ \text{Hinge}(y, \hat{y}) = \max(0, 1 - y \cdot \hat{y}) $$
- \( \hat{y} \): raw output score (not a probability)
- \( y \): true label in \{-1, 1\}
- Loss is zero when the prediction is confidently correct (i.e., margin ≥ 1)
📘 Derivative (Subgradient)
$$ \frac{d}{d\hat{y}} \text{Hinge}(y, \hat{y}) = \begin{cases} 0 & \text{if } y \cdot \hat{y} \geq 1 \\\\ - y & \text{otherwise} \end{cases} $$
Only penalizes when prediction is on the wrong side or too close to the decision boundary.
🧠 Intuition
- Enforces a margin of separation between classes
- Focuses on confidence as well as correctness
- Non-zero loss even for correct predictions if they’re too close to the margin
🧮 Use Cases
| Application | Why Hinge? |
|---|---|
| Support Vector Machines | Canonical loss for SVM training |
| Binary classifiers (with raw logits) | Encourages confident separation |
| Linear margin classifiers | Enforces geometric separation |
| Deep learning variants | Can improve robustness (e.g. max-margin loss) |
🔄 Behavior
- Zero loss for “correct and confident”
- Linear penalty as predictions drift across the boundary
- No penalty for margin > 1
⚠️ Considerations
| Pro | Con |
|---|---|
| Encourages large margin | Only works with raw scores |
| Sparse gradients | Not probabilistic |
| Robust to noisy near-margin points | Can’t use soft labels |
🔬 Visualize
- Plot \( 1 - y \cdot \hat{y} \) vs. hinge loss
- Highlight region where margin < 1 triggers penalties
📦 PyTorch Implementation
import torch
def hinge_loss(outputs, targets):
return torch.mean(torch.clamp(1 - outputs * targets, min=0))
📊 Comparison Snapshot
| Loss | Type | Probabilistic? | Margin-Aware | Use Case |
|---|---|---|---|---|
| MSE | Regression | ❌ | ❌ | Continuous targets |
| BCE | Classification | ✅ | ❌ | Probabilities |
| Hinge | Classification | ❌ | ✅ | Raw output, margin classifiers |
📚 Related Topics
📉 Binary Cross-Entropy (Log Loss)
“When your model predicts probability, BCE tells how close it came to the truth.”
📍 Definition
For binary labels \( y \in \{0,1\} \) and predicted probabilities \( \hat{y} \in (0,1) \):
$$ \text{BCE}(y, \hat{y}) = -[y \log(\hat{y}) + (1 - y) \log(1 - \hat{y})] $$
- \( y \): actual label
- \( \hat{y} \): model's predicted probability of class 1
- Loss is averaged over all samples in a batch
📘 Derivative (for backprop)
$$ \frac{d}{d\hat{y}} \text{BCE} = \frac{\hat{y} - y}{\hat{y}(1 - \hat{y})} $$
Auto-diff frameworks apply this stably using log-sum-exp tricks
🧠 Intuition
- If model says 90% for a true 1 → small loss
- If model says 10% for a true 1 → high loss
- Encourages confident, correct predictions
- Punishes overconfidence on wrong labels
🧮 Use Cases
| Application | Why BCE? |
|---|---|
| Binary classification | Standard loss for 0/1 tasks |
| Logistic regression | BCE used with sigmoid outputs |
| Anomaly detection | Score normal vs anomaly |
| Binary multi-label | Apply BCE to each label independently |
🔄 Behavior
- Loss → 0 as prediction → correct label
- Loss → ∞ as prediction → opposite of label
- Asymmetric: wrong + confident = harsh penalty
⚠️ Considerations
| Pro | Con |
|---|---|
| Probabilistic interpretation | Sensitive to extreme predictions |
| Works with sigmoid | Needs output in (0,1) |
| Good for uncertainty modeling | Can be unstable without clipping |
🔬 Visualization
- Plot BCE vs predicted probability for \( y = 0 \) and \( y = 1 \)
- Show steep curves around wrong predictions
📦 PyTorch Example
import torch.nn as nn
# For sigmoid outputs in (0, 1)
loss_fn = nn.BCELoss()
# If using raw logits (faster + numerically stable)
loss_fn = nn.BCEWithLogitsLoss()
📊 BCE Summary
| Feature | Value |
|---|---|
| Range | \[0, ∞) |
| Best value | 0 (perfect prediction) |
| Use case | Binary classification |
| Pairs with | Sigmoid activation |
| Alternate names | Log loss |
📚 Related Topics
🎯 Categorical Cross-Entropy (CCE)
“When your model picks from multiple classes, CCE rewards the sharp and correct.”
📍 Definition
For a true one-hot label vector \( \mathbf{y} \in \{0,1\}^K \) and predicted probability vector \( \hat{\mathbf{y}} \in (0,1)^K \):
$$ \text{CCE}(\mathbf{y}, \hat{\mathbf{y}}) = - \sum_{i=1}^{K} y_i \log(\hat{y}_i) $$
Only the log of the predicted probability for the correct class contributes to the loss.
📘 Derivative
$$ \frac{\partial \text{CCE}}{\partial \hat{y}_j} = \begin{cases} - \frac{1}{\hat{y}_j} & \text{if } j = \text{true class} \\\\ 0 & \text{otherwise} \end{cases} $$
Backprop is usually applied through softmax + log as one stable function.
🧠 Intuition
- Encourages high probability for the correct class
- Strongly penalizes wrong + confident predictions
- Best used with
softmaxin the output layer
🧮 Use Cases
| Application | Why CCE? |
|---|---|
| Multiclass classification | One correct label among many |
| Image recognition | 10-class tasks like CIFAR-10 |
| NLP (language modeling) | Vocabulary prediction via softmax |
| Speech, text, vision | Core loss for supervised learning |
🔄 Behavior
- Loss is low when the correct class has high predicted probability
- Loss grows sharply as correct probability shrinks
- Logarithmic curve penalizes overconfidence in wrong answers
⚠️ Considerations
| Pro | Con |
|---|---|
| Probabilistic and interpretable | Needs careful numerical stability |
| Well-understood in theory | Needs one-hot or class index labels |
| Compatible with softmax | Not for multilabel tasks (use BCE) |
🔬 Visualization Tip
- Plot softmax output vs. loss for different class positions
- Show how loss decreases as the correct class probability increases
📦 PyTorch Example
import torch.nn as nn
# CCE expects raw logits (not softmaxed), and class indices as targets
loss_fn = nn.CrossEntropyLoss()
loss = loss_fn(logits, labels) # labels: LongTensor of class indices
📊 CCE Summary
| Feature | Value |
|---|---|
| Loss range | \[0, ∞) |
| Target format | Class index or one-hot |
| Output format | Softmax probabilities |
| Common pairs | Softmax + CCE |
| Multiclass support | ✅ |
| Multilabel support | ❌ (use BCE instead) |
📚 Related Topics
🧮 Sparse Categorical Cross-Entropy
“When labels are indices, not vectors — sparse CCE steps in.”
📍 Definition
Sparse CCE shares the same formula as standard CCE:
$$ \text{CCE}(y, \hat{\mathbf{y}}) = - \log(\hat{y}_{\text{true class}}) $$
- \( y \in \{0, 1, ..., K-1\} \): class index (not one-hot)
- \( \hat{\mathbf{y}} \in (0,1)^K \): softmax probabilities
- Picks the log probability of the correct class directly
📘 Derivative
$$ \frac{\partial \text{Loss}}{\partial \hat{y}_i} = \begin{cases} - \frac{1}{\hat{y}_i} & \text{if } i = y \\\\ 0 & \text{otherwise} \end{cases} $$
Same as CCE — targets the correct class index only.
🧠 Intuition
- Skips one-hot encoding → saves memory
- Numerically identical to standard CCE
- Ideal for large output spaces (e.g., vocabularies)
🧮 Use Cases
| Application | Why Sparse CCE? |
|---|---|
| Multiclass classification | Efficient label format |
| NLP (language modeling) | Large vocab class index |
| Image classification | Avoids one-hot expansion |
| Transformer decoding | Index-based softmax loss |
🔄 Behavior
- Same gradient and output as categorical cross-entropy
- Input: logits or softmaxed probabilities
- Target: class indices like
tensor([3, 0, 1])
⚠️ Considerations
| Pro | Con |
|---|---|
| No need for one-hot labels | Can't handle multilabel targets |
| More memory efficient | Still needs softmax/logits |
| Same as CCE under the hood | Requires integer class labels |
🔬 Visualization Tip
- Show how label index directly indexes into the softmax vector
- Compare memory cost of one-hot vs sparse encoding
📦 PyTorch Example
import torch.nn as nn
loss_fn = nn.CrossEntropyLoss() # logits in, class index out
loss = loss_fn(logits, class_indices) # class_indices: LongTensor of shape [batch]
📊 Summary: CCE vs Sparse CCE
| Feature | CCE | Sparse CCE |
|---|---|---|
| Target | One-hot vector | Class index |
| Memory | Higher | Lower |
| Ease of use | More prep needed | Simpler label input |
| Best for | Small to mid class count | Large-class problems |
📚 Related Topics
📐 KL Divergence
“How much extra info is needed to use Q instead of the true P?”
📍 Definition
Given two probability distributions \( P \) (true) and \( Q \) (approximate), KL divergence is:
$$ D_{KL}(P \parallel Q) = \sum_{i} P(i) \log \left( \frac{P(i)}{Q(i)} \right) $$
For continuous distributions:
$$ D_{KL}(P \parallel Q) = \int p(x) \log \left( \frac{p(x)}{q(x)} \right) dx $$
- \( P \): ground-truth distribution
- \( Q \): approximate or predicted distribution
- Not symmetric: \( D_{KL}(P \parallel Q) \ne D_{KL}(Q \parallel P) \)
🧠 Intuition
- KL tells how much information is lost when using \( Q \) to approximate \( P \)
- If \( Q \approx P \), then KL ≈ 0
- Higher KL → worse approximation
📘 Gradient
KL divergence is differentiable with respect to \( Q \), making it usable in neural nets for training distributional outputs (e.g. VAEs, language models).
🧮 Use Cases
| Application | Why KL Divergence? |
|---|---|
| Variational Autoencoders (VAEs) | Regularizes latent space against a prior |
| Language modeling | Measures divergence from expected token distribution |
| Reinforcement Learning | Bounds policy updates (TRPO, PPO) |
| Bayesian deep learning | Compares posterior and prior distributions |
🔄 Behavior
- KL \( \geq 0 \)
- KL = 0 ⇔ \( P = Q \)
- KL → ∞ if \( Q(i) = 0 \) and \( P(i) > 0 \) (division by zero risk)
⚠️ Considerations
| Pro | Con |
|---|---|
| Captures full distribution mismatch | Not symmetric |
| Differentiable (good for training) | Sensitive to small \( Q \)-values |
| Core to VAEs, RL, NLP | Requires normalized distributions |
🔬 Visualization Tip
- Overlay distributions \( P \) and \( Q \), highlight where mismatch contributes most
- Visualize KL as "cost" of substituting \( Q \) for \( P \)
📦 PyTorch Example
import torch.nn.functional as F
# Q: log probabilities (logits passed through log_softmax)
# P: true probabilities or one-hot vectors
kl = F.kl_div(Q.log(), P, reduction='batchmean')
📊 KL Summary
| Feature | KL Divergence |
|---|---|
| Type | Asymmetric distance |
| Output | Scalar \( \geq 0 \) |
| Used with | Probabilities / Softmax |
| Common in | VAEs, RL, NLP |
| Key role | Regularization / Distribution alignment |
📚 Related Topics
🧾 Negative Log-Likelihood (NLL)
“Log the chance of the truth — then flip the sign.”
📍 Definition
For classification using log-probabilities \( \log(\hat{p}) \) and true class label \( y \):
$$ \text{NLL}(\log(\hat{p}), y) = -\log(\hat{p}_y) $$
- \( \hat{p}_y \): predicted probability for true class \( y \)
- Input must be log-probabilities → typically from
log_softmax
🧠 Intuition
- Picks the log-probability for the correct class
- Negates it → confident predictions get low loss
- Mathematically equivalent to cross-entropy, when paired with softmax
🧮 Use Cases
| Application | Why NLL? |
|---|---|
| Multiclass classification | Used with precomputed log_softmax |
| Language modeling | Predicting next word/token from vocab |
| Custom probabilistic models | Explicit work in log-probability space |
| Modular pipelines | Separation of softmax and loss computation |
🔄 Behavior
- Loss → 0 as predicted log-prob of true class → 0 (i.e., prob → 1)
- Loss grows quickly as confidence in true class drops
- Stable in log-space, great for numerical safety
⚠️ Considerations
| Pro | Con |
|---|---|
| Efficient in log-space | Requires log-probs as input |
Pairs with log_softmax | Not for raw logits directly |
| Numerically stable | Can be redundant vs CrossEntropyLoss |
🔬 Visualization Tip
- Plot log-probabilities vs NLL for the correct class
- Highlight steep rise in loss as log-prob decreases
📦 PyTorch Example
import torch.nn as nn
loss_fn = nn.NLLLoss()
log_probs = torch.log_softmax(logits, dim=1)
loss = loss_fn(log_probs, targets)
📊 NLL vs Cross-Entropy
| Feature | NLL | CrossEntropy |
|---|---|---|
| Input | Log-probabilities | Raw logits |
Requires log_softmax? | ✅ | ❌ (included) |
| Numerical stability | ✅ | ✅ |
| Flexibility | Modular, manual | Simplified, automatic |