🎯 L1 Regularization — “The Minimalist’s Guide to Learning”
"If it doesn’t spark joy (or improve prediction), throw it out."
📍 Core Idea
L1 regularization, also known as Lasso (Least Absolute Shrinkage and Selection Operator), adds the sum of absolute values of model weights as a penalty to the loss function:
\[ L = \text{Loss} + \lambda \sum_i |w_i| \]
- \( \lambda \): regularization strength — like a volume knob for punishment
- \( |w_i| \): absolute value — sharp penalties even for small deviations
🧠 Why It Works
Imagine packing for a hike. You only want to bring essentials. L1 forces your model to pack light — it drops the unimportant features entirely.
- Sparse models: many weights driven to zero
- Automatic feature selection: only the strong survive
✍️ Math Meets Geometry
Gradient:
\[ \frac{d}{dw} |w| = \begin{cases} 1 & w > 0 \\ -1 & w < 0 \\ \text{undefined (use subgradient)} & w = 0 \end{cases} \]
This leads to "dead zones" — areas where the model stops adjusting some weights. That’s how it drives weights to exactly zero.
Geometric View:
- Loss contours: ellipses
- L1 penalty shape: diamond
- Diamond corners = axes-aligned → promotes exact zeros (sparsity)
🧪 Use Cases
| Domain | Role of L1 |
|---|---|
| High-dimensional data | Kills off irrelevant features |
| Genomics, NLP, Finance | Models with thousands of variables |
| Compressed models | Leads to efficient inference |
| Interpretable ML | Easy to explain: “only these features matter” |
🧮 PyTorch Implementation
l1_lambda = 0.001
l1_norm = sum(p.abs().sum() for p in model.parameters())
loss = base_loss + l1_lambda * l1_norm
- ✔️ Works with any optimizer
- ✔️ Can be combined with L2 for Elastic Net
🎨 Visual Metaphor
Imagine sculpting a statue (the model) from a block of marble (the parameters).
L1 is the chisel: removing chunks cleanly (zeroing weights).
L2 is the sandpaper: smoothing the shape (shrinking weights).
⚠️ Limitations
| Advantage | Trade-Off |
|---|---|
| Sparse and interpretable | Non-differentiable at zero |
| Feature selection | Can be unstable with collinearity |
| Simple math | Slower convergence than L2 |
💡 Pro Tips
- Use standardization (zero mean, unit variance) before L1 — it treats all features equally
- Tune
\( \lambda \)carefully — too high = overly simplistic model - Combine with L2 to get the best of both worlds (Elastic Net)
🧲 L2 Regularization — “The Gentle Squeeze of Simplicity”
“Don't throw it away — just nudge it closer to zero.”
📍 Core Idea
L2 regularization, often called Ridge Regression, penalizes the sum of squared weights to prevent overfitting:
\[ L = \text{Loss} + \lambda \sum w_i^2 \]
- \( \lambda \): regularization strength — the "grip" that reins in complexity
- \( w_i^2 \): quadratic punishment — gentle yet firm
Unlike L1, L2 shrinks weights rather than eliminating them.
🧠 Why It Works
Think of your model like a rubber band net — L2 pulls each weight gently toward zero, keeping the net tight and controlled.
- Promotes weight stability
- Prefers many small contributions over few large ones
- Handles collinear features better than L1
✍️ Math Meets Geometry
Gradient:
\[ \frac{d}{dw} w^2 = 2w \]
- Smooth and differentiable everywhere
- Gradient gets stronger with larger weights → bigger pushback
- Encourages uniform shrinking of weights
Geometric View:
- Loss contours: ellipses
- L2 penalty shape: circle
- Solution tends to lie near the center but rarely on an axis → no exact zeros
🧪 Use Cases
| Domain | Role of L2 |
|---|---|
| Deep learning | Stabilizes large parameter spaces |
| Linear regression | Handles multicollinearity gracefully |
| Logistic regression | Improves generalization |
| Online learning | Smooth updates, avoids overcorrections |
🧮 PyTorch Implementation
L2 is built-in as weight_decay in most PyTorch optimizers:
optimizer = torch.optim.SGD(model.parameters(), lr=0.01, weight_decay=0.001)
- ✔️ Automatically applied during optimization
- ✔️ No need to modify loss function manually
🎨 Visual Metaphor
Think of L2 like a magnetized field around zero:
Weights are free to move, but always pulled softly inward
The closer to the center, the less resistance
Farther out = stronger magnetic pull
⚠️ Limitations
| Advantage | Trade-Off |
|---|---|
| Smooth convergence | Doesn’t eliminate features |
| Handles collinearity | Less interpretable |
| Efficient optimization | Can retain noisy features |
💡 Pro Tips
- Use when you expect all features to matter, but to varying degrees
- Pair with batch normalization for stable, controlled training
- Tune \( \lambda \) via cross-validation — small values often work well
🧠 L1 vs L2 — Revisited
| Aspect | L1 (Lasso) | L2 (Ridge) |
|---|---|---|
| Penalty | Absolute | Squared |
| Feature elimination | Yes | No |
| Gradient behavior | Constant | Scales with w |
| Use case | Sparse models | Stable weights |
| Geometry | Diamond | Circle |
🔗 Elastic Net — “The Balanced Negotiator Between Sparsity and Stability”
"Some features must go, others must stay small — let’s compromise."
📍 Core Idea
Elastic Net blends L1 (Lasso) and L2 (Ridge) regularization into a single penalty:
\[ L = \text{Loss} + \lambda_1 \sum |w_i| + \lambda_2 \sum w_i^2 \]
Or using mixing ratio \( \alpha \in [0,1] \):
\[ L = \text{Loss} + \lambda \left( \alpha \sum |w_i| + (1 - \alpha) \sum w_i^2 \right) \]
- \( \alpha = 1 \): pure L1 (Lasso)
- \( \alpha = 0 \): pure L2 (Ridge)
- In between: Elastic Net
🧠 Why It Works
L1 is great for sparsity, but unstable with correlated features.
L2 is stable, but retains all features.
Elastic Net combines:
- L1’s feature selection
- L2’s grouping effect (correlated features stay together)
It’s like a hiring panel:
L1 fires underperformers (weights = 0)
L2 retrains them to improve (shrinks weights)
Elastic Net does both, intelligently
✍️ Math Meets Geometry
Gradient:
\[ \frac{d}{dw} = \lambda \left( \alpha \cdot \text{sign}(w) + 2(1 - \alpha) \cdot w \right) \]
- Non-zero weights decay faster if they’re large
- Some weights are zeroed out
- Others are gently squeezed
Geometric View:
- Penalty shape: rounded diamond
- Still intersects axes → sparse
- Still smooth → stable optimization
🧪 Use Cases
| Application | Why Elastic Net? |
|---|---|
| High-dimensional models | Better than Lasso for correlated features |
| Genomics/NLP | Sparse + stable |
| Finance | Interpretability with control |
| Deep models | Manual weight control layer-wise |
🧮 PyTorch Implementation (Manual)
PyTorch does not have a built-in Elastic Net, but it’s simple to implement:
l1_lambda = 0.001
l2_lambda = 0.001
l1_norm = sum(p.abs().sum() for p in model.parameters())
l2_norm = sum(p.pow(2).sum() for p in model.parameters())
loss = base_loss + l1_lambda * l1_norm + l2_lambda * l2_norm
✔️ Can adjust \( \alpha \) via ratio of l1_lambda to l2_lambda
🎨 Visual Metaphor
Imagine training a team:
L1 says, “Fire the weak!”
L2 says, “Coach everyone to perform better.”
Elastic Net says, “Fire the worst, train the rest — we win both ways.”
⚠️ Limitations
| Pro | Con |
|---|---|
| Best of both worlds | Two hyperparameters to tune |
| Works on all feature types | Slightly more compute |
| More robust than Lasso alone | May still over-regularize |
💡 Pro Tips
- Scale your features (standardize) before applying
- Grid-search \( \lambda \) and \( \alpha \) — small changes can shift balance
- Useful as default regularizer when unsure between L1 and L2
📚 See Also
- Coordinate Descent Algorithms for Elastic Net optimization
scikit-learn’sElasticNetclass- Group Lasso — a structured variant for grouped features
- Sparse Group Lasso — even more flexible hybrid models
🌧️ Dropout — “Random Disruption for Reliable Learning”
“Train like chaos, test like a champ.”
📍 Core Idea
Dropout is a stochastic regularization technique where random units (neurons) are dropped during training:
- At each training step, neurons are independently turned off with probability \(1 - p\)
- At test time, all neurons are active, and outputs are scaled accordingly
\[ \tilde{h}_i = \begin{cases} 0 & \text{with probability } 1 - p \\ \frac{h_i}{p} & \text{with probability } p \end{cases} \]
- ✅ Keeps expected output stable
- ✅ Adds robustness to the network
🧠 Why It Works
Dropout is like ensemble learning inside one network.
- Forces redundancy — no neuron can be a “single point of failure”
- Prevents co-adaptation — neurons can’t rely on each other too heavily
- Simulates an ensemble of smaller networks → boosts generalization
Survive without your teammates — learn to be good alone.
✍️ Behind the Scenes
Forward Pass (Training):
- Apply a Bernoulli mask to hidden activations
- Scale the output to maintain expectation
Forward Pass (Testing):
- No neurons dropped
- Scale activations by \(p\) to compensate
Gradient:
- Gradients are computed only for active neurons
- Dropout acts like a sparse, randomized mask in both forward and backward passes
🧪 Use Cases
| Domain | Why Use Dropout? |
|---|---|
| Deep neural networks | Prevent overfitting in large capacity models |
| NLP (embeddings, transformers) | Combat memorization, improve generalization |
| Computer vision | Robustify against visual noise |
| Recommender systems | Helps in high-dimensional, sparse data |
🧮 PyTorch Example
import torch.nn as nn
model = nn.Sequential(
nn.Linear(256, 128),
nn.ReLU(),
nn.Dropout(p=0.5),
nn.Linear(128, 10)
)
- p = keep probability during training
- Automatically disabled during
.eval()mode
🎨 Visual Metaphor
Imagine a basketball team practicing:
Every time they train, some players are randomly benched
Others must learn to play multiple roles
Come game day: all players are active, but better prepared
⚠️ Limitations
| Advantage | Trade-Off |
|---|---|
| Easy to apply | Slows convergence (noisy updates) |
| Powerful regularizer | May underfit if overused |
| Works with SGD, Adam | Less effective with batchnorm |
💡 Pro Tips
- Start with
p = 0.5for hidden layers - Lower to
0.1–0.3in input layers (preserve input fidelity) - Combine with batch normalization carefully — use dropout after batchnorm or reconsider its necessity
- Works best with large models where overfitting is a risk
🎭 Variants & Extensions
- SpatialDropout: Drop whole feature maps (CNNs)
- DropConnect: Drop individual weights instead of activations
- Monte Carlo Dropout: Use dropout during inference to model uncertainty
- Stochastic Depth: Drop entire layers (ResNets)
📚 See Also
- Early Stopping
- L2 Regularization
- Data Augmentation
- Variational Inference (Dropout as Approximate Bayesian Inference)
🎯 Label Smoothing — “Soft Targets for Sharp Minds”
“Don’t be overconfident — be approximately right.”
📍 Core Idea
Label smoothing relaxes the traditional one-hot encoding of classification labels to prevent the model from becoming too confident.
Instead of target class = 1 and others = 0, we spread some probability mass:
\[ y_{\text{smooth}} = (1 - \epsilon) \cdot y_{\text{one-hot}} + \frac{\epsilon}{K} \]
- \(\epsilon\): smoothing factor (e.g., 0.1)
- \(K\): number of classes
- Target class gets \(1 - \epsilon\), others get \(\epsilon / K\)
🧠 Why It Works
One-hot targets train the model to assign 100% certainty — unrealistic and overfitting-prone.
Label smoothing:
- ✅ Reduces overconfidence
- ✅ Encourages generalization
- ✅ Acts like a regularizer for the output distribution
Treat labels like “soft advice” instead of rigid truth.
✍️ Math Meets Loss
Standard cross-entropy:
\[ \mathcal{L} = -\sum y_i \log(p_i) \]
With label smoothing:
\[ \mathcal{L} = -\sum \left[(1 - \epsilon) \cdot \delta_{i=y} + \frac{\epsilon}{K} \right] \log(p_i) \]
- Loss is no longer minimized by assigning probability 1 to the correct class
- Penalizes overconfident predictions, not just incorrect ones
🧪 Use Cases
| Application | Why Use Label Smoothing? |
|---|---|
| Image classification | Avoids overfitting to hard targets |
| NLP (e.g., transformers) | Improves BLEU and generalization |
| Any overconfident classifier | Improves calibration |
| Knowledge distillation | Softens teacher targets |
🧮 PyTorch Example
import torch.nn.functional as F
def label_smoothing_loss(pred, target, epsilon=0.1):
n_classes = pred.size(1)
one_hot = F.one_hot(target, n_classes).float()
smooth = one_hot * (1 - epsilon) + epsilon / n_classes
log_prob = F.log_softmax(pred, dim=1)
return -(smooth * log_prob).sum(dim=1).mean()
- ✔️ Supports batching
- ✔️ Works with logits directly
🎨 Visual Metaphor
Imagine a teacher grading an essay:
Traditional grading = 100% or 0% for each point
Label smoothing = giving partial credit to closely related ideas
This approach rewards proximity, not just perfection.
⚠️ Limitations
| Advantage | Trade-Off |
|---|---|
| Regularizes softmax outputs | Can harm performance on clean data |
| Improves calibration | May blur fine distinctions |
| Simple to implement | Harder to interpret soft targets |
💡 Pro Tips
- Typical \(\epsilon\): 0.05 to 0.2
- Helps especially with large vocabularies or imbalanced classes
- Combine with data augmentation and dropout for compounded regularization
- Monitor confidence calibration — label smoothing can reduce overconfident errors
📚 See Also
- Knowledge Distillation
- Temperature Scaling
- Confidence Calibration
- Soft Labels in Semi-Supervised Learning
- Focal Loss (focuses on hard examples — opposite flavor)
⚖️ Batch Normalization — “Stable Learning Through Internal Balance”
“Don’t just learn — learn centered and scaled.”
📍 Core Idea
Batch Normalization (BatchNorm) standardizes the inputs to each layer across a mini-batch to reduce internal covariate shift:
\[ \hat{x} = \frac{x - \mu_{\text{batch}}}{\sqrt{\sigma_{\text{batch}}^2 + \epsilon}} \]
Then apply a learned transformation:
\[ y = \gamma \cdot \hat{x} + \beta \]
- \(\mu_{\text{batch}}, \sigma_{\text{batch}}^2\): batch-wise mean and variance
- \(\gamma, \beta\): learned scale and shift parameters
- \(\epsilon\): small value for numerical stability
🧠 Why It Works
Training deep networks is challenging due to shifting distributions of activations as layers adapt. BatchNorm helps by:
- ✅ Stabilizing layer distributions
- ✅ Enabling faster and smoother training
- ✅ Reducing sensitivity to weight initialization
- ✅ Acting as implicit regularization
✍️ What Happens Step-by-Step
Training Mode
- Compute batch-wise mean and variance
- Normalize activations
- Apply learnable affine transformation: \(\gamma\), \(\beta\)
Inference Mode
Use running averages collected during training for normalization.
🧪 Use Cases
| Application | Why Use BatchNorm? |
|---|---|
| Deep CNNs (ResNet, VGG) | Faster convergence, better accuracy |
| RNNs/Transformers | Variant: LayerNorm for stable hidden states |
| GANs | Balances generator/discriminator updates |
| Any deep model | Standard practice for training stability |
🧮 PyTorch Example
import torch.nn as nn
model = nn.Sequential(
nn.Linear(256, 128),
nn.BatchNorm1d(128),
nn.ReLU(),
nn.Linear(128, 10)
)
BatchNorm1dfor dense layersBatchNorm2dfor conv layers- Switch to inference with
model.eval()
🎨 Visual Metaphor
Imagine a classroom where everyone's answers vary wildly in magnitude.
BatchNorm is like a teacher saying: “Let’s normalize everyone's answers before comparing.”
This ensures fair learning — no neuron dominates just due to scale.
⚠️ Limitations
| Advantage | Trade-Off |
|---|---|
| Enables deeper networks | Needs meaningful batch size |
| Helps generalization | Adds computation and parameters |
| Reduces internal shift | Less effective on small batches |
| Acts like regularizer | Interaction with Dropout is complex |
💡 Pro Tips
- Use after linear/conv layers, before activation
- Combine with
ReLU,GELU, etc. - With very small batches, prefer:
- GroupNorm
- LayerNorm
- Control
momentumandtrack_running_statsfor consistent inference
🔄 BatchNorm vs Alternatives
| Normalization | Applies To | Use Case |
|---|---|---|
| BatchNorm | Mini-batches | Standard in CNNs, MLPs |
| LayerNorm | Per sample | RNNs, transformers |
| GroupNorm | Channel groups | Small-batch CNNs |
| InstanceNorm | Per feature map | Style transfer, vision |
📚 See Also
- Layer Normalization
- Group Normalization
- Weight Initialization
- Learning Rate Scheduling
- Residual Networks (BN is standard in ResNets)
⏳ Early Stopping — “Quit While You’re Ahead”
“Better to stop training than to start overfitting.”
📍 Core Idea
Early Stopping halts training when validation performance stops improving. It observes:
- 🟢 Training & validation loss both drop initially
- 🔴 Validation loss rises while training loss falls → overfitting
⛔ Stop at the inflection point, before the model memorizes noise.
🧠 Why It Works
Instead of changing the loss or weights, it uses a held-out validation set to judge generalization. This prevents memorizing training noise and helps pick a simpler, better-generalizing model.
Like baking a cake — take it out when it’s done, not overcooked.
✍️ Strategy Options
- Monitor: loss, accuracy, F1, etc.
- Mode: ‘min’ for loss, ‘max’ for metrics
- Patience: how many epochs to wait before stopping
- Restore: optionally revert to best checkpoint
🧪 Use Cases
| Scenario | Why Early Stop? |
|---|---|
| Deep models | Prevents overfitting |
| Limited data | Stops memorization |
| Training cost constraints | Saves time/resources |
| Hyperparameter search | Captures optimal window |
🧮 PyTorch Example
import copy
best_loss = float('inf')
patience, counter = 5, 0
for epoch in range(epochs):
train(...)
val_loss = validate(...)
if val_loss < best_loss:
best_loss = val_loss
best_model = copy.deepcopy(model.state_dict())
counter = 0
else:
counter += 1
if counter >= patience:
print("Early stopping triggered.")
model.load_state_dict(best_model)
break
- ✔️ Tracks the best weights
- ✔️ Adjustable patience for flexibility
🎨 Visual Metaphor
Imagine training as hiking a mountain:
You ascend (training accuracy up), then begin slipping on the other side (validation falls).
Early Stopping = smart camper who stops at the peak before descending.
⚠️ Limitations
| Advantage | Trade-Off |
|---|---|
| Prevents overfitting | Needs validation data |
| Saves training time | Can stop prematurely |
| Simple to use | Requires manual tuning |
💡 Pro Tips
- Visualize training/validation loss curves
- Use checkpointing to keep best model
- Pair with learning rate schedulers
- Adapt
patiencebased on model & data complexity
🧠 Compared to Other Regularizers
| Technique | Type | Role |
|---|---|---|
| Dropout | Structural | Adds randomness to network |
| L2 Penalty | Penalty-based | Weight shrinkage |
| Label Smoothing | Output-level | Softens target class |
| Early Stopping | Process-based | Stops training early |
📚 See Also
- Learning Rate Scheduling
- Model Checkpointing
- Validation Curves
- Generalization Theory
- Cross-Validation
🪂 Stochastic Depth — “Learning to Skip to Go Faster and Generalize Better”
“Sometimes, skipping a few steps makes you stronger.”
📍 Core Idea
Stochastic Depth randomly skips full layers or residual blocks during training. It modifies the residual connection like this:
Standard Residual: $$ \text{Output} = x + f(x) $$ With Stochastic Depth: $$ \text{Output} = x + b \cdot f(x) $$
- $b \sim \text{Bernoulli}(p)$: random binary mask
- $p$: survival probability (higher for shallow layers)
🧠 Why It Works
- Trains an implicit ensemble of shallower networks
- Improves generalization and training speed
- Acts like Dropout — but for blocks, not neurons
Think: “What doesn’t kill a layer makes the model stronger.”
✍️ How It’s Used
- Used in ResNets, ResNeXts, ViTs, EfficientNets
- Survival probability often increases with layer depth
- All layers active during inference (outputs scaled appropriately)
🧪 Use Cases
| Architecture | Why Use It? |
|---|---|
| ResNet/ResNeXt | Combat overfitting in very deep networks |
| EfficientNet | Scalable regularization for small models |
| Vision Transformers | Drop full transformer blocks safely |
| NAS models | Encourages robustness and stability |
🧮 PyTorch Implementation
import torch
import torch.nn as nn
class StochasticBlock(nn.Module):
def __init__(self, block, survival_prob):
super().__init__()
self.block = block
self.survival_prob = survival_prob
def forward(self, x):
if self.training and torch.rand(1).item() > self.survival_prob:
return x # skip residual block
return x + self.block(x) / self.survival_prob
🎨 Visual Metaphor
A relay race where:
- 🏃 During training, runners (layers) might randomly skip turns
- 🏁 At test time, all runners participate, optimized by experience
⚠️ Limitations
| Advantage | Trade-Off |
|---|---|
| Reduces overfitting in deep nets | Not effective in shallow nets |
| Faster training, fewer parameters trained | Randomness may affect convergence |
| Acts like a block-wise ensemble | Needs tuned survival probabilities |
💡 Pro Tips
- Use higher survival probs for shallow blocks
- Great with
Dropout,BatchNorm, andLabel Smoothing - Try in architectures with >50 layers or transformer stacks
🔁 Compared to Other Techniques
| Method | What It Drops | Best For |
|---|---|---|
| Dropout | Neurons | General regularization |
| DropConnect | Individual weights | Fully connected layers |
| Stochastic Depth | Residual blocks | Deep CNNs / ViTs |
| LayerDrop | Transformer layers | Sequence models |
📚 See Also
- Dropout, DropPath
- ResNet, EfficientNet, ResNeXt
- ShakeDrop, Shake-Shake regularization
- LayerScale for ViTs
🧬 Mixup — “Blend the Data, Blend the Decision”
“If the world is noisy and uncertain, train your model to see gradients, not edges.”
📍 Core Idea
Mixup creates new samples by linearly interpolating between random pairs of training inputs and labels:
$$ \tilde{x} = \lambda x_i + (1 - \lambda) x_j $$ $$ \tilde{y} = \lambda y_i + (1 - \lambda) y_j $$
- $\lambda \sim \text{Beta}(\alpha, \alpha)$ — usually $\alpha \in [0.1, 0.4]$
- Creates smooth, soft-labeled samples
🧠 Why It Works
- Promotes linearity in learned space
- Blurs class boundaries for better generalization
- Defends against adversarial examples by making gradients smoother
- Reduces overfitting on small or noisy datasets
✍️ Training Dynamics
- Loss is computed using interpolated labels
- Can use
CrossEntropywith soft targets (if supported) - Effective in both image and embedding space
🧪 Use Cases
| Application | Why Use Mixup? |
|---|---|
| Image classification | Smoother decision boundaries |
| NLP embeddings | Data augmentation in vector space |
| Audio/speech | Regularizes spectrogram predictions |
| Adversarial defense | Reduces gradient sensitivity |
🧮 PyTorch Example
def mixup_data(x, y, alpha=0.4):
lam = torch.distributions.Beta(alpha, alpha).sample().item()
index = torch.randperm(x.size(0))
mixed_x = lam * x + (1 - lam) * x[index]
y_a, y_b = y, y[index]
return mixed_x, y_a, y_b, lam
def mixup_loss(loss_fn, pred, y_a, y_b, lam):
return lam * loss_fn(pred, y_a) + (1 - lam) * loss_fn(pred, y_b)
🎨 Visual Metaphor
Think of mixing paints: red + blue = purple. Mixup teaches the model that:
- Classes are not isolated — they're continuous blends
- Learning happens between the data, not just on it
⚠️ Limitations
| Advantage | Trade-Off |
|---|---|
| Great generalization | Soft targets may confuse early |
| Robust to noise | Less effective for non-continuous data |
| Simple to add | Can slow initial convergence |
💡 Pro Tips
- Start with $\alpha = 0.2$ for general tasks
- Combine with CutMix or Cutout in vision
- Apply
input normalizationafter mixing - Use with SGD+momentum for best results
🔁 Related Techniques
| Method | Type | Description |
|---|---|---|
| Mixup | Interpolation | Blend samples and targets |
| CutMix | Mask fusion | Mix regions from different inputs |
| Manifold Mixup | Feature-space mix | Mix hidden representations |
| Label Smoothing | Output smoothing | Encourages soft targets |
📚 See Also
- CutMix, Manifold Mixup
- Vicinal Risk Minimization
- Adversarial & Virtual Adversarial Training