Dropout
Properties
Dropout is a regularization technique that randomly sets a fraction of neuron activations to zero during each forward pass of training. This prevents neurons from co-adapting and forces the network to learn redundant representations ( Srivastava et al., 2014).
# Mechanism
During training, each activation \(h_i\) is independently zeroed with probability \(p\) (the drop rate):
At inference, all neurons are active but outputs are scaled by \((1 - p)\) to match expected training values (or equivalently, training uses inverted dropout — scale by \(\frac{1}{1-p}\) at train time so no change is needed at test time).
# Ensemble Interpretation
Training with dropout approximates training an ensemble of \(2^n\) thinned networks (where \(n\) is the number of units). At test time, the full network with scaled weights approximates geometric mean of the ensemble ( Deep Learning, Ch. 7).
# Properties
- Computationally cheap: random masking adds minimal overhead
- Works well for fully-connected and recurrent layers; less effective for convolutional layers (use Batch Normalization instead)
- Typical drop rates: \(p = 0.5\) for hidden layers; \(p = 0.1\)–\(0.2\) for input/embedding layers
- Transformers use attention dropout and residual dropout in addition to MLP dropout
- Creates a form of noise injection during training, improving robustness
# Variants
- DropConnect — drops individual weights rather than activations
- Spatial Dropout — drops entire feature maps (used in CNNs)
- Monte Carlo Dropout — keep dropout active at test time to estimate prediction uncertainty