Cross-Entropy
Properties
created
30.09.2024, 14:36
modified
26.07.2026, 13:36
published
Empty
sources
ChatGPT
topics
Loss Functions
authors
Jakub
ai-assisted
Yes
- Cross-entropy measures how different two probability distributions are
- Unlike Entropy which measures the uncertainty inside a single probability distribution
- Use for multi-class classification tasks i.e. labels are dependent. Because in multi-class classification, each sample belongs to exactly one class from a set of mutually exclusive classes (e.g. image with cat can not contain a dog).
$$ CCE = - \sum_{i=1}^{n} y_i \log(\hat{y}_i) $$
- \(y_i\) = the true label for class \(i\).
- Typically encoded as a one-hot vector (so \(y_i = 1\) for the correct class, and \(0\) otherwise).
- \(\hat{y}_i\) = the predicted probability that the model assigns to class \(i\).
- These are the outputs of a softmax layer, i.e. \(\hat{y}_i \in [0,1]\) and \(\sum_i \hat{y}_i = 1\).
- So in words:
- \(y_i\) is the ground-truth distribution.
- \(\hat{y}_i\) is the model’s predicted distribution.
- The loss penalizes the model if it assigns low probability to the true class.
# Sources
# Intuition
# Why Cross-Entropy and Not MSE?
- MSE treats all errors as equally costly in a quadratic sense — it’s designed for continuous outputs where the distance between prediction and target is meaningful
- In classification, we output probabilities, not continuous values, so the geometry is fundamentally different
- Vanishing gradients with MSE + softmax:
- When the predicted probability is far from the true label (model is very wrong), the sigmoid/softmax saturates and MSE produces near-zero gradients → training stalls
- Cross-entropy avoids this: its gradient with respect to the logit simplifies to just \(\hat{y}_i - y_i\), which is large when the model is confidently wrong
- Cross-entropy directly measures probabilistic mismatch, making it the natural fit for the softmax output layer
# Why Logarithm?
- Directly maximizing predicted probability \(\hat{y}_c\) (where \(c\) is the correct class) would work in theory, but log turns that into a sum, which is numerically stable and mathematically convenient
- \(\log\) heavily penalizes confidently wrong predictions: if the model assigns \(\hat{y}_c = 0.01\) to the correct class, \(-\log(0.01) = 4.6\); if it assigns \(0.99\), \(-\log(0.99) \approx 0.01\)
- The sharp penalty forces the model to be calibrated — not just right, but confidently right
# Information-Theoretic View
- From information theory: cross-entropy \(H(p, q) = -\sum_i p_i \log q_i\) is the expected number of bits needed to encode events from distribution \(p\) using a code optimized for \(q\)
- Minimizing cross-entropy = minimizing the “extra cost” of using the model’s predicted distribution \(\hat{y}\) instead of the true distribution \(y\)
- This equals KL Divergence \(D_{KL}(y \| \hat{y})\) up to a constant (the entropy of the true labels), so minimizing cross-entropy also minimizes the KL divergence between the true and predicted distributions
# Connection to Maximum Likelihood Estimation
- Cross-entropy loss is equivalent to negative log-likelihood under a categorical distribution
- Minimizing the average cross-entropy over all training examples = maximizing the likelihood that the model would have generated the observed labels
$$ \mathcal{L} = -\frac{1}{N}\sum_{n=1}^{N} \log \hat{y}_{c_n} $$
where \(c_n\) is the correct class index for sample \(n\) — this is just the log-probability the model assigned to the right answer, averaged across the dataset