Search

Search IconIcon to open search

Adamw

Last updatedUpdated: by Jakub Žovák · 2 min read

Properties
created 09.03.2026, 10:00
modified 06.09.2026, 10:24
published Empty
topics Optimizers, Weight Decay
authors Jakub
ai-assisted Yes
  • AdamW is a variant of Adam that decouples weight decay from the gradient-based update. Introduced by Loshchilov & Hutter (2019).

# Motivation: L2 Reg ≠ Weight Decay in Adam

  • In SGD, adding L2 regularization (\(\frac{\lambda}{2}\|\theta\|^2\) to the loss) is mathematically equivalent to applying weight decay directly to the weights
  • In Adam this equivalence breaks: the adaptive scaling by \(\hat{v}_t\) modifies the effective magnitude of L2 regularization per parameter, so parameters with large gradients receive less regularization than parameters with small gradients
  • This means Adam + L2 does not regularize uniformly, leading to suboptimal generalization

# Decoupled Weight Decay

AdamW separates the weight decay step from the gradient update:

$$ \theta_t = \theta_{t-1} - \alpha \left( \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon} + \lambda\, \theta_{t-1} \right) $$

where:

  • \(\theta_t\) is the parameter vector being updated
  • \(\alpha\) is the learning rate
  • \(\hat{m}_t\) is the bias-corrected first moment (mean of gradients) from Adam
  • \(\hat{v}_t\) is the bias-corrected second moment (mean of squared gradients) from Adam
  • \(\epsilon\) is a small constant for numerical stability (avoids division by zero)
  • \(\lambda\) is the weight decay coefficient applied directly to the weights, independent of the adaptive gradient scaling

# Difference from Adam

AspectAdamAdamW
RegularizationL2 added to loss (gradient)Weight decay applied directly to weights
Effective decayScaled by \(1/\sqrt{\hat{v}_t}\) per paramUniform \(\lambda\) per param
GeneralizationWeakerStronger

# Usage

  • AdamW is the default optimizer for most modern LLM training (GPT, BERT, LLaMA)
  • It is also used as the baseline optimizer in hybrid setups like Muon, which applies AdamW to embedding and normalization parameters