Neural Networks Components
Properties
created
30.09.2024, 15:08
modified
26.06.2026, 09:49
published
Empty
sources
Empty
topics
Neural Networks Components, Attention, Normalization, Residual Blocks
authors
Empty
ai-assisted
No
- Residual Blocks
- Self-Attention
- Layer Normalization
- Gated Linear Unit
- MiniMax Sparse Attention
- Lightweight block indexer selects the most relevant KV blocks so the model attends only to them — 14x faster prefill, 7x faster decode at 109B scale