Medusa: Parallel Decoding with Multiple Heads¶
Authors: Cai et al. Year: 2024 ArXiv/Link: https://arxiv.org/abs/2401.10891
Summary¶
Multiple prediction heads enable parallel token prediction, achieving significant speedups in autoregressive generation.
Key Concepts¶
- Parallel decoding
- Multi-head prediction
- Token parallelism
- Speedup mechanisms
- Efficient generation
Impact¶
Parallelized generation for faster inference
Category¶
Decoding Strategies