LongNet: Scaling Transformers to 1M tokens¶
Authors: Ding et al. Year: 2023 ArXiv/Link: https://arxiv.org/abs/2307.02486
Summary¶
Scales transformers to million-token context lengths through efficient attention mechanisms.
Key Concepts¶
- Long context
- Attention efficiency
- Token scaling
- Extended sequences
- Memory efficiency
Impact¶
Enabled processing of very long documents
Category¶
Long Context & Scaling