Skip to content

LongNet: Scaling Transformers to 1M tokens

Authors: Ding et al. Year: 2023 ArXiv/Link: https://arxiv.org/abs/2307.02486

Summary

Scales transformers to million-token context lengths through efficient attention mechanisms.

Key Concepts

  • Long context
  • Attention efficiency
  • Token scaling
  • Extended sequences
  • Memory efficiency

Impact

Enabled processing of very long documents

Category

Long Context & Scaling