SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems¶
Authors: Wang et al. Year: 2019 ArXiv/Link: https://arxiv.org/abs/1905.00537
Summary¶
Harder tasks for NLU evaluation when GLUE started to saturate, maintaining challenge as models improved.
Key Concepts¶
- Advanced NLU
- Challenging benchmarks
- Harder tasks
- Benchmark evolution
- Performance saturation
Impact¶
Extended evaluation as models saturated GLUE
Category¶
Benchmark Papers