Skip to content

SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Authors: Wang et al. Year: 2019 ArXiv/Link: https://arxiv.org/abs/1905.00537

Summary

Harder tasks for NLU evaluation when GLUE started to saturate, maintaining challenge as models improved.

Key Concepts

  • Advanced NLU
  • Challenging benchmarks
  • Harder tasks
  • Benchmark evolution
  • Performance saturation

Impact

Extended evaluation as models saturated GLUE

Category

Benchmark Papers