Skip to content

MT-Bench: A Benchmark for Evaluating Language Model Instruction Following

Authors: Zheng et al. Year: 2023 ArXiv/Link: https://arxiv.org/abs/2306.05685

Summary

Benchmark for instruction following with multi-turn conversations and LLM-based evaluation.

Key Concepts

  • Instruction following
  • Multi-turn evaluation
  • LLM-based scoring
  • Conversation quality
  • Practical evaluation

Impact

Key benchmark for evaluating instruction-tuned models

Category

Benchmark Papers