Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

Pedro Rodriguez,Joe Barrow,Alexander Miserlis Hoyle,John P. Lalor,Robin Jia,Jordan Boyd-Graber

Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

2021

Pedro Rodriguez
Joe Barrow
Alexander Miserlis Hoyle
John P. Lalor
Robin Jia
Jordan Boyd-Graber

Leaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). Rather than replace leaderboards, we advocate a re-imagining so that they better highlight if and where progress is made. Building on educational testing, we create a Bayesian leaderboard model where latent subject skill and latent item difficulty predict correct responses. Using this model, we analyze the ranking reliability of leaderboards. Afterwards, we show the model can guide what to annotate, identify annotation errors, detect overfitting, and identify informative examples. We conclude with recommendations for future benchmark tasks.

Keywords:

Field (computer science)
Annotation
Artificial intelligence
simplicity
Reliability (statistics)
Benchmark (surveying)
Ranking
Computer science
Overfitting
Natural language processing
Bayesian probability

Correction
Source
Cite
Save
Machine Reading By IdeaReader

100

References

Citations