Generative Data Augmentation for Commonsense Reasoning

Yiben Yang,Chaitanya Malaviya,Jared Fernandez,Swabha Swayamdipta,Ronan Le Bras,Ji-Ping Wang,Chandra Bhagavatula,Yejin Choi,Doug Downey

Generative Data Augmentation for Commonsense Reasoning

2020

Recent advances in commonsense reasoning depend on large-scale human-annotated training sets to achieve peak performance. However, manual curation of training sets is expensive and has been shown to introduce annotation artifacts that neural models can readily exploit and overfit to. We propose a novel generative data augmentation technique, G-DAUGˆC, that aims to achieve more accurate and robust learning in a low-resource setting. Our approach generates synthetic examples using pretrained language models and selects the most informative and diverse set of examples for data augmentation. On experiments with multiple commonsense reasoning benchmarks, G-DAUGˆC consistently outperforms existing data augmentation methods based on back-translation, establishing a new state-of-the-art on WinoGrande, CODAH, and CommonsenseQA, as well as enhances out-of-distribution generalization, proving to be robust against adversaries or perturbations. Our analysis demonstrates that G-DAUGˆC produces a diverse set of fluent training examples, and that its selection and training approaches are important for performance.

Keywords:

Correction
Cite
Save
Machine Reading By IdeaReader

References

Citations