Generating an Entailment Corpus from News Headlines

John D. Burger,Lisa Ferro

Generating an Entailment Corpus from News Headlines

2005

John D. Burger
Lisa Ferro

We describe our efforts to generate a large (100,000 instance) corpus of textual entailment pairs from the lead paragraph and headline of news articles. We manually inspected a small set of news stories in order to locate the most productive source of entailments, then built an annotation interface for rapid manual evaluation of further exemplars. With this training data we built an SVM-based document classifier, which we used for corpus refinement purposes---we believe that roughly three-quarters of the resulting corpus are genuine entailment pairs. We also discuss the difficulties inherent in manual entailment judgment, and suggest ways to ameliorate some of these.

Keywords:

Computer science
Logical consequence
Artificial intelligence
Natural language processing
Textual entailment
Paragraph
Support vector machine
Small set
Classifier (linguistics)
Headline
Annotation
Training set

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations