Extracting structured data from unstructured document with incomplete resources

Hervé Déjean

Extracting structured data from unstructured document with incomplete resources

2015

Hervé Déjean

We present a method for extracting structured elements of information, called structured data (sdata), from ocr'ed pages. The method first analyzes the layout of the page, building several concurrent layout structures. Then a tagging step is performed in order to tag textual elements based on their content. Combining the layout structures and the tagged elements, layout models for representing the structured data are inferred for the current page. These models are used to correct or tag some elements missed by the tagging step. The final set of structured data is extracted. An evaluation is presented.

Keywords:

Data model
Document layout analysis
Computer science
Data extraction
Information retrieval

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations