Extraction of text content from PDF documents based on automaton theory
2012
The existing methods of extracting text content from a PDF file,such as the one adopted by the PDFBox library,are not efficient enough to handle the high-speed network traffic.Moreover,these methods cannot extract the contents streamingly from partial PDF packets in transfer.This paper proposed a new method based on automaton theory.The method adopted a hierarchical keyword Deterministic Finite Automaton(DFA) to extract information from complete or incomplete PDF files.The experimental results show that the response time of the proposed method is about 17%-37% of the algorithm used by PDFBox when processing PDF files in Chinese or English.
Keywords:
- Correction
- Source
- Cite
- Save
- Machine Reading By IdeaReader
0
References
0
Citations
NaN
KQI