Extraction of text content from PDF documents based on automaton theory

Liu Jin-gang

Extraction of text content from PDF documents based on automaton theory

2012

Liu Jin-gang

The existing methods of extracting text content from a PDF file,such as the one adopted by the PDFBox library,are not efficient enough to handle the high-speed network traffic.Moreover,these methods cannot extract the contents streamingly from partial PDF packets in transfer.This paper proposed a new method based on automaton theory.The method adopted a hierarchical keyword Deterministic Finite Automaton(DFA) to extract information from complete or incomplete PDF files.The experimental results show that the response time of the proposed method is about 17%-37% of the algorithm used by PDFBox when processing PDF files in Chinese or English.

Keywords:

Computer science
Deterministic finite automaton
Response time
Network packet
Theoretical computer science
Automaton

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations