Extraction of text content from PDF documents based on automaton theory

2012 
The existing methods of extracting text content from a PDF file,such as the one adopted by the PDFBox library,are not efficient enough to handle the high-speed network traffic.Moreover,these methods cannot extract the contents streamingly from partial PDF packets in transfer.This paper proposed a new method based on automaton theory.The method adopted a hierarchical keyword Deterministic Finite Automaton(DFA) to extract information from complete or incomplete PDF files.The experimental results show that the response time of the proposed method is about 17%-37% of the algorithm used by PDFBox when processing PDF files in Chinese or English.
    • Correction
    • Source
    • Cite
    • Save
    • Machine Reading By IdeaReader
    0
    References
    0
    Citations
    NaN
    KQI
    []