How to extract Structured text from PDF files in Java (Tutorial)

Developers hoping to extract content from PDF documents whilst maintaining the structure of the text should follow this tutorial. Some (but not all) PDF files contain text content which can be extracted in a structured format, retaining paragraphs and other layout and formatting information.

How to extract Structured text from PDF files in Java

Download JPedal trial jar.
Choose output format
Create a File handle, InputStream or URL pointing to the PDF file
Include a password if file password protected
Open the PDF file
Extract the Document text
Close the PDF file

Java code to extract Structured Text…

ExtractStructuredTextProperties properties = new ExtractStructuredTextProperties();

 properties.setFileOutputMode(OutputModes.XML);

 //properties.setFileOutputMode(OutputModes.HTML);

 ExtractStructuredText extract = new ExtractStructuredText("C:/pdfs/mypdf.pdf", properties);

 //extract.setPassword("password");

 if (extract.openPDFFile()) {

     Document anyStructuredText = extract.getStructuredTextContent();

 }


 extract.closePDFfile();

How do I know if a PDF file contains Structured text?

You can find out if it is present by reading is this blog post.

What is Structured text?

When Adobe created the PDF file format it was designed as an end file format, not one for editing and reusing. It works like a vector graphics file not a text document – so it contains ‘draw’ commands for images, text and shapes not any details of structures – there are no styles, line or paragraph markers or even spaces. It looks perfect but the structure is added by your brain looking at the display – there is nothing in the file.

It turned out that lots of people wanted to extract text from PDF files and were very disappointed by what they got back. So Adobe added some additional functionality into the spec so that you could add extra metadata into the file to preserve all this information and easily retrieve it. This is called Marked Content and the results are very good, but it needs to be added into the PDF when it is created.