How are images displayed in a PDF file?

Images are not stored inside a PDF file as Tiff or PNG or JPG images. They are stored as the binary pixel data along with the Colorspace used by that data. This allows a lot of flexibilty. For example, a CMYK image can be stored as a block of binary data (4 bytes for each pixel) and a specified as using a CMYKColorspace. The actual image data can be compressed in different ways to best suit the data (DCT for colour images, CCITT or JBIG2 for black and white 1 bit images). The image is scaled to fit the slot of the page so it can often be of a higher resolution.

There are 2 image commands for drawing images (ID and DO). The ID command allows the binary image data to be embedded in the command stream. This is not as flexible as the DO command which stores the image in a separate PDF object of type XObject or XForm. So the DO command tells to be far more common. It allows better data compression, offers more functionality and you can edit the image object without having to alter the command stream.

Each image has a name (like Im4). In the stream, you would see the command

/Im4

which draws the image at this point with the current graphics Matrix.

The actual image IM4 is defined in a separate object which is listed in the Resources table. In this case it is Object 20 0 R.

XObject<</Im4 20 0 R/Im3 21 0 R>>

Object 20 contains the information on the image and the compressed binary pixel data

20 0 obj <<

/Filter/DCTDecode

/Type/XObject/

Length 33555/

Height 413/

BitsPerComponent 8/

ColorSpace 17 0 R/

Subtype/Image/

Width 633

stream (binary pixel data follows)

Our software libraries allow you to

Convert PDF files to HTML

Use PDF Forms in a web browser

Convert PDF Documents to an image

Work with PDF Documents in Java

Read and write HEIC and other Image formats in Java

7 Replies to “How are images displayed in a PDF file?”

Can “/Resources” point to an array of indirect object references that are composed of dictionaries? For example, if I have Font dictionary indirectly references as 20 0 R, and I have another dictionary (“<< /XObject <> >>”) referenced indirectly as 21 0 R, can my resources like like this: “/Resources [ 20 0 R 21 0 R ]”?

Sorry, my previous comment was sloppy and full of typos:

Can “/Resources” point to an array of indirect object references that are composed of dictionaries? For example, if I have a Font dictionary indirectly referenced as 20 0 R, and I have another dictionary (“<< /XObject <> >>”) referenced indirectly as 21 0 R, can my resources look like this: “/Resources [ 20 0 R 21 0 R ]”?

It cannot take an Array. You either give it the Dictionary ref or directly embed the Dictionary object

Forrest Short says:
July 14, 2017 at 8:44 pm
Thanks for the feedback, Mark!
I have another question due to the fact that I can’t get an image to actually show in my file. If I use a Java File object to open a .jpeg file and pass that object as a parameter to a new FileInputStream object, which in turn is passed as a parameter to a new BufferedInputStream object, can I use all the bytes from the BufferedInputStream’s internal byte buffer as the stream for an XObject? That’s what I’m currently doing, but Adobe Reader won’t render the image. I saved the same image as a pdf using Word, opened it up in a simple text editor, and noticed that the image stream data was half the length as mine. The length of my image stream, by the way, is the same number of bytes as the file itself.

You cannot directly embed an image in an XObject. You need to provide the raw data (with an encoding), ColorSpace and other details. When you save a Doc as PDF, this is being done for you automatically.

It is a lot simpler to add an image with a library like IText.

You might find these 2 posts helpful to understand XObjects

https://blog.idrsolutions.com/2010/04/understanding-the-pdf-file-format-how-are-images-stored/
https://blog.idrsolutions.com/2010/09/understanding-the-pdf-file-format-images/

Since both PNG and PDF formats use DEFLATE algorithms I’m wondering are these algorithms are somewhat compatible. I.e. is it possible to select PNG compression parameters which will later allow to directly copy binary data while making PDF file? Of course there is much more to do to create valid PDF file, but I would like to avoid decompressing and compressing of image data.

Mark Stephens says:
March 29, 2019 at 3:03 pm
Your problem with all these strategies is that there is no header on the data and you need to factor in the ColorSpace. You do not want to be creating CMYK PNGs for example.

Comments are closed.

How are images displayed in a PDF file?

Our software libraries allow you to

What is JPEG XL?

How to read JPEG XL images in Java

What is the AVIF image format?

7 Replies to “How are images displayed in a PDF file?”