Skip to content
100% local · 0 KB uploadedCompress PDF

How PDFs Store Text and Images (Under the Hood)

A PDF is a text file with binary blobs inside. The structure is surprisingly readable: a header, a list of numbered objects, a cross-reference table, and a trailer. This page walks through that structure and shows where text, fonts, and images actually live. Written for developers and anyone who wants to understand what a PDF tool is doing when it reads or writes a file.

Overview

The four parts of a PDF

Every PDF file has the same four-part layout:

  1. Header — one line, e.g. %PDF-1.7, declaring the version.
  2. Body — a sequence of numbered indirect objects. This is where everything lives: pages, fonts, images, metadata.
  3. Cross-reference table (xref) — byte offsets of every object, so a reader can find them without scanning the whole file.
  4. Trailer — points to the root object (the document catalog), the xref location, and optional encryption info. The reader reads the trailer first, then jumps to the root.

The trailer-first design is why a PDF can be read in constant time without loading the whole file — important for huge documents and streaming readers.

Body

Indirect objects

Each object in the body is numbered and looks like this:

3 0 obj
<< /Type /Page /MediaBox [0 0 612 792] /Contents 4 0 R /Resources << /Font << /F1 5 0 R >> >> >>
endobj

The 3 0 obj says "this is object 3, generation 0". The body inside the << ... >> is a dictionary — a list of key/value pairs. Values can be numbers, strings, names (prefixed with /), arrays, other dictionaries, or references like 4 0 R ("go look at object 4"). This is how a page points to its content stream and its fonts without inlining them.

The generation number is usually 0. It increments when an incremental update deletes and rewrites an object — a feature that lets PDFs be edited by appending changes to the end of the file without rewriting everything. Most modern writers rewrite the whole file instead, so generations stay at 0.

Content

How text is stored

Page content lives in a content stream — an object whose value is a stream of bytes (often Flate-compressed) containing PDF drawing operators. Text is drawn with operators likeBT (begin text), Tf (set font and size), Td (move position), Tj (show text), and ET (end text). A simple "Hello" looks roughly like:

BT
/F1 24 Tf
100 700 Td
(Hello) Tj
ET

The string (Hello) uses the font's encoding. For simple Latin text that is often WinAnsiEncoding or the font's built-in encoding. For Unicode text, the string is a 16-bit-encoded byte string and the font has a/ToUnicode mapping that lets readers recover the actual code points for copy-paste and search. If a PDF has garbled copy-paste but visible text, the /ToUnicodemap is missing or wrong — a common export bug.

Text is not stored as HTML or paragraphs. There is no semantic "paragraph" object. Lines, line breaks, and positioning are all explicit drawing commands. This is why reflowing a PDF to a different page width is hard: the reader would have to reverse-engineer the layout from drawing commands.

Fonts

Font embedding

A font object in PDF has two parts: a font dictionarythat names the font and points at its data, and the font program itself (Type 1, TrueType, or OpenType) embedded as a stream. The PDF spec requires fonts to be embedded for the document to render identically everywhere — but does not enforce it, which is why some PDFs substitute fonts on machines that don't have them.

Subsetting embeds only the glyphs actually used in the document. A 250 KB font used for a single heading becomes an 8 KB subset. Most modern exporters subset by default. Older exporters or certain "Print to PDF" paths sometimes embed the full font, which is a common cause of bloated files. See ourfile size guidefor the full breakdown.

Images

How images are stored

Images are stored as XObjects (external objects) — streams of raw or compressed pixel data with a dictionary describing width, height, color space, bits per component, and filter. The page content stream then draws the image with the Do operator.

PDF supports several image filters (compression):

  • DCTDecode — JPEG. Lossy, good for photos. The most common source of size in image-heavy PDFs.
  • FlateDecode — zlib/deflate. Lossless, good for line art and images with few colors.
  • JPXDecode — JPEG2000. Better compression than JPEG at high quality, but slower and less universally supported.
  • JBIG2Decode — for bitonal (1-bit) scanned text. Can shrink scanned pages dramatically.
  • CCITTFaxDecode — classic fax compression for bitonal images.

An image-heavy PDF is mostly a sequence of compressed image streams. Re-encoding those streams with a more aggressive JPEG quality is the single biggest size lever for scanned documents.

Index

xref table and trailer

The xref table lists the byte offset of every indirect object. A reader opens the file, reads the trailer (which is at a known offset from the end of the file), finds the xref, and uses it to jump directly to any object without parsing the whole file. This is what makes PDF random-access.

PDF 1.5 added cross-reference streams, which compress the xref table itself. Combined with object streams (many small objects packed into one compressed stream), this is the structural compression that makes modern PDFs smaller than their 1.4 equivalents. See ourcompression explainer for what this does in practice.

The trailer also points to the/Root (the document catalog, which lists the pages tree), the /Info dictionary (Title, Author, CreationDate — the metadata), and optionally the/Encrypt dictionary if the file is encrypted.

Compression

Streams and filters

Any object whose value is binary data (a content stream, an image, an embedded font) is a stream: a dictionary followed by stream, raw bytes, andendstream. The dictionary's /Filterentry tells the reader how to decode the bytes —FlateDecode for zlib, DCTDecode for JPEG, and so on. Multiple filters can be chained, e.g. an image that is decoded then downsampled.

This is why "compressing a PDF" is not one operation. Each stream is independently compressed with its own filter. A compress tool that re-saves with object streams shrinks the structural overhead; re-encoding images with a more aggressive JPEG quality shrinks the image streams. They are independent levers.

In practice

Why this matters for PDF tools

When a tool like IXPDF merges, splits, or rotates a PDF, it parses the xref, walks the pages tree, copies or rearranges the page objects, and rewrites the xref and trailer. The content streams (text and images) are copied byte-for-byte — no re-encoding, no quality loss. This is why structural operations are fast and lossless: they shuffle object references, they do not re-render the page.

Compression is the exception: it re-serializes the structure with object streams. Even then, the image and font streams are copied as-is unless the user explicitly asks to re-encode them.

All of this happens in your browser via a Web Worker — the file is parsed, modified, and re-serialized locally. Your PDF never leaves your device.

Keep reading

Related resources