Client-Side PDF Processing: How Browser-Local PDF Tools Work
A client-side PDF tool does the same work a server would do, but inside the user's browser. This page is for developers and IT admins who want to understand the architecture: which browser APIs are involved, which libraries do the heavy lifting, where the file actually goes, and what the honest limits are. IXPDF is the running example, but the model applies to any browser-local PDF tool.
The pipeline
From file input to download, with no server
The end-to-end flow for a client-side PDF operation is:
- File input. The user drops a file into a dropzone. The browser hands the page a
Fileobject (a pointer to the file on disk, not the bytes themselves). - Read into memory. The page calls
file.arrayBuffer()to read the bytes into anArrayBufferin the page's memory. This is the only point where the file's bytes exist in the page; they are not yet in any worker. - Validate. The page checks the file type (magic bytes, extension) and size before doing any work. Invalid files fail fast on the main thread and never reach the worker.
- Transfer to a Web Worker. The page posts the
ArrayBufferto a worker, using a transferable transfer so the bytes move (not copy) to the worker's heap. The main thread no longer has access to the buffer. - Process in the worker. The worker calls a PDF library (pdf-lib, PDF.js, or WASM-compiled qpdf/mupdf) to parse and modify the PDF. This is the heavy work and it runs off the main thread so the UI does not freeze.
- Return the result. The worker produces a new
ArrayBuffer(orBlob) and posts it back to the main thread, again as a transferable. - Offer as download. The main thread wraps the result in a
Blob, callsURL.createObjectURL(blob), and offers it as a download via an<a download>element. The object URL is revoked after the download starts to avoid a memory leak.
At no point in this flow does the file leave the device. There is no fetch, no XMLHttpRequest, no sendBeacon with file content. The only network requests the page makes are for static assets (HTML, JS, CSS, fonts) when the page first loads.
Building blocks
The browser APIs and libraries involved
A client-side PDF tool stands on a small set of browser APIs and a small set of JavaScript libraries. Each does one job.
- File API — the
FileandBlobinterfaces.file.arrayBuffer()reads bytes into memory;new Blob([bytes], { type: 'application/pdf' })packages bytes for download. - Web Workers — a background thread with its own heap, no DOM access. The PDF parsing runs here so the main thread stays responsive. Communication is via
postMessage;ArrayBuffercan be transferred (moved, not copied) using the second argument. - URL.createObjectURL / URL.revokeObjectURL — produces a temporary
blob:URL that points at the in-memory result, so an<a download>element can offer it as a download. The URL must be revoked or it leaks. - pdf-lib — a pure-JavaScript PDF library that can create, modify, and re-save PDFs. Used for merge, split, rotate, watermark, page operations, and metadata edits. Does not support encryption (a known limit — see below).
- PDF.js — Mozilla's PDF renderer, used for rendering pages to canvas (for PDF-to-image conversion) and for parsing PDFs that pdf-lib cannot read.
- WASM (optional) — for libraries that compile native code to run in the browser at near-native speed. qpdf and mupdf can be compiled to WASM; IXPDF uses this path selectively for operations pdf-lib cannot do.
- CompressionStream — the native browser API for streaming compression (gzip, deflate). Useful for some size-reduction paths without shipping a JS compression library.
Why a worker
Why PDF parsing does not run on the main thread
PDF parsing is CPU-heavy and synchronous in shape — you read bytes, decode objects, follow cross-references, build an in-memory tree. A 50 MB PDF can take hundreds of milliseconds to parse, and a 200 MB PDF can take seconds. If that work runs on the main thread, the page is frozen for the duration: no animation, no scroll, no input. The browser may show a "page is not responding" dialog.
Moving the work to a Web Worker solves this. The main thread stays responsive; the worker chews on the PDF in the background; the page shows a progress indicator and updates the UI when the worker posts a result. The trade-off is that the worker has no DOM access, so any UI work has to happen on the main thread after the result comes back.
Workers also have their own memory budget. A large PDF transferred to a worker lives in the worker's heap, not the page's. This matters for very large files: the page can stay light while the worker holds the heavy buffer, and the buffer is freed when the worker is done.
Transferable objects
Moving bytes, not copying them
When an ArrayBuffer is posted to a worker without transfer, the bytes are copied — the worker gets its own copy, and the main thread keeps the original. For a 100 MB PDF that means 200 MB of memory used and a noticeable copy time.
The fix is the transfer list: the second argument to postMessage. Listing the buffer in the transfer list moves the bytes to the worker — the main thread's reference is detached (zero length), and the worker takes ownership. No copy, no doubled memory. The same trick is used to send the result back.
This is a small detail that matters a lot at scale. Without transferables, client-side PDF processing of large files would be impractical.
Honest limits
What a browser cannot do well
Client-side processing is not a silver bullet. Some operations are hard or impossible to do well in a browser today, and an honest tool says so rather than faking them.
- OCR. Optical character recognition on scanned PDFs needs a trained model and serious CPU. Browser OCR exists (Tesseract.js) but quality is well behind desktop Tesseract or cloud OCR. IXPDF marks OCR as Research Required.
- PDF encryption (password protection). pdf-lib does not yet support writing encrypted PDFs. Adding it requires either a different library or WASM-compiled qpdf. IXPDF's Protect PDF tool is Research Required.
- PDF to Word/Excel/PowerPoint. Converting a PDF to an editable Office file is a layout-reconstruction problem that needs a backend. Browser libraries that attempt it produce low-fidelity output. IXPDF does not offer these conversions.
- AI summarization / Q&A. Needs a model and a backend. Cannot be done client-side at acceptable quality in 2026. IXPDF does not offer AI features.
- Very large files. A browser tab has a memory budget (typically a few GB). A 2 GB PDF will OOM in most browsers. Server-side tools handle this with streaming; browser tools cannot stream a PDF the same way.
The honest move is to mark these as Planned,Research Required, or Future Backend Candidatein the UI, and not ship a fake tool that returns empty output and pretends it succeeded. This is the line between a local-first tool and a content-farm "tool" page.
Why it matters
The privacy property is structural, not a promise
The reason client-side processing is a meaningful privacy property — not just a marketing claim — is that the no-upload guarantee is enforced by the structure of the code, not by a policy. There is no server to upload to. The page has no fetch call that carries file content. A user can open DevTools, go to the Network tab, run an operation, and verify that no request carries the file. The guarantee is auditable.
This is different from a server-side tool that promisesto delete your file after 24 hours. That promise is enforced by policy and infrastructure you cannot see. Client-side processing is enforced by the absence of a server. For sensitive files, this is the difference that matters.
Read more in the IXPDF privacy pageand about IXPDF.
Keep reading