The libraries, each with one job
None of this is a PDF engine written from scratch, and none of it is fetched from somewhere else at runtime. The libraries below are bundled with the site and served from this domain, which is why the tools keep working with the network switched off. Three of them are WebAssembly modules rather than plain JavaScript, and each of those loads on demand only when a file needs it.
- pdf-lib: rewriting the structure
- Reads a PDF's object tree so pages can be merged, reordered, rotated, split, or dropped, and metadata edited out. This is the lossless route: nothing about a page's contents is reinterpreted, so text stays text and fillable fields stay fillable.
- pdf.js (Mozilla): rendering pages
- Draws a page onto a canvas to produce pixels. Everything that needs to see the picture rather than the data goes through it: the side-by-side comparison, burning a black bar into the page image so a redaction is real, and re-encoding scans for compression.
- JSZip: packing results
- When a tool produces many files at once, say a PNG per page, this bundles them into a single archive for download.
- qpdf (compiled to WebAssembly): decryption fallback, rebuild, and encryption
- A C++ PDF toolkit compiled to a
.wasmmodule. It isn't used for ordinary files; it's fetched on demand only when the built-in handler can't read a file's encryption structure, when a file's structure is being rebuilt, or when you encrypt a file yourself. It still runs entirely in your tab, and the file never leaves the browser. - MozJPEG (compiled to WebAssembly): re-encoding scans
- The JPEG encoder behind the compressor's image route. It only comes into play once pages are being redrawn as pictures, and there it beats the browser's own encoder: a page full of photographic detail is where it earns its keep.
- tesseract (compiled to WebAssembly): reading a scan
- The recognition engine behind the text extractor's “OCR scanned pages” option. It's the largest download on the site — roughly 7 MB, fetched only when you switch that option on — and it reads printed English. Handwriting and poor scans come back imperfect, which is why the tool warns you before it runs.
That's where the “pure JavaScript” line this page used to take ended, and it's worth saying why. Not because JavaScript can't do the job — it does most of it here. It's that all three Wasm modules are mature C and C++ projects, each built for exactly the awkward cases a from-scratch implementation handles worst: damaged structures, unusual encodings, scans that were never meant to be read by a machine. Compiling them puts already-tested code on that work, instead of asking you to trust a brand-new parser with the file you least want to lose.
One file, start to finish
- You choose a file. The browser hands the page a handle to a file on your disk. Nothing has moved yet; the page can't even see the bytes without asking.
- The page reads those bytes into memory in your tab. This is the step that replaces an upload on a normal site: same local operation, no network involved.
- The relevant library parses the document there. Any error about a malformed or encrypted file comes from your own browser reading your own file.
- Your operation runs. Either the structure route rewrites the object tree, or, if it needs the picture, pages are drawn to canvases and rebuilt from those images. Both outcomes are described below, including what each one costs.
- The result becomes a Blob in that same tab and is offered as a download from a local object URL. No server is waiting for it, so there's no copy anywhere else and nothing to recover later.
Check it yourself
- Unplug. Load any tool page, switch off Wi-Fi, then run a file through it. Every operation should complete, because the dependencies are already on your machine by then.
- Watch the network panel. Open DevTools, switch to Network, tick "Preserve log", drop a file in, and work through the tool. Every request you see resolves to this domain. A request to any other domain would be the thing this site promises not to do.
- Read what shipped. Open
/sw.js. It's served like any other file and contains the complete list of everything the site caches for offline use: site assets only, pages, stylesheets, scripts. Auditing it takes no imagination, because the list is the list.
The parts that cost you something
Read this before you use a tool for something that has to survive scrutiny. These are trade-offs built into doing PDF work inside a browser, and each one shows up in the tool itself too.
- Structure edits lose nothing; rasterized pages lose their text layer
- Drawing pages as images is what makes a redaction permanent and lets a scan be compressed hard. It also means those pages stop being searchable and selectable, and that fillable form fields, comments, links, and bookmarks don't survive, because there's nothing left in the file to hold them.
- Compression runs two routes and keeps the smaller result
- The lossless route and the image route are both attempted, and you get whichever comes out smaller. On a fixed set of sample documents, the difference wasn't subtle: an image-heavy page with form fields shrank by roughly 95% and lost its eight form fields, while a twenty-page text-only sample grew by a factor of 461 on the image route and was never offered. Numbers that large are why the tool shows you both instead of picking quietly. Where keeping fields matters, there's a control to skip that route entirely.
- Very large pages get quietly scaled down
- Above roughly 24 megapixels per page, rasterizing reduces the resolution instead of failing. That keeps the browser alive, but an enormous page comes back smaller than it went in: fine for reading, wrong as an archival master.
- OCR is one tool, in one language
- Everywhere except the text extractor, the tools read text that's already in the file: a scan nobody has recognized has none, so extraction returns nothing and searching finds nothing. The extractor's “OCR scanned pages” option does recognize glyphs, but only printed English, and only as well as the scan allows — handwriting and poor scans come back wrong. Treat that output as a first pass you check, not a transcript you send.
- Encryption stays encryption
- Password-protected files open with the password you already have. Nothing here recovers one; that boundary is in the format, not a shortcut taken here.
- The ceiling is your own machine
- Everything is one tab doing the work, limited by the memory that tab can get. Work yields to the interface regularly so it never looks frozen, but thousands of pages at once is asking a lot; splitting usually gets there.
What is stored on your device
Almost nothing, but not literally nothing, so it's worth being precise: a light/dark theme preference, offline copies of the site's own files, and, when you start on one page and continue into a tool, the file's bytes handed over through local storage for at most ten minutes. Everything stays on your machine and disappears when you clear site data. The privacy policy names each one so you can go look.
