A PDF looks like a fixed page, but the file behind it keeps a record of how it was made. The document properties name the author and the program that made the file and give the exact dates. A second copy of the same facts, written as XML, sits in a separate block and carries a unique ID for the document. Comments carry the name of whoever wrote them. A photo placed on a page often still has the GPS position where it was taken. A PDF that was saved several times can even contain the text of earlier versions, which stays in the file after it disappears from the page.
This page opens your PDF in your browser and lists everything of this kind, in plain language and by risk. Then it rewrites the file with only what the pages need, reads the result back, and checks that every page is the same and none of the listed data is left. Nothing is uploaded, and the code behind the page makes no network requests. It is meant for the moment before you send a contract, a CV, a report, or a scan to someone outside your organization.
How to use the Remove Metadata from PDF: Author, Software, and Hidden Data
- Add your PDFDrop one or more PDF files on the box, or choose them from your device. Each file is read on your device and nothing leaves it. A file that is protected with a password has to be unlocked first, for example with the Unlock PDF tool on this site.
- Read the reportThe page says what the file reveals and groups it: document properties, XMP metadata, document ID, earlier versions, comment authors, photo data, attached files, and more. Each group says whether cleaning removes it.
- Check the settingsAuthor, software, dates, ID, thumbnails, and earlier versions are always removed. Names on comments and data inside photos are removed by default. Attached files and scripts stay unless you switch them on.
- Remove the hidden dataPress the button. The page rewrites the file, opens the result again, scans it, and compares it with the original. You see what was removed and the checks that passed.
- Download the clean copyDownload the file, or all files as a ZIP. The clean copy has the same pages. Keep your original, because the clean copy is a new file, and a digital signature on the original will not be valid on it.
Where a PDF keeps its metadata
The PDF standard, ISO 32000-1, describes two places for metadata. The first is the document information dictionary, an optional entry in the file's trailer. It holds Title, Author, Subject, Keywords, Creator, Producer, CreationDate, and ModDate, plus any other entry a program adds. The standard defines Author as the name of the person who created the document, Creator as the product that created the original document, and Producer as the product that converted it to PDF. This is the part you see in the Properties window of a PDF viewer, and it is the only part most metadata cleaners touch.
The second place is a metadata stream, which is XML in the XMP format. The standard says the contents of a metadata stream are metadata represented in XML, and a note adds that tools that do not understand PDF can read it as plain text when the stream is neither compressed nor encrypted. A stream can describe the whole document, and it can also be attached to a page or to a single picture. It repeats the author, title, and dates, and adds the program that created the file, the unique ID of the document, the ID of this version, and often a history of the editing programs and the times they saved the file.
A third item is the file identifier, an optional entry in the trailer that holds two byte strings. The standard says the first is a permanent identifier based on the contents when the file was first created, and that it does not change when the file is updated. Two files with the same first string came from the same original, even if their content differs. That is useful for tracing a document and unnecessary for showing one, so it is removed.
Earlier versions and unused data
Many programs save a PDF by adding the changes to the end of the file. The standard describes this as an incremental update: changes are appended to the end of the file, leaving its original contents intact, deleted objects are marked as deleted but left unchanged in the file, and a file that has been updated several times may contain several copies of an object with the same number. That is what makes a quick save fast, and it is also why a name you replaced, a paragraph you removed, or a page you deleted can still be there. A cleaner that only changes the document properties adds one more copy on top.
This page does not edit the file in place. It reads the version that a viewer shows, follows every reference from the document catalog to find the objects the pages use, and writes only those into a new file with a single cross-reference table. Everything it did not reach, such as older copies and objects nothing refers to, is not carried over. The report counts the saves, the older copies, and the unused objects, so you can see how much was left behind before you clean the file.
Comments, photos, attachments, and scripts
A comment, a highlight, or a sticky note stores the name of the person who made it and the time. The page lists them by page and removes the name and the time by default. The text of the comment is content, so it stays. Turn the setting off if the names are part of the record, for example in a review where the reviewers must be known.
A JPEG picture placed in a PDF is usually stored exactly as the camera or phone wrote it, with its EXIF block. That block can hold the GPS position, the camera model, and the time. The page finds these pictures, reports the location it can read, and removes the photo data with the same code as the photo cleaner on this site. The compressed picture data inside the JPEG is not touched, and the page checks that the bytes of the picture are identical before it keeps the result.
A PDF can also carry attached files and JavaScript. Both are reported with their names or the start of the script. They are not removed unless you switch the setting on, because an attachment or a form calculation is often the point of the file.
What the research found
Researchers at Inria in France collected 39,664 PDF files that 75 security agencies in 47 countries had published, and measured what the files revealed. Only 7 of the agencies sanitized any of their files before publishing, and in 65% of those sanitized files the researchers still found sensitive information. They concluded that removing the data at the surface is not enough: all the hidden data has to go. They also describe eleven kinds of hidden data and embedded content that the US National Security Agency lists for PDF files, among them metadata, attached files, scripts, hidden layers, stored form data, comments, update data, obscured text and images, and unreferenced data.
This page covers the kinds on that list that can be removed without changing what a page shows: the metadata, attached files, scripts, comment details, update data, and unreferenced data. It reports filled-in form data and signatures. It does not search a page for text that is hidden behind a picture or in a layer that is switched off, because that text is part of what the page contains. The limitations below say so plainly.
How the cleaner was tested
Sample files were made with libraries that are not part of this page: reportlab, pikepdf with qpdf, and pypdf. They cover a classic cross-reference table, object streams with a cross-reference stream, a fast-web-view file, an incremental update, a damaged index, a file cut short, comments, an attached file, JavaScript, a filled form, a JPEG with GPS and camera data, and thumbnails with private data. Each was cleaned, and the result was checked with qpdf, which found no structural problem, and with PyMuPDF, which found the same page count and the same text and rendered every page to identical pixels. The output was searched, after qpdf decoded every stream, for each name, program, and phrase that had been put in. None was left, and a control run on the uncleaned files showed that the search does find them.
The page was also run on 75 real PDF files from different programs, between 2 KB and 16 MB. All of them cleaned. Text was identical in all of them, the first six pages and the last page rendered to identical pixels at low resolution, and none of the 186 values from their original document properties was left in the file. A final test damaged the sample files in 300 random ways. Each damaged file was either cleaned or refused with a clear message, and none caused an unexpected error.
Limits and accuracy
- A PDF that is protected with a password cannot be read. Unlock it first with the Unlock PDF tool on this site, or save a copy without the password from your PDF viewer, and add that copy.
- The clean copy is a new file, so a digital signature on the original shows as invalid on it. The page warns you when a file is signed.
- Text and pictures that are part of the pages stay, including text hidden behind a picture, text in a white color, and layers that are switched off. The page does not redact anything. Use a redaction tool for that.
- Names inside the visible content, such as a signature block or a header, are content and are not touched. Filled-in form values are listed but kept.
- Only JPEG pictures are cleaned. Other picture types in a PDF are stored as raw or compressed pixels and carry no photo block.
- The clean copy has no XMP block, so a PDF/A file loses the label that says it is PDF/A, and a file without a title no longer shows one in the browser tab or to a screen reader.
- A file that stored its objects in compressed object streams is written out with plain objects, so the clean copy can be larger. In the tests the largest growth was about a fifth.
- The page cleans what it can read. Streams that use a filter it does not support are copied unchanged, and form data written as XFA is not examined.
Frequently asked questions
How do I remove metadata from a PDF?
Add the file on this page, read the report of what it hides, and press the button. The page rewrites the file without the document properties, the XMP block, the document ID, the earlier versions, and the other hidden data, then downloads a clean copy. Nothing is uploaded.
Is removing the properties in a PDF viewer enough?
Often it is not. The properties window shows the document information dictionary, but the same author and software are usually repeated in an XMP block. A file that was saved incrementally can also keep the old values in earlier copies of the data. This page removes all three.
Will the PDF look the same after cleaning?
Yes. The page content, fonts, and pictures are copied byte for byte. The page also opens the clean file again and checks that it has the same number of pages. In the tests with real files, the rendered pages were identical to the originals.
Does it remove the location from photos in the PDF?
Yes, for JPEG pictures. The page lists the GPS position and camera it finds and removes the photo block, leaving the compressed picture data untouched. The setting can be switched off. Other picture types carry no photo block inside a PDF.
Why is my password-protected PDF refused?
An encrypted file has to be decrypted to read its objects, and this page does not ask for or try passwords. Unlock it first with the Unlock PDF tool on this site, or save a copy without the password from your PDF viewer, and add that copy here.
Does cleaning remove the author name from comments?
By default, yes. The name, the subject line, and the time of each comment, highlight, and note are removed, and the comment text stays. You can switch this off in the settings if the reviewers need to stay named.
Is my PDF uploaded to a server?
No. The file is read and rewritten in your browser, and the code behind the page makes no network requests. The clean copy is created on your device and downloaded from there.
Research and references
This page was written and checked against the sources below.
- Library of Congress: PDF, Version 1.7 (ISO 32000-1:2008)
- Adobe: XMP Specifications (Part 3, Storage in Files)
- Adhatarao and Lauradoux: Exploitation and Sanitization of Hidden Data in PDF Files (ACM IH&MMSec 2021)
- pikepdf documentation: a Python wrapper around the qpdf library, used to check the cleaned test files


