Privacy

Redact PDF: Remove Text for Good, Right in Your Browser

Redact PDF files in your browser: search for names, emails, and numbers or draw boxes, and the text under them is removed for good. Nothing is uploaded.

Use the Redact PDF: Remove Text for Good, Right in Your Browser

Add a PDF, then mark what to black out

Search for names, numbers, emails, and other data, or draw boxes by hand. Every page with a box is turned into a picture with the boxes painted on it, so the text under a box is removed from the file instead of being covered. Afterward the page opens the new file and shows what it still holds.

Everything runs in your browser. What you enter is never uploaded or stored.

A black rectangle over a word in a PDF changes how the page looks, not what the file contains. The word is usually still there under the box, and anyone who selects the area, copies it, or searches the file can read it. This has exposed names and figures in court filings and government releases more than once. Real redaction removes the content itself, and that is the only kind this page does.

Add a PDF, then mark what to black out. You can search for words and for kinds of data, or draw boxes by hand. When you redact, each page that has a box is drawn as it looks, the boxes are painted on it, and the page is replaced by that picture. The text under a box is not hidden. It is not in the file at all. The page then opens the new file and shows what it still holds. Nothing is uploaded, and the code behind the page makes no network requests.

How to use the Redact PDF: Remove Text for Good, Right in Your Browser

  1. Add your PDFDrop a PDF on the box, or choose it from your device. A file that is protected with a password has to be unlocked first, for example with the Unlock PDF tool on this site.
  2. Find what to black outType words or phrases, one per line, and tick the kinds of data to look for. Press Find on all pages. Every match gets a box, and the page jumps to the first one. Look at each box on the page. Remove any that is wrong.
  3. Draw boxes where the search cannotDrag on the page to draw a box around anything else, such as a signature, a logo, or text in a scanned page. Drawn boxes can be removed from the list under the page.
  4. Choose the picture quality and press RedactSharp keeps every pixel and makes a larger file. JPEG makes a smaller one. Pick the resolution, decide whether the document properties and attached files go too, and press the button.
  5. Read the checks and downloadThe page lists what was checked in the new file. If one of the words you redacted is still readable on a page that has no box, it says which page. Download the file when the checks pass.

Why a black box is not redaction

A PDF keeps the text of a page and the way it is drawn as separate things. A rectangle drawn on top, whether it is a shape, a highlight, or a comment, only adds a shape to the page. The text underneath stays in the page content, where a copy command, a text search, or a script can read it. The US National Security Agency describes redaction as the process of selectively removing visible and non-visible classified or sensitive information from a document, and its guide for Adobe Acrobat is built around the step that removes the content itself, followed by a step that removes metadata, hidden content, and scripts.

Researchers at Inria in France found the same problem on a large scale. They collected 39,664 PDF files that 75 security agencies in 47 countries had published, and in 65% of the files that had been sanitized, they still found sensitive information. The lesson is that removing data at the surface is not enough. Everything that could carry it has to go.

What this page does to a redacted page

The page is drawn in your browser by the PDF.js library from Mozilla, the one that Firefox uses to show PDFs, and the boxes are painted on it in black. The result is a picture of the page. The new file keeps the page's size and puts that picture on the page, and the old content, fonts, comments, links, and form fields of that page are not carried over. A PDF reader finds nothing on the page except the picture, so there is nothing to select, copy, or search.

Pages without a box are copied as they are, so their text stays selectable and the file stays small. The cost on a redacted page is that its remaining text is part of the picture too. If you need that page searchable after redacting, run it through an OCR program afterward. A fresh text layer made from the redacted picture can only contain what the picture shows.

Finding what to redact

Words and phrases are matched as you typed them, ignoring case unless you ask for it, and a phrase matches even when a line break falls in the middle. The kinds of data are checked by their rules, so that the page does not black out every long number. A card number has to pass the Luhn check, the check digit that the standard for card numbers uses. An IBAN has to pass the mod 97 check that ISO 13616 defines. A Social Security number must have an area, group, and serial number that can be issued, so 000, 666, and 900 to 999 as the first part are skipped. A phone number needs 8 to 15 digits with separators, and dates are not mistaken for it. Email addresses and IP addresses are found by their format. Text in form fields and comments is not part of the page text, so it is searched separately, and a field or comment that matches gets a box over the whole of it.

These are the kinds of data that court rules ask people to remove. US Federal Rule of Civil Procedure 5.2 says that a filing with a Social Security number, a taxpayer identification number, a birth date, the name of a minor, or a financial account number may include only the last four digits of the numbers, the year of the birth date, and the minor's initials. This page blacks out the whole match. If a rule lets you keep the last four digits, draw a box over the rest by hand.

Each match is turned into a box on the page. Where in a line a word sits is worked out from the width of its characters, using the widths of Helvetica and Times, scaled to the length of the line, and a margin that grows with the font size is added so that no edge of a letter is left. Fonts with other proportions, such as bold text, can move a box by a point or two, which is why you see every box on the page before you redact, and why the margin can be made wider.

What else can hold the text

The page content is not the only place a name can be. A comment on another page may quote it. A bookmark may carry it as a title. A form field keeps its value in the field, away from the page. Tags for screen readers hold a copy of each page's text. The document properties and the XMP block may name it in the title or subject. An attached file may be a whole copy of the document.

When you redact, the page replaces the pages you boxed, removes the comments, links, and form fields on them, and drops the form fields that belonged to them. It removes the accessibility tags and the article threads, since both can hold a copy of the text. It looks for every word you redacted in the rest of the file, and it removes a comment or a bookmark that holds one. The document properties and XMP can be removed with a setting, and attached files with another. Attached files are kept unless you ask, since they may be the point of the file, but the page warns you that they may hold the text.

What the check after redacting shows

The new file is opened in PDF.js a second time, from its own bytes. The page counts the text it finds on every page you redacted, which should be nothing, and it looks for each redacted word on the pages you did not box. A name that you blacked out on page 2 may be written again on page 9. That is not an error in the file, but it is something you would want to know, so the page lists it, and you can add a box and redact again.

A separate search covers the rest of the file: properties, XMP, bookmarks, comments, form fields, attached files, and any page content that is stored as plain characters. The result says either that none of the words were found, or where each one still is. The check cannot read text that a font encodes in its own way, so the second reader's view of each page is the main evidence for the pages.

How the redactor was tested

A sample PDF was made with reportlab and pikepdf, which are not part of this page. It has the same name, email address, phone number, Social Security number, card number, IBAN, and IP address in Helvetica, Times, Courier, and bold type on a page with a 90 degree rotation, plus a bookmark, a link, a form field holding the name, a comment that quotes it, an attached file, and document properties. The boxes this page made for each match were compared with the exact word boxes that PyMuPDF reads from the file. Of the 23 boxes compared, the 16 on upright pages reached 1.5 to 1.8 points past their word on both sides. On the rotated bold page, where the character widths are a guess, the worst edge fell 0.2 points short of the line box that PyMuPDF reports, and the furthest reached 6 points beyond it.

A run in Chrome then redacted pages 1 to 4 of the sample, with every kind of data ticked and one box drawn by hand. The file it downloaded was drawn with PyMuPDF and compared with the original pixel by pixel: no ink of any of the 65 blacked-out words showed, and the 77 words next to them were left alone. The redacted files were also checked with qpdf, which found no structural problem, and with PyMuPDF, which found no text on any redacted page, the same page sizes as before, one picture on each, and identical text and pixels on the page that was not redacted. The form field, the bookmark, the comment, the attached file, and the properties that named the person were gone, and when every page was boxed, the name appeared nowhere in the bytes of the file. The check for card numbers, IBANs, and Social Security numbers was run against published test numbers and against the same numbers with one digit changed.

Limits and accuracy

  • A redacted page becomes a picture, so its remaining text can no longer be selected or searched. Run an OCR program on it if you need that.
  • The search needs text it can read. A scanned page has none, so draw its boxes by hand. The page tells you when a page has no searchable text.
  • The position of a found word is worked out from character widths, so bold text and unusual fonts can be a point or two off. Look at each box before you redact, and widen the margin if needed.
  • The pictures are drawn by PDF.js, so a page may differ slightly from how another program draws it, for example when a font is not embedded in the file.
  • A PDF that is protected with a password has to be unlocked first. Use the Unlock PDF tool on this site.
  • A digital signature on the original does not apply to the redacted copy, because it is a new file. The page warns you when a file is signed.
  • Attached files are kept unless you switch the setting on, and they may contain the text. The same goes for text hidden in the pages you did not box, such as a white-on-white line or a switched-off layer.
  • Redaction removes what is in the file. It cannot remove what was already copied, indexed, or sent before. If a file with the text has been shared, treat the text as exposed.

Frequently asked questions

How do I redact a PDF?

Add the file on this page, search for the words and kinds of data to remove or draw boxes by hand, check each box on the page, and press Redact. Each page with a box becomes a picture with the boxes painted on, so the text under them is gone. Download the new file.

Is a black box over text enough?

No. A rectangle only covers the text on the screen. The text usually stays in the file, where it can be selected, copied, or searched. Real redaction removes the text itself, which is what this page does by turning the page into a picture.

Can someone copy the redacted text from the new PDF?

Not from a redacted page. The page holds one picture and nothing else, and a second reader finds no text on it. The page also looks for the words in the rest of the file and tells you if one is still readable on a page that has no box.

Does it find names and numbers by itself?

It finds the words you type and these kinds of data: emails, phone numbers, Social Security numbers, card numbers, IBANs, and IP addresses. Numbers have to pass their own checks. It does not guess names, so type the names you want to remove.

Why is the text on a redacted page not selectable any more?

Because the page is now a picture. That is how the text under the boxes is removed for good. Pages without a box keep their text. If you need a searchable redacted page, run an OCR program on the result.

Does it work on scanned PDFs?

Yes, with boxes you draw by hand. A scan has no text for the search to read, so the page says so and you drag boxes over what to remove. The result is a picture either way.

Is my PDF uploaded to a server?

No. The file is opened, drawn, and rewritten in your browser, and the code behind the page makes no network requests. The redacted copy is created on your device and downloaded from there.

Research and references

This page was written and checked against the sources below.

  1. NSA: Redaction of PDF Files Using Adobe Acrobat Professional X (Information Assurance Directorate)
  2. Adhatarao and Lauradoux: Exploitation and Sanitization of Hidden Data in PDF Files (ACM IH&MMSec 2021)
  3. Federal Rules of Civil Procedure, Rule 5.2: Privacy Protection for Filings Made with the Court
  4. Library of Congress: PDF, Version 1.7 (ISO 32000-1:2008)
  5. PDF.js: the Mozilla library that draws the pages in this tool