Text

Email Extractor: Find Email Addresses in Text and Files

Email extractor that finds addresses in text, Word, Excel, and mail files, checks each one, and sets aside file names. Nothing is uploaded.

Use the Email Extractor: Find Email Addresses in Text and Files

Searched in your browser. Nothing is uploaded or stored.

Settings · what counts as an address

Paste text or add files to find the email addresses in them.

Everything runs in your browser. What you enter is never uploaded or stored.

An email address is easy to see and harder to find. In running text it is wrapped in brackets, quotes, and punctuation, it sits in links and tables, and some of it is written to confuse programs, such as name [at] example [dot] com. A simple pattern also finds things that are not addresses: the file name logo@2x.png, a user name inside a web address, a social media handle, a price written 5@10.

This page finds the addresses in text you paste or files you add, checks each one against the rules for addresses, and sets aside what is not. You get a list of different addresses with the number of times each appeared and the line where it was first seen. Everything is searched in your browser, nothing is uploaded, and the code behind the page makes no network requests.

How to use the Email Extractor: Find Email Addresses in Text and Files

  1. Paste text or add filesPaste text in the box, or drop files on the box above it. Text, CSV, HTML, mail messages saved as .eml, Word, Excel, PowerPoint, and OpenDocument files are read. You can add up to 50 files.
  2. Check the settingsOpen Settings to choose whether to look for disguised addresses, accept international letters, ignore capital letters, write the addresses in lower case, or leave out role addresses such as info@.
  3. Read the listThe tiles show how many different addresses were found and how many were repeats. The table shows each address, how many times it appeared, and the line where it was first seen.
  4. Look at what was set asideCandidates that have the form of an address but end in a domain that does not exist are listed apart, and candidates that are not addresses are listed with the reason. Check them if you expected more.
  5. Copy or downloadCopy the list with new lines, commas, or semicolons, copy it in groups of 50 to 500 for a Bcc field, or download it as text or as a CSV file with the counts.

What an address looks like

The format of an email address is defined by RFC 5322. An address is a local part, an @ sign, and a domain. The local part is usually a dot-atom: one or more characters from a set called atext, which is letters, digits, and the characters ! # $ % & ' * + - / = ? ^ _ ` { | } ~, with single dots between the groups. The RFC defines dot-atom-text as one or more atext characters followed by any number of a dot and one or more atext characters, so a dot cannot be first, last, or next to another dot. RFC 5321 sets the limits for mail transport: 64 octets for the local part and 255 for the domain.

RFC 6531 extends the format for international mail. The atext definition is extended to allow a UTF-8 string, so a name such as jörg or 张伟 is valid, and the domain may be written as a Unicode name or in its ASCII form. The page accepts both and shows the ASCII form that DNS uses, so jörg@münchen.de and jörg@xn--mnchen-3ya.de are counted as one address.

Why a simple pattern is not enough

The characters that are legal in a name include / = ? & and %, which are also the characters of web addresses. A pattern that accepts all of them reads ?to=anna@example.com as the address to=anna@example.com. The page ends a name at those characters unless you turn that setting off. It also ends a name at brackets, braces, and quotes, which people use to wrap an address in text, and it takes the trailing full stop, comma, or semicolon off the end.

Some @ signs are not part of an address at all. In ftp://user:secret@files.example.com the part before the @ is a user name, and the page skips it. A candidate with two @ signs in a row is skipped. A word that starts with @ is a handle. The page lists the candidates that it rejects, with the reason, so you can see what was left out.

File names that look like addresses

Web pages and style sheets are full of names such as logo@2x.png, which have the form of an address and are not one. The page checks the end of the domain against the IANA list of top-level domains, the root zone of the DNS, which had 1,437 entries in the version used here. A domain that ends in png, js, or jpg is not in it, so the candidate is listed apart as probably not an address, with a note when the ending is a common file extension. International top-level domains are in the list in their xn-- form, and the page converts a domain before it checks it. An address at a domain such as .local or .invalid, which is correct in form and does not exist on the internet, is listed apart in the same way.

Addresses written to confuse programs

Some pages write an address as name [at] example [dot] com or name(at)example(dot)com so that programs do not collect it. The page reads these forms by default, and marks the address as disguised. It decodes HTML entities such as @ for the @ sign, and the %40 code in a mailto link. A third setting also reads name at example dot com, which can find things that are not addresses, so it is off by default. An address that is made unreadable on purpose, for example by adding letters that a person is expected to remove, is not recovered.

Files

A Word, Excel, PowerPoint, or OpenDocument file is a ZIP archive of XML parts. The page reads the text of the parts, and also the targets of the links, so an address that is hidden in a mailto link under other words is found. A mail message saved as .eml is read with its encodings undone: quoted-printable text, which splits long lines with a trailing equals sign, and base64 text. An old binary Word or Excel file and a PDF cannot be read yet, and the page says so. Copy the text from your viewer and paste it.

How the extractor was tested

The form of an address was compared with the Python email-validator library, which was run without any DNS checks, on 6,000 generated addresses: names with every atext character, dots in the wrong places, lengths around 64, international names, and domains with long parts, bad hyphens, and different top-level domains. On the 3,160 where both programs apply the same rules, they agreed on every one, and on the form of the international domains as well. The other 2,840 were left out for three reasons. This page rejects a name longer than 64 octets, as RFC 5321 says, and the library does not. It requires a top-level domain of two or more letters, or the xn-- form. And the library rejects reserved names such as .test, which this page accepts in form and lists apart.

The search was tested on 400 documents in which addresses were planted in 19 ways, such as in angle brackets, in quotes, in mailto links, in markdown, in HTML, after a question mark, and in entity form, along with decoys. Every planted address was found and no decoy was reported. It was also run on 40 mail messages made with Python's email package, with quoted-printable and base64 parts, and found every address that Python's parser finds. On 3,396 real text files from one computer it found 1,325 addresses that a one-line regular expression also finds. The 32 that the expression found and this page did not were file names, domains that do not exist, names over 64 octets, and an address that had been scrambled on purpose.

Limits and accuracy

  • The page checks the form of an address. It does not check that the mailbox exists, because that needs a network request, and it never sends mail.
  • PDF files and old binary .doc and .xls files cannot be read. Copy the text and paste it, or save the file in a newer format.
  • An address broken over two lines, as in a text that was wrapped, is not joined. Mail messages are the exception, because their line breaks are undone first.
  • A quoted local part, such as "john smith"@example.com, and an address with an IP number in brackets are allowed by the RFCs and are not extracted.
  • The list of top-level domains is a copy from the date shown on the page. A domain that was added later is listed apart until the copy is updated.
  • The setting that reads name at example dot com can find text that is not an address. Check the list when you use it.
  • Collecting addresses is not the same as having permission to write to them. The laws on commercial mail, such as the GDPR in Europe and the CAN-SPAM Act in the United States, and the rules of your mail provider, may apply to a list you collect.

Frequently asked questions

How do I extract email addresses from text?

Paste the text in the box on this page, or drop a file on it. The page lists every different email address, with the number of times it appeared and the line where it was first seen, and you can copy the list or download it as text or CSV.

Can it extract emails from a Word or Excel file?

Yes. It reads the text of .docx, .xlsx, .pptx, and OpenDocument files, and also the targets of links, so an address hidden under other words in a mailto link is found. Old .doc and .xls files and PDF files cannot be read yet.

Why are some results listed as not addresses?

They have the form of an address but are probably something else. The end of the domain may not be a real top-level domain, as in logo@2x.png, or the candidate may be a user name in a web address, or have two @ signs. Each is listed with the reason.

Does it find addresses written as name [at] example [dot] com?

Yes, by default. It reads brackets and parentheses, HTML entities, and mailto codes. A setting for the plain words at and dot is off by default, because it can find text that is not an address.

Does it check that an address works?

No. It checks that the address has the right form and that its domain ends in a real top-level domain. Checking that a mailbox exists needs a network request to the mail server, and this page makes none.

Can I get the list ready for a Bcc field?

Yes. Choose semicolons for Outlook or commas for most other programs, and choose groups of 50 to 500, since many providers limit the number of recipients of one message. Each group has its own copy button.

Is my text or file uploaded?

No. The text and the files are read and searched in your browser, and the code behind the page makes no network requests. Nothing is saved, so copy the list before you close the page.

Research and references

This page was written and checked against the sources below.

  1. RFC 5322: Internet Message Format (addr-spec, atext, dot-atom)
  2. RFC 5321: Simple Mail Transfer Protocol (size limits)
  3. RFC 6531: SMTP Extension for Internationalized Email
  4. RFC 5891: IDNA 2008, Protocol (hyphen restrictions)
  5. IANA: list of top-level domains