Engineering guide

URL Encoding Explained: Percent-Encoding and When to Use It

What is URL encoding? Learn how percent-encoding works, which characters must be encoded, encodeURI vs encodeURIComponent, and common mistakes to avoid.

An address bar with highlighted percent-encoded segments above three characters mapping to percent codes

Paste a link with a space or an ampersand in the wrong place and it can break, or worse, quietly point somewhere you did not intend. URLs can only safely carry a limited set of characters, and URL encoding, also called percent-encoding, is the mechanism that lets everything else travel safely.

This guide explains how percent-encoding works, which characters need it, how non-English characters are handled, the difference between encodeURI and encodeURIComponent in JavaScript, and the mistakes that cause broken or double-encoded links. It draws on the URI standard (RFC 3986) and Mozilla's documentation, linked at the end.

The short answer: URL encoding replaces each unsafe character with a percent sign followed by two hexadecimal digits, so a space becomes %20 and an ampersand inside a value becomes %26. In JavaScript, use encodeURIComponent() on each value you insert into a URL.

What is URL encoding?

Mozilla describes percent-encoding as a mechanism to encode 8-bit characters that have specific meaning in the context of URLs, and notes that it is sometimes called URL encoding. RFC 3986 defines the format precisely: a percent-encoded octet is a character triplet consisting of the percent character followed by the two hexadecimal digits that represent that octet's numeric value.

In practice that means a space, which has no place in a URL, is written as %20, because the space character has the byte value 20 in hexadecimal. The browser or server reverses the process, called decoding, to recover the original text.

Which characters need to be encoded?

RFC 3986 sorts URL characters into groups. Unreserved characters, which need no encoding, are uppercase and lowercase letters, digits, the hyphen, the period, the underscore, and the tilde. Reserved characters are the ones with a special job in URL syntax: the general delimiters : / ? # [ ] @ and the sub-delimiters ! $ & ' ( ) * + , ; =.

A reserved character can be used for its special purpose, such as a slash separating path segments, or it can appear as plain data inside a value. When you mean it as data, you must percent-encode it. The standard is explicit that URIs differing in the replacement of a reserved character with its percent-encoded form are not equivalent, so an encoded slash is not the same as a slash. Anything outside these sets, such as spaces, quotation marks, angle brackets, and non-ASCII letters, must always be encoded.

  • Space becomes %20.
  • Ampersand (&) becomes %26.
  • Equals sign (=) becomes %3D.
  • Question mark (?) becomes %3F.
  • Hash (#) becomes %23.
  • Slash (/) becomes %2F.
  • Plus (+) becomes %2B.
  • Percent (%) becomes %25.

How non-ASCII characters are encoded

Letters like é, ü, or 日 are not ASCII, so they are first converted to bytes using UTF-8, and then each byte is percent-encoded. RFC 3986 says that text should first be encoded as octets according to UTF-8 and that only the octets that do not correspond to unreserved characters should then be percent-encoded.

The letter é is two bytes in UTF-8, C3 and A9, so the word café becomes caf%C3%A9. Mozilla describes the JavaScript behavior the same way: encodeURIComponent replaces each character with one, two, three, or four escape sequences representing its UTF-8 encoding.

Space: %20 or plus?

A space can appear as %20 or as a plus sign, and the difference depends on context. Mozilla notes that a space translates to a plus in application/x-www-form-urlencoded messages, the format HTML forms use, and to %20 in standard URLs.

This is why a plus in a query string is ambiguous. Mozilla's documentation for URLSearchParams says it encodes the space character as a plus, so a value like More webdev appears as More+webdev. If you need a literal plus sign inside a value, encode it as %2B, otherwise a decoder that follows form rules will read it as a space.

encodeURI vs encodeURIComponent

JavaScript provides two built-in encoders, and choosing the wrong one is a classic bug. According to Mozilla, encodeURI preserves characters that have special meaning in URL syntax, including ; / ? : @ & = + $ , and #, while encodeURIComponent encodes nearly all of them. Neither escapes the unreserved letters, digits, hyphen, underscore, or period.

Mozilla's guidance is that encodeURI is for encoding a URL as a whole, assuming it is already well formed. To assemble dynamic values into a URL, use encodeURIComponent on each value so syntax characters do not end up in unwanted places. Its own example makes the failure clear: encodeURI on a link containing the name Ben & Jerry's leaves the ampersand untouched, which breaks the query parameter, while encodeURIComponent turns it into %26 and keeps the value intact.

A couple of details are worth knowing. encodeURIComponent does not escape the characters ! ' ( ) and *, although RFC 3986 reserves them, so strict implementations add extra escaping. And it throws a URIError if you pass it a lone surrogate, which is half of an emoji or other character that is not part of a valid pair.

  • Use encodeURIComponent for each individual query value or path segment you insert into a URL.
  • Use encodeURI only when you already have a complete, well-formed URL that needs light cleanup.
  • Never run encodeURIComponent on a complete URL, because it also encodes the colon and slashes that make the URL work.

The easier way: let the platform build the URL

Hand-assembling query strings invites mistakes. Mozilla documents URLSearchParams as a built-in interface for building and parsing query strings, and it percent-encodes special characters for you. Using it, or the URL-building helpers in your language or framework, means you rarely have to call an encoder yourself.

Remember that URLSearchParams follows form-encoding rules, so spaces appear as plus signs. That is normal and decoders understand it, but it explains why the same text can look different depending on how a URL was built.

Worked example: building a search URL

Suppose you want a link to a search for the phrase rock & roll. The raw text contains two characters that cannot appear unencoded in a query value: a space, and an ampersand that would otherwise start a new parameter.

Run the value through encodeURIComponent and you get rock%20%26%20roll, so the finished link ends in ?q=rock%20%26%20roll. Build the same query with URLSearchParams and you get q=rock+%26+roll instead, with the spaces written as plus signs. Both decode back to exactly rock & roll, because the server knows which convention the query string uses.

Now compare the mistake. If the ampersand is left raw, the link ends in ?q=rock & roll, and the server sees a parameter named q with only the value rock, followed by a stray second parameter with no value. The search silently runs for the wrong thing. This is why encoding each value, not the whole URL, is the habit worth building.

Common URL encoding mistakes

Most URL bugs come from the same few mistakes, and recognizing them makes debugging much faster.

  • Double encoding: encoding a value that is already encoded turns %20 into %2520, because the percent sign itself becomes %25. If you see %25 where you expected %, something was encoded twice.
  • Encoding a whole URL: running encodeURIComponent on https://example.com/page produces https%3A%2F%2Fexample.com%2Fpage, which is no longer a working link.
  • Forgetting to encode & = or # in values: an unencoded ampersand starts a new parameter, and an unencoded hash starts the fragment, so the rest of your value may never reach the server.
  • Decoding too early: split a query string into parameters first, then decode each piece, otherwise an encoded & or = in a value is mistaken for a separator.
  • Mixing plus and %20: be consistent about which context you are in, and encode a literal plus as %2B.
  • Lowercase hex digits: they work, but RFC 3986 recommends uppercase hexadecimal digits for consistency, so %2F is better than %2f.
  • Treating encoding as security: percent-encoding is not encryption and does not make input safe, so anyone can decode it and your server must still validate what it receives.

How to decode a URL

Decoding reverses the process: each percent sign followed by two hexadecimal digits is converted back to its byte, and the bytes are read as UTF-8. In JavaScript, decodeURIComponent does this for a single value.

Be ready for malformed input. A stray percent sign that is not followed by two valid hexadecimal digits makes the JavaScript decoder throw an error, so wrap decoding of untrusted input in error handling rather than assuming every URL is well formed.

Privacy and security notes

URLs are routinely stored in browser history, server logs, and analytics, so avoid putting passwords, tokens, or personal data in them, encoded or not. Encoding only changes how characters are written, not who can read them.

When you use an online encoder or decoder, remember that a query string you paste may contain exactly that kind of private data. Prefer tools that run in your browser, or the console and URL helpers you already have, and avoid pasting live tokens into sites you have not verified.

Practical checklist

  • Encode data values with encodeURIComponent before inserting them into a URL.
  • Use encodeURI only on a whole, already well-formed URL.
  • Remember that non-ASCII characters are encoded as UTF-8 bytes, such as é becoming %C3%A9.
  • Encode & = ? # / and + when they are data rather than syntax.
  • Use URLSearchParams or your framework's URL helpers instead of concatenating strings.
  • Split query strings into parameters before decoding each value.
  • Watch for %25 in output, which is a sign of double encoding.
  • Keep secrets out of URLs, and use local tools for sensitive query strings.

Research and references

This guide was prepared from the authoritative references below.

  1. RFC 3986: Uniform Resource Identifier (URI): Generic Syntax
  2. MDN Web Docs: Percent-encoding
  3. MDN Web Docs: encodeURIComponent()
  4. MDN Web Docs: encodeURI()
  5. MDN Web Docs: URLSearchParams