Open an email's raw source, a data URL, or an authentication header and you will often see a long run of letters and digits ending in one or two equals signs. That is Base64, one of the most widely used and most misunderstood encodings on the web.
This guide explains what Base64 is, how it converts bytes to text step by step, why the result is about a third larger, how the URL-safe variant differs, where Base64 is used, and why it must never be mistaken for encryption. It draws on RFC 4648, the standard that defines it, and on Mozilla's documentation, linked at the end.
The short answer: Base64 represents binary data as plain text using 64 safe characters, so it can travel through systems that only handle text. It is a way of writing data down, not a way of hiding it.
What is Base64?
MDN describes Base64 as a group of similar binary-to-text encoding schemes that represent binary data in an ASCII string format by translating it into a radix-64 representation. It exists because many systems, including email and parts of the web, were designed for text and can corrupt or reject raw binary bytes.
By converting the bytes into a limited, safe alphabet, Base64 makes sure the data remains intact without modification during transport over media that can only handle ASCII text. MDN lists email via MIME, storing complex data in XML, and encoding binary data for data URLs as common uses.
How Base64 works, step by step
RFC 4648 defines the process: take the input as 24-bit groups, split each group into four 6-bit pieces, and write each piece as one character from the alphabet. Because 6 bits can represent 64 values, each character stands for a number from 0 to 63.
Take the three letters Man. In binary, M is 01001101, a is 01100001, and n is 01101110. Joined, that is 24 bits: 010011010110000101101110. Split into four groups of 6 bits, you get 010011, 010110, 000101, and 101110, which are the values 19, 22, 5, and 46. In the Base64 alphabet those are T, W, F, and u, so Man becomes TWFu.
The Base64 alphabet
The standard alphabet from RFC 4648 uses 64 characters: uppercase letters A to Z for values 0 to 25, lowercase a to z for 26 to 51, the digits 0 to 9 for 52 to 61, then plus for 62 and slash for 63. A 65th character, the equals sign, is used only for padding.
Every character is plain ASCII and safe in text systems, which is the whole point. The characters were chosen so they survive being passed through software that treats text specially.
What is the equals sign padding?
Input is processed in groups of three bytes, but real data is rarely an exact multiple of three. When the last group is short, Base64 pads the output with equals signs to complete the final group of four characters. RFC 4648 says implementations must include the appropriate pad characters at the end of encoded data unless a specification says otherwise.
You can see the pattern in short examples. Man is three bytes and encodes to TWFu with no padding. Ma is two bytes and becomes TWE=, with one equals sign. M is a single byte and becomes TQ==, with two. The word hello is five bytes, a full group of three plus two leftover bytes, so it encodes to aGVsbG8= with one equals sign.
Why Base64 makes data about 33% larger
Every 3 bytes of input become 4 characters of output, so the encoded form is roughly a third larger than the original. MDN says the Base64 version of a string or file is typically roughly a third larger than its source.
The math is easy to check. 3,000 bytes become 4,000 characters, an increase of 33.3 percent. A 1 MiB file, which is 1,048,576 bytes, becomes 1,398,104 characters. Very short inputs show more overhead because of padding: a single byte becomes four characters.
Base64 vs base64url
The standard alphabet contains plus and slash, which cause trouble in URLs and filenames. RFC 4648 therefore defines a URL-safe variant, often called base64url, that replaces plus with a minus sign and slash with an underscore, noting that the slash can be problematic in file names and URLs.
As an example, the four characters ?>?> encode to Pz4/Pg== in standard Base64. In base64url they become Pz4_Pg, and the padding is often omitted. The two variants are not interchangeable, so a decoder must know which one it is reading.
Where Base64 is used
Base64 appears wherever binary data has to travel as text. Understanding the common cases helps you recognize it when you see it.
- Email attachments, through MIME, which carries binary files inside text messages.
- Data URLs, which embed a small file directly in a page or stylesheet, such as data:text/plain;base64,aGVsbG8= for the word hello.
- JSON and XML, which are text formats and cannot hold raw bytes, so binary values are encoded as strings.
- HTTP Basic authentication, where the username and password are joined with a colon and Base64 encoded, so user:pass becomes dXNlcjpwYXNz.
- Tokens and identifiers that must be safe in URLs, which usually use the base64url variant.
Base64 is encoding, not encryption
This is the most important point in the guide. Anyone can decode Base64 instantly, with no key. RFC 4648's security considerations say so plainly: base encoding visually hides otherwise easily recognized information, such as passwords, but does not provide any computational confidentiality.
The Basic authentication example makes the risk concrete. The string dXNlcjpwYXNz looks scrambled, but it decodes to user:pass in a single step. Treat Base64 as a way to carry data, never to protect it. If you need secrecy, use real encryption over a secure connection, and do not store passwords in Base64.
Encoding, encryption, and hashing compared
These three terms are often mixed up, and keeping them straight prevents real security mistakes.
- Encoding changes the representation of data so it can be handled by a system, and is reversible by anyone. Base64 is encoding.
- Encryption scrambles data with a key so only authorized parties can read it. It is reversible only with the key.
- Hashing produces a fixed-size fingerprint that cannot be reversed to the original, and is used to verify integrity or store passwords with the right algorithm.
Worked example: decoding Base64 by hand
Decoding reverses the encoding. Take TWFu. Look up each character's value in the alphabet: T is 19, W is 22, F is 5, and u is 46. Write each as 6 bits: 010011, 010110, 000101, and 101110. Join them into 24 bits and regroup into three bytes of 8 bits: 01001101, 01100001, and 01101110. Those are 77, 97, and 110, the character codes for M, a, and n, so TWFu decodes to Man.
Padding works the same way in reverse. Take TQ==. T is 19 and Q is 16, which as 6-bit groups give 010011 and 010000, twelve bits in all. The first eight bits, 01001101, are 77, the letter M. The remaining four bits are only filler, so they are discarded, and the two equals signs simply tell the decoder how many bytes the final group really holds.
Base64 in JavaScript: btoa, atob, and Unicode
Browsers provide btoa to encode and atob to decode. MDN's example is btoa("hello") returning aGVsbG8=. But btoa treats each character as one byte, so each character must have a code point below 256. MDN says characters outside that range throw an InvalidCharacterError.
That limit catches people out with text. The word café encodes to Y2Fm6Q== if you let btoa treat it as Latin-1, but to Y2Fmw6k= if you first convert it to UTF-8, which is what most other systems expect. MDN's recommended approach for arbitrary Unicode is to convert the string to UTF-8 bytes with TextEncoder first, and then encode those bytes. Use the same approach in reverse with TextDecoder, so both sides agree on the encoding.
When not to use Base64
Base64 solves a transport problem, and it has costs. Avoid it when a better option exists.
- When you can send the bytes directly, such as a file upload or a binary API, Base64 only adds size and processing time.
- For large images embedded as data URLs, the page grows by a third and the image can no longer be cached separately from the document.
- As protection for secrets, since it provides none.
- As a compression step, since it makes data bigger, not smaller.
Common Base64 mistakes
Most Base64 bugs come from a few patterns that are easy to avoid once you know them.
- Treating Base64 as encryption, and exposing credentials that are merely encoded.
- Mixing the standard and URL-safe alphabets between encoder and decoder.
- Dropping the equals padding when the decoder requires it, or adding it when it does not.
- Passing Unicode text to btoa without converting it to UTF-8 bytes first.
- Adding line breaks or spaces that a strict decoder rejects.
- Assuming the decoded bytes are text, when they may be an image, a file, or a key.
A privacy note about online Base64 tools
Base64 strings frequently contain sensitive material, such as tokens, credentials, certificates, and file contents. Decoding is instant and requires no secret, which also means anyone you hand the string to can read it. Use a decoder that runs in your browser or a command on your own machine, and avoid pasting live credentials into any website you have not verified.
Practical checklist
- Use Base64 only to carry binary data through text-only channels.
- Never rely on it to protect secrets, because it is trivially reversible.
- Expect output about 33 percent larger than the input.
- Choose base64url for URLs and filenames, and keep encoder and decoder consistent.
- Keep or restore the equals padding as the receiving system requires.
- Convert Unicode text to UTF-8 bytes before using btoa.
- Decode sensitive strings locally rather than in unknown websites.
Research and references
This guide was prepared from the authoritative references below.



