Text

Fix Garbled Text and Weird Characters (Mojibake Fixer)

Fix garbled text in your browser. Repair weird characters, see exactly what went wrong, and open a file to find its real encoding.

Use the Fix Garbled Text and Weird Characters (Mojibake Fixer)

Processed in your browser. Nothing is uploaded or stored.

Paste text with strange characters like é, ’, or ? marks to see what went wrong and get it back.

Everything runs in your browser. What you enter is never uploaded or stored.

Garbled text has a recognizable look: an é that shows as é, a curly apostrophe that turns into ’, a name written in question marks, or a whole line of Russian or Chinese that has become a mess of accented Latin letters. It is called mojibake, and it happens when text saved in one encoding is read as another. The damage looks random, but it is not. It follows exact rules, so in most cases it can be undone exactly.

This page undoes it. Paste the damaged text and it works out which mix-up caused the damage, reverses it, and shows what changed. If you still have the file, open the file instead: it holds the original bytes, so the right encoding can be found rather than guessed. Everything runs in your browser, and the code behind the page makes no network requests, so the text you paste or open is not uploaded.

How to use the Fix Garbled Text and Weird Characters (Mojibake Fixer)

  1. Paste the text, or open the filePaste the garbled text into the box, or switch to the file tab and open the file itself. A file is the better choice when you have it, because pasting a copy can lose information that the file still holds.
  2. Read what the tool foundThe result says which mix-up caused the damage, for example UTF-8 text that was read as Windows-1252, and how many places it fixed. If the text was damaged more than once, it says so.
  3. Check the resultThe fixed text appears below, and a list shows each change, such as é becoming é. Text that was already correct is never changed, so a line that is half good and half damaged keeps its good half.
  4. Try the other possibilities if it is wrongIf the result does not read naturally, open the list of other possibilities. It tries the other likely encoding pairs and shows each result, and you can choose the one that is spelled correctly.
  5. Copy it or save it as UTF-8Copy the result, or download it as a UTF-8 file. For a CSV that opens wrongly in Excel on Windows, add the byte order mark, which tells Excel the file is UTF-8.

What garbled text is, and why it happens

A computer stores text as bytes, and an encoding is the agreement about which bytes mean which characters. Nearly all modern text is UTF-8, which RFC 3629 defines: it writes the characters of ASCII as one byte each and everything else, such as é, ’, 你, or 😀, as two to four bytes. Older programs, files, and web pages used single-byte encodings such as Windows-1252, where each byte is one character.

Mojibake appears when bytes written in one encoding are read as another. The word is Japanese for character transformation. The classic case is a UTF-8 é, which is the two bytes C3 and A9. A program that assumes Windows-1252 reads C3 as à and A9 as ©, and shows é. Nothing was lost. The two bytes are still there, shown as two wrong characters, which is why the damage can be reversed.

  • A web page, e-mail, or database that says one encoding but contains another.
  • A CSV or text file saved as UTF-8 and opened by a program that assumes an older encoding, which is the usual story with Excel on Windows.
  • Text copied through several programs, each of which converts it, so that it is damaged more than once.
  • A file from an old system in a legacy encoding such as Windows-1251 or GBK, opened as if it were UTF-8 or Windows-1252.

The patterns to recognize

The common patterns are easy to spot once you know them. Because UTF-8 uses lead bytes that read as Ã, Â, â, Ð, and Ñ in Windows-1252, damaged text is full of those letters followed by symbols.

  • é è ü ñ ç and similar: accented Latin letters in French, German, Spanish, and Portuguese, read as Windows-1252.
  • ’ “ †– …: curly quotes, dashes, and ellipses in the same situation. A ? or � may stand in for the last character.
  • © ® °: a symbol that was preceded by an extra Â.
  • привет and similar: Cyrillic text read as Windows-1252.
  • ä½ å¥½ and similar: Chinese, Japanese, or Korean text read as Windows-1252.
  • 😀 and similar: an emoji that has been through the same mix-up.
  • � (a diamond with a question mark): the reader found bytes it could not use and replaced them. These are lost characters.

How the repair works

The repair reverses the mix-up exactly. It turns each damaged character back into the byte it stood for, using the encoding the text was wrongly read as, and then reads those bytes as UTF-8. If the result is valid UTF-8, the odds are overwhelming that this was the damage, because random text almost never forms valid UTF-8 sequences. If it is not valid, the text is left alone.

This is done one damaged spot at a time. A line such as Café menu: crème brûlée has one correct é and several damaged ones. The correct é is not changed, and each damaged group of characters is repaired by itself. Text that was damaged twice has each layer peeled off in turn, up to four times. Every line is handled separately, so one damaged line in a good document does not affect the rest.

  • Windows-1252 and Latin-1 are tried first, because they cause most of the damage. Latin-1 shows bytes 80 to 9F as invisible control characters, which are handled too.
  • A quote that lost its last byte, such as †followed by a question mark, is recovered, because only one byte value could have been there.
  • A space after à or  is read as the byte A0, a non-breaking space that a web page turned into a plain space.
  • UTF-8 read as Windows-1250, 1251, 1253, Mac Roman, KOI8-R, and several other single-byte encodings is found by trying each and keeping the one that explains the most.
  • UTF-8 read as GBK, Big5, Shift JIS, EUC-JP, or EUC-KR is repaired when no bytes were lost along the way.
  • Text from an older encoding that was read as Windows-1252, such as Russian in Windows-1251 or Chinese in GBK, is read again with the right encoding.

Why correct text is not changed

A repair tool that changes text which was fine is worse than none. A word like Größe has two accented letters in a row and is entirely correct. A name such as Ñandú has letters that also appear in damaged text. This tool changes a spot only when turning its characters back into bytes and reading them as UTF-8 produces something valid and sensible, and when another encoding does not explain the line better.

The tool was checked against a collection of correct text in more than 16 languages, including Cyrillic, Greek, Arabic, Hebrew, Thai, Hindi, Chinese, Japanese, and Korean, and against text that had been damaged in many ways by independent software. Correct text came out unchanged every time, and the damaged text was repaired in the cases above.

What cannot be repaired

The repair is exact only when no information was lost. Some damage destroys it. If a program could not read a byte, it may have replaced it with �, or with a question mark, and the original byte is gone. A question mark in the place of a character, as in a name that has become K?ln, cannot be turned back into ö. The tool reports how many � characters it finds and does not invent replacements.

Pasting is also lossy in a different way. Many programs change characters when you copy them, for example by turning a non-breaking space into a plain one or by normalizing quotes. If you have the original file, open it in the file tab. The file still has its original bytes, and the right encoding can be found from them exactly.

Open the file, not a copy

The file tab reads the raw bytes of a text file and works out which encoding they are in. It checks for a byte order mark, recognizes UTF-16 even without one, accepts valid UTF-8, and otherwise tries the common older encodings, ranked by how natural the text looks. The Unicode standard's FAQ describes the UTF-8 byte order mark, EF BB BF, as a signature that is not required for UTF-8 but marks a file as UTF-8.

Some encodings look alike. Central European and Western European Windows encodings differ in only a few letters, so the tool shows the top readings with a preview of each and the words that differ between them. Choose the one that is spelled correctly. If a file is valid UTF-8 but the text inside has been damaged twice, which happens when a file is converted two times, the tool says so and repairs the text as well.

Making a CSV open correctly in Excel

A common complaint is a CSV file that looks fine in a text editor but shows é in Excel. The file is UTF-8, and many versions of Excel on Windows open a CSV without a byte order mark using an older encoding. Saving the file as UTF-8 with the byte order mark, which the download here can add, tells Excel the file is UTF-8 and the characters come out right. Importing the file through the data import feature and choosing UTF-8 as the file origin also works.

The byte order mark is a signature of three bytes at the start of the file. It is harmless to most programs that read UTF-8, but some tools, such as scripts and parsers that expect the first characters to be specific text, may be confused by it. So the tool adds it only when you ask, and it suggests it by default only for files that end in .csv or .tsv.

Other codes that look like weird characters

Not every strange sequence is mojibake. HTML entities such as é or é, percent codes such as %C3%A9 in web addresses, backslash escapes such as \u00e9 and \xc3\xa9 in code and logs, and quoted-printable codes such as =C3=A9 in e-mail are all ways of writing a character as ASCII. The tool detects them and offers to decode them. It does not decode them on its own, because in web pages and source code they are correct and decoding them would change the meaning.

They can also combine with mojibake. An entity such as é is the two characters à and ©, which are themselves the damaged form of é. Decoding the entities first and then repairing the text gives é.

How to prevent it

Most mojibake is preventable by using UTF-8 everywhere and saying so. Declare UTF-8 in HTML with a meta charset tag and in the HTTP Content-Type header, save source files and exports as UTF-8, configure databases and their connections for a UTF-8 character set, and when exchanging CSV with Excel users add the byte order mark. When a program asks for the encoding of a file, choose UTF-8 unless you know the file is old.

The WHATWG Encoding Standard, which browsers follow, even treats the labels ISO-8859-1 and ASCII as Windows-1252. That is why a page that declares itself Latin-1 shows the extra characters of Windows-1252 such as curly quotes, and why so much mojibake follows the Windows-1252 pattern.

Limits and accuracy

  • Characters that were already lost, shown as � or as a question mark, cannot be recovered from a pasted copy. The tool reports them and leaves them as they are.
  • Pasted text can differ from the original, because programs change characters when text is copied. The original file gives a more reliable result.
  • The repair covers UTF-8 text read as Windows-1252, Latin-1, and the other single-byte encodings in the list, text damaged up to four times, UTF-8 read as the East Asian encodings when no bytes were lost, and older-encoding text read as Windows-1252. Other combinations are not found automatically, although the list of other possibilities may help.
  • Telling similar encodings apart, such as Windows-1250 and Windows-1252, or Chinese, Japanese, and Korean encodings from each other, depends on how natural the text looks. This is a good guide and can be wrong, so the tool shows the other readings and the words that differ.
  • Encodings that browsers do not support, such as IBM437 and IBM850 for old DOS files, are not covered, because the page uses the browser's own encoding tables.
  • A file may be up to 20 MB, and pasted text up to 2 million characters.
  • The tool restores the text. It does not translate it, correct spelling, or turn text into a different script.

Frequently asked questions

What is mojibake?

It is text that shows as the wrong characters because bytes saved in one encoding were read as another, for example é appearing as é. The name comes from Japanese and means character transformation. The damage follows exact rules, so it can usually be reversed.

Why does é turn into é and ’ into ’?

Those characters are stored in UTF-8 as two or three bytes. A program that assumes Windows-1252 shows each byte as a separate character: the two bytes of é become à and ©, and the three bytes of the apostrophe become â, €, and ™. The text was saved correctly and read incorrectly.

How do I fix weird characters in a CSV opened in Excel?

Save the CSV as UTF-8 with a byte order mark, or import it with the file origin set to UTF-8. Many versions of Excel on Windows open a CSV without the mark in an older encoding. If the file is already damaged, open it in the file tab, which finds the real encoding and can save a fixed copy.

Can every garbled text be fixed?

No. It can be fixed when every damaged character still stands for the byte it came from, which is the usual case. It cannot when bytes were lost, which shows as � or a question mark. In that case, only the original file or another copy of the text can help.

Will it change text that is already correct?

No. A spot is repaired only when reversing it gives valid, sensible text, and correct text such as Größe, Zażółć, or Привет is left alone. It was checked against correct text in many languages, and none of it was changed.

What does the � symbol mean?

It is the Unicode replacement character. A program shows it when it meets bytes it cannot read in the encoding it is using, and it replaces them. The original characters are gone from that copy, so they cannot be recovered from it.

Is my text uploaded when I use this page?

No. The repair runs in your browser and the code behind the page makes no network requests, so what you paste or open stays on your device. You can check this in the Network panel of your browser's developer tools.

Why should I open the file instead of pasting the text?

A file holds the original bytes, so its real encoding can be identified exactly. A pasted copy has already been through your browser and clipboard, which may have changed some characters, so some damage becomes impossible to undo.

Research and references

This page was written and checked against the sources below.

  1. WHATWG Encoding Standard
  2. RFC 3629: UTF-8, a transformation format of ISO 10646
  3. Unicode FAQ: UTF-8, UTF-16, UTF-32 and BOM
  4. ftfy documentation: fixes text for you
  5. Wikipedia: Mojibake
  6. MDN Web Docs: TextDecoder