Mojibake Converter & Fixer — Character Encoding Repair Tool
Paste garbled (mojibake) text to instantly detect the encoding mismatch that caused it, with ranked repair candidates. Also converts explicitly between UTF-8, Shift_JIS, EUC-JP, and JIS (ISO-2022-JP) — free, no sign-up required.
What causes mojibake, and how this tool fixes it
Mojibake happens when text saved in one character encoding gets read back using a different one. The underlying bytes never change, but the "translation rule" applied to them does — so a perfectly valid file can come out as a wall of nonsense symbols the moment it's opened with the wrong assumption about its encoding.
This tool doesn't guess blindly. It re-decodes your pasted text under every plausible combination of UTF-8, Shift_JIS, EUC-JP, and JIS (ISO-2022-JP), then scores each result by how closely it resembles real Japanese or ASCII. That said, not every case is fixable: if the original bytes were already discarded during the very first bad decode, you'll typically see repeating "?" or "�" characters, and no amount of re-decoding can bring back information that's already gone. If you already know exactly which two encodings are involved, the manual tab skips the guessing entirely.
How to use the mojibake fixer
- Paste the garbled text Drop the broken text into the "Auto-fix mojibake" tab. Everything is processed locally, so nothing leaves your browser.
- Check the ranked candidates Look through the list of repair candidates and find the one labeled with the "Best match" badge — it's usually the correct reading.
- Switch to manual mode if needed If none of the automatic candidates look right, open the "Manual conversion" tab and explicitly choose the "appears as" and "actually is" encodings yourself.
- Reopen the original file if it still fails When even manual conversion can't recover readable text, the bytes may already be lost. Go back to the source file and reopen it in a text editor with the correct encoding set.
Tips for getting more out of it
- When multiple candidates appear, the one marked "Best match" — usually the most natural-looking Japanese or ASCII — is normally the correct one.
- The classic "縺薙繧薙?..." pattern seen in emails and CSV files almost always resolves with the UTF-8 → Shift_JIS pair.
- If the text has turned into a string of "?" or "�" characters, part of the original byte sequence has likely been lost, making a full text-level recovery difficult.
- Use the "Manual conversion" tab for a one-step result when you already know which encoding pair caused the garbling.
When this comes in handy
A partner's CSV file won't display correctly
When a spreadsheet opens a CSV from a business partner and shows garbled characters, paste one of the broken cells here first to confirm the real content before you ask them to resend the file.
An old email is unreadable
Subject lines or bodies from older or unusual mail clients sometimes decode incorrectly. You can often recover the original text yourself instead of writing back to the sender.
Server logs or database records look broken
When a system's encoding setting changed at some point in its history, older log lines or database rows can end up mojibake while newer ones are fine. This tool helps you read the affected records without re-exporting everything.
Learning how encoding mismatches behave
Try the manual tab with different "appears as" and "actually is" combinations to see firsthand which mismatches are recoverable and which patterns of garbling they each produce.
Glossary
- Character encoding
- The rule that maps characters to specific byte sequences (and back again). Text data is really just bytes; without agreeing on the encoding, there's no way to know which characters those bytes represent.
- UTF-8
- The dominant encoding on the modern web, capable of representing virtually every character in every language. Japanese characters typically take up three bytes each in UTF-8.
- Shift_JIS
- A Japanese encoding that predates UTF-8 and was the long-time default on Windows in Japan. Still found in older systems and some legacy file formats, it's also the encoding most often confused with UTF-8.
- EUC-JP
- A Japanese encoding historically favored on Unix and Linux systems, including many older Japanese web servers. It uses different byte patterns from Shift_JIS, so mixing the two up tends to produce especially severe garbling.
- JIS (ISO-2022-JP)
- An encoding built entirely from 7-bit characters, historically the standard for Japanese email because it could pass safely through mail systems that only handled 7-bit data. It's less common in web content today.
- Irreversible mojibake
- A case where the original bytes were already thrown away during the very first incorrect decode. It usually shows up as a run of "?" or "�" characters, and because the source information itself is gone, no string-level tool can restore it.
Frequently Asked Questions
Side Note — Why "縺薙繧薙?縺ォ縺。縺ッ" became the face of Japanese mojibake
Search for Japanese mojibake and you will almost certainly run into "縺薙繧薙?縺ォ縺。縺ッ" (which should read "こんにちは," or "hello") — it has become something of a cultural icon among Japanese internet users. It's the textbook result of opening UTF-8 text in an application that assumes Shift_JIS, such as an old version of Windows Notepad or various legacy systems.
This particular pairing shows up constantly because a Shift_JIS decoder forcibly reinterprets the 3-byte sequences UTF-8 uses for Japanese characters as if they were 2-byte Shift_JIS characters. By coincidence, most of those reinterpreted sequences happen to fall within the range of valid Shift_JIS characters, so no error occurs — the text just looks garbled while secretly remaining fully recoverable. That property is exactly why an automatic fixer like this one succeeds so often.
By contrast, opening Shift_JIS text as if it were EUC-JP tends to produce invalid byte sequences that get discarded outright, destroying information the moment it happens — recovery at the string level then becomes impossible in principle. This split between "recoverable" and "unrecoverable" garbling is a big part of why Japanese character-encoding issues have quietly frustrated engineers for decades.