🧬

Unicode Escape Converter

Convert text to \uXXXX / \UXXXXXXXX escapes and back, with correct surrogate pairs.
Unicode 📂 Encode 👤 ToolWorld 🏷️ v1.1.12
Loading...
Direction
Format
Plain text {chars} chars / {words} words / {lines} lines
Escaped result 0

How to Use

📖 Tool Introduction


The Unicode Escape Converter is a free online transcoding tool that moves between human-readable text and the backslash escape sequences you meet constantly in source code. You have certainly seen strings like \u4f60\u597d and \U0001F600: they are the way JavaScript and Python respectively write out Unicode characters. This tool turns ordinary text into such escape sequences and restores escape sequences back to the original text. The crucial part is how it handles surrogate pairs: characters whose code point exceeds FFFF, such as emoji, must be split into a high and a low surrogate unit under the \uXXXX convention, for example the grinning face becomes \uD83D\uDE00. The tool splits them correctly on encoding and merges the two units back into one character on decoding, so you never end up with a broken half character. Besides the four-digit \uXXXX form, it also supports the eight-digit \UXXXXXXXX code-point notation used by Python and Java. You may choose to leave printable ASCII untouched so the output stays clean. Everything is computed locally in the browser, which means no network request and no upload of any kind. The two sliding pills on top pick the direction and the notation, and the whole workflow from paste to copied result takes only a few seconds, which makes this a favourite tool for front-end and script engineers who inspect and author escaped literals every day. It is equally handy when you need to embed emoji inside a JSON configuration that is transmitted as plain ASCII, or when you want to prove to a colleague that a mysterious box glyph really stands for a specific code point rather than a missing font. The keep-ASCII option is especially useful for long English sentences, because it leaves the ordinary letters readable while still escaping every symbol that would otherwise break a source file. Decoding is just as forgiving: a string copied out of a config file rarely needs manual cleanup before it comes back to life as normal readable text.

✨ Key Features


  • Correct surrogates: above-BMP chars like emoji split into high/low pairs automatically
  • Two notations: \uXXXX (JS style) and \UXXXXXXXX (Python style)
  • Bidirectional: one sliding pill escapes or restores text
  • Keep ASCII: optionally leave printable ASCII unescaped for readability
  • Live counters: character statistics update as you type
  • Fully local: pure browser computation, no network calls

📝 How to Use


  1. Use the direction pill to pick escape or restore
  2. Use the format pill to pick \uXXXX or \UXXXXXXXX
  3. Tick whether to keep printable ASCII unescaped
  4. Paste text or an escape string into the input area
  5. Click the button and copy the result from the output

⚠️ Notes


  • Digit count: \u takes exactly four hex digits, \U exactly eight
  • Pairing: a lone high or low surrogate cannot form a char and is kept as-is
  • Prefix only: only \u and \U prefixes are recognised; other backslash forms stay untouched
  • Language fit: \U is illegal in JS strings, choose the notation your target language expects
  • Long input: split very long payloads into chunks

Related

⏫︎