Character Encoding Viewer
| # | Char | Code point | UTF-8 bytes | UTF-16 units |
|---|---|---|---|---|
| The analysis appears here, one row per character | ||||
How to Use
📖 Tool Introduction
The Character Encoding Viewer is a free online tool that helps developers and encoding enthusiasts understand Unicode and how characters are actually encoded. The same character can look radically different under different encoding schemes: a Latin letter is one byte in both ASCII and UTF-8, a common Chinese character occupies three bytes in UTF-8, and an emoji needs two UTF-16 code units that together form a surrogate pair. When you are chasing garbled text, diagnosing an encoding mismatch, writing escaping logic, revising for an interview or preparing a teaching demo, you often urgently need a side-by-side table that goes character by character. This tool accepts any Unicode text and then splits it apart by code point, giving you three views of every single character: its Unicode code point as a U-plus hexadecimal number, its UTF-8 byte sequence after encoding, and its code units under UTF-16. It pays special attention to the boundary between the Basic Multilingual Plane and the supplementary planes: whenever a code point exceeds FFFF, it automatically extracts both the high and low UTF-16 surrogate units instead of wrongly counting the emoji as two separate characters. The table is laid out row by row, the character column shows the visible glyph directly, and the code point and byte columns use a monospace font so you can compare values line by line. You can also copy the whole table as tab-separated text and paste it into your notes or documents. Every calculation runs locally in the browser, so your text is never uploaded and you can safely paste any content for debugging. It is equally useful when a string looks fine on screen but breaks the moment it crosses a network boundary, such as a Chinese name turning into question marks after a redirect, or an emoji being miscounted as two broken boxes by a length check. By exposing the exact bytes and code units, the tool demystifies those confusing mismatches and lets you confirm that the encoding on both ends of your system actually agrees.
✨ Key Features
- Per-character breakdown: iterates by Unicode code point and handles surrogate pairs
- Three views at once: code point, UTF-8 bytes and UTF-16 units
- Visible glyphs: the character column shows the actual glyph for easy recognition
- Hexadecimal consistency: code points and bytes are shown in uppercase hex
- Table export: copy the whole table as tab-separated text in one click
- Fully local: your text never leaves the browser, keeping it private
📝 How to Use
- Type or paste the text you want to inspect into the input box
- Click Analyze, or press Ctrl+Enter
- Read the table row by row to see each character's three encodings
- Notice how astral characters such as emoji show two UTF-16 units
- Click Copy table to save it as tab-separated text for your notes
⚠️ Notes
- Surrogate pairs handled: characters above FFFF use two UTF-16 units and are correctly merged into one row
- Byte order: UTF-16 units are shown to explain code point structure, not for transmission
- Code point width: points in the BMP use four hex digits, while astral points use five
- Invisible chars: spaces, tabs and other invisible characters are listed by their code points
- Very long text: avoid pasting extremely long input in one go to keep the table readable