Unicode Character Inspector
Paste text to see how visible characters relate to Unicode code points, UTF-16 units, and grapheme clusters.
UNICODE CHARACTER INSPECTOR
Paste text to inspect each Unicode code point, encoding unit, and grapheme cluster.
Waiting for text.
Input is limited to 100,000 UTF-16 code units; the report shows at most 2,000 rows.
| № | Text | Code point | UTF-16 | Scalar | Grapheme | Name |
|---|
A QUICK WALKTHROUGH
How to use this tool
- Paste up to 100,000 UTF-16 code units.
- Inspect each code point and its grapheme position.
- Use the report to distinguish encoding units from scalar values and visible clusters.
Three useful boundaries
A Unicode code point is a scalar value or an unpaired surrogate in the input. UTF-16 code units describe JavaScript string storage. Grapheme clusters approximate user-perceived characters and may contain several code points.
Names are deliberately conservative
The report does not invent or guess Unicode character names. It shows a name only when a reliable bundled or system source is available; this build has no such source, so names are marked unavailable.
Local limits
Input is limited to 100,000 UTF-16 code units and the report to 2,000 grapheme rows. Processing stays in this browser.
GOOD TO KNOW
Common questions
Is a grapheme cluster the same as a code point?
No. Combining marks, emoji sequences, and joined scripts can make one grapheme from several code points.
Why can a code point use two UTF-16 units?
Supplementary Unicode values above U+FFFF are encoded as a surrogate pair in UTF-16.
Are character names always shown?
No. The tool refuses to guess names when no reliable local data source is available.