Text tool

Unicode Character Inspector

Paste text to see how visible characters relate to Unicode code points, UTF-16 units, and grapheme clusters.

In-browser processingNo account requiredPrivacy details ↗

UNICODE CHARACTER INSPECTOR

Paste text to inspect each Unicode code point, encoding unit, and grapheme cluster.

Waiting for text.

Input is limited to 100,000 UTF-16 code units; the report shows at most 2,000 rows.

A QUICK WALKTHROUGH

How to use this tool

  1. Paste up to 100,000 UTF-16 code units.
  2. Inspect each code point and its grapheme position.
  3. Use the report to distinguish encoding units from scalar values and visible clusters.

Three useful boundaries

A Unicode code point is a scalar value or an unpaired surrogate in the input. UTF-16 code units describe JavaScript string storage. Grapheme clusters approximate user-perceived characters and may contain several code points.

Names are deliberately conservative

The report does not invent or guess Unicode character names. It shows a name only when a reliable bundled or system source is available; this build has no such source, so names are marked unavailable.

Local limits

Input is limited to 100,000 UTF-16 code units and the report to 2,000 grapheme rows. Processing stays in this browser.

GOOD TO KNOW

Common questions

Is a grapheme cluster the same as a code point?

No. Combining marks, emoji sequences, and joined scripts can make one grapheme from several code points.

Why can a code point use two UTF-16 units?

Supplementary Unicode values above U+FFFF are encoded as a surrogate pair in UTF-16.

Are character names always shown?

No. The tool refuses to guess names when no reliable local data source is available.