Code tool

HTML Entity Encoder & Decoder

Choose basic encoding, non-ASCII encoding or decoding. Unicode is processed by code point, locally in your browser. Results are displayed as text.

In-browser processingNo account requiredPrivacy details ↗

Entity text is not an XSS sanitizer. Results are plain text; no HTML or JavaScript is executed.

Input stays in this browser; nothing is sent.

A QUICK WALKTHROUGH

How to use this tool

  1. Paste up to 100,000 UTF-16 code units of text or entities.
  2. Choose Decode, Basic encode, or Basic + non-ASCII encode.
  3. Convert, review the plain-text result, and select it to copy.

Two encoding modes

Basic mode replaces &, <, >, double quotes and apostrophes with &amp;, &lt;, &gt;, &quot; and &#39;. Non-ASCII mode also turns every code point above U+007F into a decimal numeric entity. Emoji are encoded as one code point, not two surrogate units. Existing entity ampersands are encoded again.

Native HTML decoding

Decoding uses the browser’s HTML character-reference rules in a detached textarea, in a single pass. Named, decimal and hexadecimal references are supported. Unknown names remain text; missing semicolons follow HTML rules. Zero, surrogate and out-of-range numeric values become U+FFFD; some control references are remapped according to HTML. Newline and invalid input normalization can prevent exact round trips. This is not a JavaScript escape parser.

Plain-text boundary

Only the detached textarea parses entity content; the visible result is assigned to textarea.value. No output is rendered as HTML. Encoding is context-dependent and does not sanitize HTML or make arbitrary content safe to insert in scripts, URLs or attributes. Output is limited to 1,000,000 UTF-16 code units.

GOOD TO KNOW

Common questions

How are the encoding modes different?

Basic encodes the five HTML-sensitive characters. Non-ASCII mode additionally encodes all non-ASCII code points as decimal references, including emoji.

What happens to invalid numeric references?

Native HTML decoding replaces zero, surrogates and values above U+10FFFF with U+FFFD, and remaps certain control values. It does not use JavaScript Unicode rules or promise lossless decoding.

Does this sanitize HTML or upload input?

No. It processes entity text locally and displays plain text. It does not provide XSS sanitization.