Text tool

Character Encoding Detector

Paste text or bytes to inspect browser-detectable character encoding evidence and decode a safe preview.

In-browser processingNo account requiredPrivacy details ↗

Character encoding detector

This browser-native tool reports BOM, ASCII, and strict UTF-8 evidence. It does not guess legacy encodings or convert arbitrary text.

Enter text or bytes to inspect.

Detected evidence

Preview (replacement shown for malformed bytes)

A QUICK WALKTHROUGH

How to use this tool

  1. Choose text or hex bytes.
  2. Paste up to 100,000 characters and inspect locally.
  3. Read the candidates, byte evidence, and explicit limitations.

What is actually detected

The tool identifies UTF-8, ASCII, UTF-16LE/BE, and UTF-32LE/BE when a recognized BOM is present. Without a BOM it validates strict UTF-8 and labels ASCII as a subset, rather than claiming the original encoding.

Legacy encodings are not guessed

ISO-8859-1 and Windows-1252 can share the same bytes and browser TextDecoder support is not a universal detector. This page does not infer them from language or byte frequency, and it does not convert arbitrary encodings.

Invalid input and replacement

Hex input must contain complete byte pairs. Strict UTF-8 validation rejects malformed sequences; the preview uses replacement behavior only to show a readable diagnostic, never as proof that the source was valid.

Local and bounded

Text and hex input are limited to 100,000 characters. Nothing is uploaded; this browser-native implementation handles byte inspection and UTF-8/UTF-16 preview only.

GOOD TO KNOW

Common questions

Does this identify every text encoding?

No. It reports BOM evidence, ASCII, and strict UTF-8 evidence. No-BOM legacy encodings remain ambiguous.

What happens to an invalid UTF-8 sequence?

It is marked invalid. The optional preview replaces malformed bytes with U+FFFD so the diagnostic remains visible.

Can I convert ISO-8859-1 or Windows-1252 text?

No. This tool detects limited evidence and previews Unicode decodings; it is not an arbitrary encoding converter.