Character Encoding Detector
Paste text or bytes to inspect browser-detectable character encoding evidence and decode a safe preview.
Character encoding detector
A QUICK WALKTHROUGH
How to use this tool
- Choose text or hex bytes.
- Paste up to 100,000 characters and inspect locally.
- Read the candidates, byte evidence, and explicit limitations.
What is actually detected
The tool identifies UTF-8, ASCII, UTF-16LE/BE, and UTF-32LE/BE when a recognized BOM is present. Without a BOM it validates strict UTF-8 and labels ASCII as a subset, rather than claiming the original encoding.
Legacy encodings are not guessed
ISO-8859-1 and Windows-1252 can share the same bytes and browser TextDecoder support is not a universal detector. This page does not infer them from language or byte frequency, and it does not convert arbitrary encodings.
Invalid input and replacement
Hex input must contain complete byte pairs. Strict UTF-8 validation rejects malformed sequences; the preview uses replacement behavior only to show a readable diagnostic, never as proof that the source was valid.
Local and bounded
Text and hex input are limited to 100,000 characters. Nothing is uploaded; this browser-native implementation handles byte inspection and UTF-8/UTF-16 preview only.
GOOD TO KNOW
Common questions
Does this identify every text encoding?
No. It reports BOM evidence, ASCII, and strict UTF-8 evidence. No-BOM legacy encodings remain ambiguous.
What happens to an invalid UTF-8 sequence?
It is marked invalid. The optional preview replaces malformed bytes with U+FFFD so the diagnostic remains visible.
Can I convert ISO-8859-1 or Windows-1252 text?
No. This tool detects limited evidence and previews Unicode decodings; it is not an arbitrary encoding converter.