Extract Unique Words
Paste text to keep one copy of each word while preserving the order in which it first appears.
Unicode letters and numbers are tokenized; punctuation and whitespace separate tokens.
A QUICK WALKTHROUGH
How to use this tool
- Paste up to 100,000 characters.
- Choose whether case should be significant.
- Extract and copy the first-appearance list.
Word tokenization
Words use browser Intl.Segmenter word granularity when available and keep segments marked word-like. The fallback keeps contiguous Unicode letters and numbers. Punctuation, symbols, and whitespace are separators and are not included.
Case and order
Case-sensitive mode compares the exact token. Case-insensitive mode compares each token after Unicode toLocaleLowerCase mapping, while the first observed spelling is retained. Results always stay in first-appearance order.
Unicode and whitespace
Unicode letters and numbers are supported through the browser tokenizer. Whitespace includes spaces, tabs, line breaks, and other separators; it only separates tokens and is not preserved in the result.
GOOD TO KNOW
Common questions
Are words sorted alphabetically?
No. Each unique token stays where it first appeared.
Does case matter?
Yes by default. Turn on Ignore case to merge tokens whose Unicode lowercase forms match.
Is my text uploaded?
No. Extraction runs locally in this browser.