Extract Words by Length
Paste text, set an inclusive minimum and maximum length, and receive the first occurrence of each matching word.
A QUICK WALKTHROUGH
How to use this tool
- Paste up to 50,000 characters.
- Set inclusive code-point limits from 1 to 1,000.
- Extract, copy, or download the matching words.
Unicode word tokens and length
Word tokens are the word-like segments returned by Intl.Segmenter with granularity word. Punctuation and whitespace are not tokens. Length counts Unicode code points with [...token].length, so it is not UTF-16 units and not displayed grapheme clusters.
Inclusive filtering and stable deduplication
Minimum and maximum limits are inclusive whole numbers from 1 through 1,000. Filtering happens while scanning input order; after a token passes the range, exact duplicate text is skipped and the first occurrence remains. Case variants remain distinct.
Local processing
The text stays in the current browser and is never uploaded. Inputs over 50,000 UTF-16 characters are rejected.
GOOD TO KNOW
Common questions
Are limits inclusive?
Yes. A token whose Unicode code-point length equals either limit is included.
How are duplicates removed?
Exact token text is deduplicated after length filtering, preserving the first occurrence in input order. Case differences remain separate.
Is an emoji one character?
Length uses Unicode code points. A visible emoji sequence can contain multiple code points, while grapheme clusters are not used.