Text tool

Extract Words by Length

Paste text, set an inclusive minimum and maximum length, and receive the first occurrence of each matching word.

In-browser processingNo account requiredPrivacy details ↗

Up to 50,000 characters. Processing stays in your browser.

Paste text and choose a word-length range.

Word tokens follow Intl.Segmenter word boundaries. Length is Unicode code points; duplicate removal keeps the first occurrence in input order.

A QUICK WALKTHROUGH

How to use this tool

  1. Paste up to 50,000 characters.
  2. Set inclusive code-point limits from 1 to 1,000.
  3. Extract, copy, or download the matching words.

Unicode word tokens and length

Word tokens are the word-like segments returned by Intl.Segmenter with granularity word. Punctuation and whitespace are not tokens. Length counts Unicode code points with [...token].length, so it is not UTF-16 units and not displayed grapheme clusters.

Inclusive filtering and stable deduplication

Minimum and maximum limits are inclusive whole numbers from 1 through 1,000. Filtering happens while scanning input order; after a token passes the range, exact duplicate text is skipped and the first occurrence remains. Case variants remain distinct.

Local processing

The text stays in the current browser and is never uploaded. Inputs over 50,000 UTF-16 characters are rejected.

GOOD TO KNOW

Common questions

Are limits inclusive?

Yes. A token whose Unicode code-point length equals either limit is included.

How are duplicates removed?

Exact token text is deduplicated after length filtering, preserving the first occurrence in input order. Case differences remain separate.

Is an emoji one character?

Length uses Unicode code points. A visible emoji sequence can contain multiple code points, while grapheme clusters are not used.