Text Similarity Checker
Choose multiset Dice n-grams, sentence Jaccard between two texts, or single-text descriptive statistics.
A QUICK WALKTHROUGH
How to use this tool
- Paste a source and a text to check, or one text for statistics.
- Select a mode and its applicable n-gram size or threshold.
- Review metrics and sentence matches; cancel or edit to discard an in-progress calculation.
Sentence lexical overlap
Tokens are NFKC-normalized lowercase Unicode letter/number runs. Split sentences at . ! ? 。 ! ? or line breaks, including in abbreviations and decimals. Unspaced Chinese, Japanese and Korean may form long tokens; no language-specific segmentation or semantic comparison is provided.
Mode
Dice compares n-gram occurrence counts, using unigrams when either passage is shorter than n. Sentence mode selects each checked sentence’s best source-set Jaccard and reports the share meeting the threshold. Empty sets score 0 and cannot be flagged. Statistics show type/token ratio, unique four-token grams/all four-token grams, and population standard deviation of sentence token lengths/mean. Missing denominators or fewer than two CV samples show N/A.
Single-text statistics
Per text: 50,000 UTF-16 characters, 500 sentences and 5,000 tokens. Sentence mode: 250,000 pairs and a conservative 5,000,000 set-token work budget, with cooperative cancellation. Output: 500 sentence rows, 100 Dice matches and 1,000,000 characters.
Text Similarity Checker
These are lexical statistics, not semantic similarity, plagiarism or AI detection, or proof of originality. No API, upload, web search or persistent storage is used.
GOOD TO KNOW
Common questions
Can these results establish originality or plagiarism?
No. They describe supplied wording; no composite originality score or authorship probability is produced.