What Is a Similarity Score?
A similarity score tells you what percentage of your text matches existing sources — and why that number alone doesn’t tell the full story.
Check Your Paper FreeA similarity score is the percentage of your text that matches existing sources in a plagiarism checker’s database. It shows how much of your writing overlaps with other content — but it does not automatically mean plagiarism.
How a Similarity Score Is Calculated
Plagiarism checkers don’t read your writing the way a human does. Instead, they break your text into small overlapping fragments — typically sequences of 8 to 12 consecutive words called n-grams — and compare each fragment against a large database of web pages, published journal articles, and previously submitted student work.
The resulting score represents the percentage of your text that matches at least one source in that database. A 20% similarity score means roughly one-fifth of your document contains strings of words found verbatim somewhere else.
Important: Different tools scan different databases. The same document can return 12% in one checker and 28% in another — not because one is wrong, but because their source libraries differ. Never compare scores across platforms; compare within the same tool over time.
Similarity Score vs. Plagiarism — They Are Not the Same
This is the most misunderstood aspect of plagiarism checking, and it matters enormously. A similarity score is a technical measurement. Plagiarism is an ethical and often legal violation. The two overlap, but they are not the same thing.
Similarity means: text in your document matches text elsewhere. Plagiarism means: you presented someone else’s ideas or words as your own without giving credit.
High similarity that is not plagiarism
- Properly formatted quotations — a cited quote from a source is correctly attributed, even if it triggers a match.
- Bibliography and reference lists — most modern checkers exclude these automatically, but not all do.
- Standard methodology phrases — formulas like “This study aims to…” or “Data were collected using…” appear in thousands of papers and don’t constitute copying.
- Self-citation of your own previous work — overlaps with your own earlier papers may flag as similarity but aren’t plagiarism in the traditional sense (though some institutions have self-plagiarism policies).
The counter-intuitive case: A document with only 5% similarity can still contain real plagiarism — if that 5% is one uncited paragraph lifted directly from a key source. A high number isn’t always damning; a low number isn’t always safe.
What Is a Good Similarity Score? Thresholds by Document Type
There is no single universal threshold. Different institutions, journals, and content types apply different standards. The table below reflects widely cited guidelines — but always check your own institution’s policy.
| Document Type | Generally Acceptable | Needs Review | Concern |
|---|---|---|---|
| Undergraduate essay | < 15% | 15–25% | > 25% |
| Master’s thesis | < 10% | 10–20% | > 20% |
| PhD dissertation | < 10% | 10–15% | > 15% |
| Journal article | < 5% | 5–10% | > 10% |
| Blog / web content | < 20% | 20–30% | > 30% |
These are widely cited guidelines, not universal rules. Always check your institution’s specific policy.
Why 0% Is Not Always Good
A zero similarity score might sound ideal, but in academic writing it can actually raise red flags. Proper academic work requires citing sources — if a dissertation or thesis shows 0%, it may mean the student used no external literature at all, which is a problem in itself.
A second concern: some students rewrite text specifically to avoid detection, stripping out citations along the way. A 0% result achieved through excessive paraphrasing without attribution is still academic misconduct — it just looks cleaner on a report.
Most experienced reviewers expect to see a similarity score of roughly 5–15% in a well-written, properly sourced academic paper. A small, managed overlap — mostly from cited material and standard phrases — is a sign of normal scholarly engagement, not a warning sign.
Factors That Artificially Inflate Your Score
Before panicking about a high score, check whether any of these common sources are skewing the result:
- Reference list included in the scan — if the checker didn’t exclude your bibliography, every cited title adds matches. Export the report and check whether references are flagged.
- Assignment prompt copied into the document — if you pasted the question or brief at the top, it will match any other student who did the same.
- Standard academic methodology language — phrases common across your discipline will match hundreds of papers. These are not plagiarism.
- Previously submitted versions of the same work — if you submitted a draft earlier, newer submissions may match against your own prior file.
- Common definitions and widely known facts — stating that “photosynthesis is the process by which plants convert sunlight into energy” will generate a match. It is not plagiarism.
How Different Tools Calculate Similarity Scores Differently
This is something most guides don’t explain clearly, and it’s the reason students are often confused when they get wildly different numbers from different checkers.
Every plagiarism tool has its own database — and the composition of that database is the single biggest driver of your score. Turnitin, for example, has indexed over 1.6 billion student submissions in addition to its web and journal archives. A paper submitted to Turnitin is matched against that enormous pool; the same paper run through a smaller tool with a web-only index will produce a much lower percentage simply because fewer matching sources exist in the database.
The n-gram length used for comparison also varies between tools, as does how they handle formatting, punctuation, and excluded content like bibliographies. A match that one tool filters out as boilerplate, another may count.
The practical takeaway: don’t compare your score between platforms. If your institution uses Turnitin, Turnitin’s score is the one that matters. Running your paper through a different tool first won’t give you a preview of that number — it gives you a different measurement of a different thing.
WriteCheck.pro provides a free baseline similarity check that’s useful as a first-pass review before your final submission — it helps you identify obvious overlaps and flag sections worth paraphrasing or re-citing before you submit anywhere official.
What To Do If Your Score Is Too High
- Open the full similarity report — don’t react to the percentage alone. The report shows you exactly which fragments matched and which source they matched against.
- Identify whether flagged sections are properly cited — if a highlighted passage is a direct quote with a citation, it may be fine. Your institution may allow you to request those sections be excluded from the score.
- Check whether your reference list was included — if bibliography entries are in the matched section, ask your instructor or submission system whether that’s expected, or resubmit with references excluded.
- Paraphrase uncited overlapping sections in your own words — for passages that aren’t quotes and don’t have citations, rewrite them. The goal isn’t to beat the algorithm — it’s to express the idea yourself.
- Don’t attempt to spoof the checker — substituting look-alike Unicode characters, adding invisible whitespace, or writing in white text on a white background are all flagged by modern tools. Instructors who see an anomalous detection pattern may escalate more seriously than if you’d just had a high score.
Not sure where your paper stands? Get a free similarity score on WriteCheck.pro — no account needed.
Check My Paper FreeFrequently Asked Questions
Check Your Paper — Free
Get your similarity score in seconds. No account, no upload limits, no paper stored.
Run a Free Check Now