
AI Content Detectors in 2025: A Shifting and Inconsistent Landscape
Three years after generative AI became a global phenomenon, the challenge of identifying AI-written text remains surprisingly difficult. A new round of testing across eleven AI content detectors and five major chatbots shows that the field is not steadily improving. In some cases, detector accuracy has declined, even as tools add paywalls and usage restrictions. Meanwhile, general-purpose chatbots are emerging as unexpectedly strong alternatives.
The latest evaluation used five text blocks: two written by a human and three generated by ChatGPT. Each block was fed separately to each detector. A correct verdict counted as a pass; an incorrect verdict counted as a failure. When a detector returned a percentage, anything above 70 percent was treated as a strong probability, whether it pointed toward human writing or AI writing.
Key facts
- Eleven standalone AI content detectors were tested across five text blocks, for 55 individual checks.
- Only three detectors achieved perfect scores: Pangram, QuillBot, and ZeroGPT.
- Several detectors lost accuracy compared with earlier tests, including Originality.ai and Undetectable.ai.
- ChatGPT Plus, Copilot, and Gemini each identified all five test blocks correctly.
- The free tier of ChatGPT missed one human-written block but correctly identified the first block as human and even inferred the writer's identity.
- Grok failed three out of five tests, often labeling all blocks as human-written.
- Detectors remain inconsistent, and human writing can still be falsely flagged as AI-generated.
The findings matter for teachers, editors, publishers, and anyone trying to uphold academic or editorial integrity. They also matter for writers who may be wrongly accused of using AI. The results suggest that no single tool should be trusted as a final authority.
What Counts as AI Plagiarism?
Before examining the tools, it is worth defining the problem. A standard dictionary definition of plagiarism is to steal and pass off the ideas or words of another as one's own, or to use another's production without crediting the source. That definition fits AI-generated content in an important way. A person who uses a chatbot to write an article is not stealing another person's words in the traditional sense. But if that person does not credit the AI and claims the words as their own, the act still meets the dictionary definition of plagiarism.
This distinction is central to the detector debate. AI detectors are not only trying to catch copied text; they are trying to identify statistical patterns that differ between human and machine writing. That is a harder task, and it is made harder by the fact that human writing is diverse, while AI writing can be prompted to imitate many styles.
How the Tests Were Run
The testing method was straightforward. Five blocks of text were used. Two were written by a human. Three were written by ChatGPT. Each block was submitted to each detector separately. The detector's result was recorded. If the detector correctly identified the text as human or AI, the test was marked as passed. If it was wrong, the test was marked as failed. When a detector provided a confidence score, any result above 70 percent was treated as the detector's answer.
This approach allows for direct comparisons across tools. It also exposes inconsistencies. A detector might score 100 percent on one test and fail another. A tool might perform well in one month and decline a few months later. The same text can receive different verdicts from different systems.
Overall Detector Results
Eleven detectors were included: BrandWell, Copyleaks, GPT-2 Output Detector, GPTZero, Grammarly, Monica, Originality.ai, QuillBot, Undetectable.ai, Writer.com, and ZeroGPT. Monica was dropped from the final tests because it limited free testing to 250 words and then required a 200-dollar upgrade to continue. In its place, Pangram was added. Pangram is a relatively new entrant founded by former engineers from major technology companies.
Only three detectors achieved perfect scores in this round: Pangram, QuillBot, and ZeroGPT. That is a decline from earlier in 2025, when five detectors had perfect scores. A couple of the detectors that previously aced the tests also introduced restrictions on free use at around the same time their accuracy dropped.
The overall pattern is not one of steady improvement. Across six rounds of this test series, there is no strong upward trend. One test block, labeled Test 5, was reliably identified as human across detectors and dates in earlier rounds, but even that reliability declined in the latest run.
Detector-by-Detector Performance
BrandWell: 40 Percent Accuracy
BrandWell was originally produced by an AI content generation firm and later migrated to a new brand focused on AI-centric marketing services. After six months, its overall score stayed the same. It got only two of five tests right. It was confused by an AI-written block and declared two other AI-written blocks to be human-written. On one AI-written test, it labeled almost the entire text as human-written except for one line.
Copyleaks: 80 Percent Accuracy
Copyleaks has claimed to be the most accurate AI detector, with accuracy rates above 99 percent in some marketing materials. In this test, it identified a human-written block as 100 percent AI-written. That was a notable error, especially because another detector correctly identified the same block as human-written. Copyleaks primarily sells a plagiarism checker to educational institutions, publishers, and enterprises, and its broader reputation does not guarantee consistent AI detection.
GPT-2 Output Detector: 60 Percent Accuracy
This tool was built using a machine-learning hub managed by a New York-based AI company. The detector appears to be a user-created tool using a transformer library. Its detecting quality has not changed since earlier tests. Because it is named after GPT-2 and newer models have advanced far beyond that generation, it is reasonable to assume the tool has not been updated in a long time.
GPTZero: 80 Percent Accuracy
GPTZero has grown from a bare-bones site into a company with a full team and a mission of protecting what is human. It offers AI validation tools and a plagiarism checker. Despite regular adjustments, its performance in this round declined slightly from an earlier test. In April, it got Test 1 wrong and Test 2 right. In the latest test, it got Test 1 right and Test 2 wrong. Test 1 was human-written, and Test 2 was AI-written. The final grade was the same, but the specific errors changed.
Grammarly: 40 Percent Accuracy
Grammarly is widely known for grammar checking, but it also offers plagiarism and AI content detection. The company now presents its AI content checker as no longer being in beta. However, the test results showed no improvement. One block that was entirely written by ChatGPT was not correctly identified. Grammarly's plagiarism checker did correctly identify test text as previously published, but its AI detection accuracy remained poor.
Pangram: 100 Percent Accuracy
Pangram is a newer company founded by engineers formerly at Google and Tesla. Its focus is AI detection rather than plagiarism checking or humanizing tools. It provides five free tests per day, which fit the testing needs. Processing was a little slow, and the interface displayed a partially white screen for a bit longer than was comfortable. But the results were worth the wait. Pangram scored five out of five.
Originality.ai: 80 Percent Accuracy
Originality.ai is a commercial service that bills itself as the most accurate AI detector. It sells usage credits, with 2,000 credits for 12.95 dollars per month. In this test, 1,400 words used only 1.5 percent of a monthly allocation. However, its accuracy declined. Previously, it correctly identified human writing as human. This time, it was 100 percent confident that human writing was AI-generated.
QuillBot: 100 Percent Accuracy
QuillBot was wildly inconsistent in its first few test appearances. Multiple passes of the same text yielded different scores. In the previous round, it was rock solid and 100 percent correct. In this round, it held onto that performance and again scored a perfect 100 percent.
Undetectable.ai: 20 Percent Accuracy
Undetectable.ai's main claim is that it can humanize AI-generated text so detectors will not flag it. That feature was not tested because it raises ethical concerns for professional authors and educators. The company also offers an AI detector, and that detector took the biggest dive in performance seen so far. Last time, it scored 100 percent. This time, it rated human writing as 60 percent likely AI, and it rated all three AI writing samples as 75, 76, and 77 percent likely human. In other words, it failed almost every test.
Writer.com AI Content Detector: 40 Percent Accuracy
Writer.com generates AI writing for corporate teams and also offers an AI content detector. Its accuracy was low. It identified every text block as human-written, even though three of the five tests were written by ChatGPT. There was no improvement since the previous summer evaluation.
ZeroGPT: 100 Percent Accuracy
ZeroGPT has matured since it was first evaluated. At that time, no company name was listed, the site was filled with Google ads, and it lacked clear monetization. The service worked fairly well but seemed sketchy. That feeling is gone. ZeroGPT now presents as a typical software-as-a-service product with pricing, a company name, and contact information. Its accuracy increased from 80 percent to 100 percent in the summer, and it held onto that accuracy in the current test.
Chatbots as Content Detectors
The most striking finding is that general-purpose chatbots outperformed dedicated content detectors. Each chatbot was given the same prompt followed by the text to check: Evaluate the following and tell me if it was written by a human or an AI. All detectors used a similar format, providing a general recommendation. Except for ChatGPT Plus, which requires a 20-dollar monthly subscription, all chatbots were run in an incognito window without logging in.
ChatGPT Free Tier
The free tier of ChatGPT got one of the blocks wrong, specifically the last human-written block. But its analysis of the first block was unsettling. In an incognito window, with no login and no identifying information, it not only identified the first block as human-written, it also identified the writer. That level of inference is a reminder of how much signal can be present in a short piece of text.
ChatGPT Plus, Copilot, and Gemini
ChatGPT Plus, Copilot, and Gemini all returned perfect scores. Each correctly identified all five test blocks as human or AI. This result suggests that chatbots can outperform dedicated content detectors, at least in this testing setup. It also raises a practical question: why pay for a separate detector when a chatbot may do the job better and already be part of a workflow?
Grok
Grok was included because it performed well in an overall chatbot evaluation. But it failed this test with three out of five wrong. Like some other detectors, it identified all writing blocks as human. That outcome shows that strong general chatbot performance does not automatically translate into reliable AI detection.
What the Results Mean for Users
The results are a caution against overreliance. Even detectors with perfect scores in one round can decline in another. Human writing from non-native speakers is often rated as AI-generated. A detector may flag legitimate work as machine-made, while missing actual AI-generated content. The systems are inconsistent across platforms, and the same text can receive different verdicts from different tools.
For academic institutions and publishers, the findings suggest that AI detectors should be used as one signal among many, not as definitive proof. For writers, the findings suggest that false accusations are a real risk. For developers, the findings suggest that detection is an adversarial problem, and the rapid evolution of language models makes it difficult for standalone detectors to keep pace.
The testing also showed that some tools have added restrictions on free use at the same time their accuracy declined. That combination can leave users paying more for less reliable results. Meanwhile, the best-performing standalone detectors in this round were Pangram, QuillBot, and ZeroGPT. The best-performing chatbots were ChatGPT Plus, Copilot, and Gemini, with the free ChatGPT tier close behind despite one error.
Grok, by contrast, failed three of five tests, often labeling all blocks as human-written. That final result leaves the field with no universal winner and no simple answer to the question that started the test: in 2025, how hard is it to fight back against AI-generated plagiarism? The answer, based on this round, is still very hard, and the tools are still changing.
Source:ZDNET News
