Does Grammarly Detect AI Text? Official Claims vs Independent Tests
Grammarly claims 99% accuracy and #1 on the RAID benchmark. Independent testing (PCWorld) found 37% on a known AI-generated text. Here's what each source actually shows.
Grammarly added a dedicated AI Detector to its editor, separate from the Humanizer agent, and it's easy to conflate the two since both live inside the same account. This piece is about the detector only: what it actually checks, what Grammarly itself claims about its accuracy, and what happens when someone runs a known AI-generated text through it. Short answer up front — yes, it detects AI text, and how well depends heavily on which source you trust.
How Grammarly's AI Detector actually works
Paste text in, and the detector segments it into sections and checks each one against patterns associated with AI generation — sentence structure and predictability, repetition and uniformity, and comparison against known outputs from models it's aware of. Grammarly's own page lists ChatGPT, Gemini, and Claude as models it identifies text from. It returns a single percentage: how much of the submitted text appears AI-generated.
A free scan is available to anyone. Grammarly Pro unlocks an "AI Detector agent" with more detail — phrase-level identification of which specific sections triggered the score, plus rewrite and citation suggestions. Grammarly's own support documentation is direct about the ceiling here: "no AI detector is 100% accurate," and the tool "should never be used as a standalone verification method." That disclaimer is worth taking seriously, because it's not just boilerplate — the testing below backs it up.
The official claim: 99% accuracy, #1 on RAID
Grammarly's marketing cites two numbers: 99% detection accuracy, and a #1 ranking on what it calls "RAID's independent benchmark." RAID is real — it's not a made-up credibility prop. Researchers from the University of Pennsylvania, University College London, King's College London, and Carnegie Mellon built it as the largest public benchmark for AI-text detection, over 10 million documents across 11 language models, 11 genres, and a dozen adversarial attack types, published at ACL 2024. Citing it is a legitimate move — RAID is exactly the kind of shared, academically-reviewed dataset the detection industry needed.
Here's the part worth flagging honestly, though: Grammarly is not the only detector vendor citing a top RAID result. Originality.ai — a direct competitor in this exact category — has published its own claim of ranking highest on the same benchmark. When two competitors both claim the top spot on the same academic dataset, that's a reason to treat vendor-stated "#1" claims as marketing framing of a real study, not as a settled, independently-confirmed fact. RAID itself, as a benchmark, is credible. Any single company's self-reported ranking on it deserves a second look before you take it at face value — ours included, for whatever it's worth, since we don't cite a specific detector ranking anywhere on this site.
What independent testing found
PCWorld ran a direct test in late 2024: a short story written entirely by Google Gemini, submitted to Grammarly's AI Detector twice to check consistency. Both runs came back at 37% — on text that was, by construction, 100% AI-generated. The tester's own reaction captures it: "This story is a complete fabrication by Google Gemini, so seeing such a low percentage is surprising." The result was at least consistent — the same score twice suggests the tool isn't just noisy, it's calibrated to under-report on this kind of text specifically.
Other independent reviews report a similar pattern of Grammarly running conservative — meaning it tends to under-flag AI content rather than over-flag human writing. That's a real design tradeoff, not necessarily a flaw: a detector that leans conservative produces fewer false accusations against genuine human writers, at the direct cost of missing more actual AI text. Whether that tradeoff is right for you depends on what you're using the detector for — screening your own drafts before submission is a different use case than an institution trying to catch violations at scale.
Paraphrased AI content is where the gap widens further across most reviews we found: text that's been run through a paraphrasing tool before hitting Grammarly's detector gets missed noticeably more often than raw, unedited AI output. That tracks with how most detectors work generally — they're trained on the statistical fingerprint of raw model output, and a paraphrasing pass changes enough of that fingerprint to slip under a conservative threshold.
AI Detector vs. AI Humanizer — these are different features
Let's clear this up directly, because it's a genuinely common mix-up: Grammarly's AI Detector and Grammarly's AI Humanizer are two separate tools on two separate pages, and searching for one often surfaces content about the other. The Detector scans text and estimates AI likelihood — it's a checker. The Humanizer rewrites AI-drafted text to sound more natural — it's a rewriter. You can use either without the other; they don't share a workflow. If you landed here wanting the rewriting tool specifically, we cover Grammarly's Humanizer feature and pricing separately.
What triggers a high score
Based on Grammarly's own description of the mechanism, three things move the number: uniform sentence structure (LLMs tend to produce sentences that sit in a narrow length band), repetitive phrasing patterns, and similarity to known outputs from the specific models the detector is trained to recognize. None of this is unique to Grammarly — it's the same general family of signals every major detector reads, described in more technical depth on our own breakdown of what detectors actually flag.
There's a limitation here if you're testing very recent AI output: detector training data has a cutoff, and models evolve. A detector tuned on older model outputs can miss newer models whose statistical fingerprint has shifted, which is part of why Grammarly's own disclaimer about non-100%-accuracy holds up under testing rather than reading as legal cover-your-bases language.
Want your AI draft to read naturally instead of gaming a detector score?
Refrazr rewrites AI drafts into natural, human-sounding writing — one clean pass, meaning fully preserved, no detector-bypass claims. Free tier is 500 words/day.
Try Refrazr free → Check your own AI score firstBottom line
Grammarly's AI Detector is real, free to use at a basic level, and built on a legitimate methodology — the RAID benchmark it cites is a credible academic dataset, whatever you make of the "#1" framing. But the gap between the marketed 99% and PCWorld's observed 37% on a known-AI text is large enough that treating the detector as a pass/fail verdict on anything that matters — academic, professional, or otherwise — would be a mistake. Grammarly's own documentation says as much. Use it as one data point, run a second detector if the stakes are real, and remember that a low score doesn't clear a text any more than a high score convicts one.
Related reading
Frequently asked
Does Grammarly detect AI-generated text?
How accurate is Grammarly’s AI detector?
What’s the difference between Grammarly’s AI Detector and AI Humanizer?
Is the RAID benchmark Grammarly cites legitimate?
Can paraphrased AI text fool Grammarly’s detector?
Does using Grammarly’s grammar-check features get my writing flagged as AI?
Is Grammarly’s AI detector free?
Try it free
Humanize your text now
500 words free every day. No sign-up required to try. Paste your AI draft and see how it reads rewritten.
Need more words? View pricing — packs from $1.99, never expire.