The question gets a clean answer as soon as you stop treating Grammarly as one thing. What ships under that logo is a bundle of features running on different machinery, and the machinery is what a detector reacts to. One clarification before anything else, because it derails half the conversations on this topic: Grammarly also sells an AI detector, a separate product that scans text you paste into it and estimates whether a model wrote it, and that product is not the subject here. The subject is whether using Grammarly on your own writing changes how somebody else’s detector reads it afterwards, which is the question that put one Georgia undergraduate in front of her university’s integrity office.
Which Grammarly Features Get Flagged as AI?
Grammarly runs two kinds of machinery under one logo, and only one of them writes. On one side sit the red and blue underlines: spelling, grammar, punctuation, a clarity nudge. Grammarly’s own support answer on whether its use triggers detectors describes that technology as integral to the suggestions the product makes and says it “does not change the substance of the writing”. On the other side sit the generative features, which the same page says “can generate content and meaningfully change writing”. Two categories, one interface, and almost every argument about this topic is two people describing different halves of it.
The company is unusually direct about what happens to the second half. Grammarly’s AI Detector user guide states that if you work with its rewriting agents, its own detection agent will likely mark the result as machine-generated, because “rewrites from the Grammarly, Paraphraser, and Humanizer agents come from our LLM”. The same document draws the opposite line for the underlines: “traditional, non-generative corrections in the form of Grammarly red and blue underlines should typically not impact the percentage score”. Read the hedges rather than around them. Typically, and should, are not the same word as never, and the document is describing Grammarly’s own detector rather than Turnitin’s or GPTZero’s.
Then the part worth sitting with. On the marketing page for that detector , the company advertises that the product “Achieves 99% detection accuracy and ranks #1 on RAID’s independent benchmark”. In the support documentation for the same product, it says the score “should not be used as a definitive assessment of whether AI-generated text is present”. Both sentences are Grammarly’s, about one tool, written for two different rooms. Whichever of the two you believe, you have just been told by the vendor that a percentage is not a finding, which is the single most useful sentence on this page if somebody is currently holding a percentage over you.
The practical translation is short. What matters is not whether Grammarly was open while you worked. It is whether any sentence in the document came out of a model. Accepting a comma is a different act from accepting a rewritten paragraph, and the value of knowing which one you did is that it makes your account specific, checkable and boring, which is what a credible account looks like.
There is also a switch, and it matters most if your course draws the line at corrections only. That same support article is Grammarly’s own answer to how you turn those features off , and it is worth following in its current form rather than from anybody’s summary, because the feature set keeps moving and the instructions move with it. The point of doing it is not the number on somebody’s report. A generative feature you have switched off is a rule you cannot break by accident at one in the morning, and that is a sturdier plan than relying on your own restraint at the end of a long draft.
Nobody has published a test of this that you can check. The most granular public attempt found for this page is a July 2026 post on an anonymous detector-review blog, aibusted, which ran one 400-word sample through five Grammarly modes and five detectors and printed the grid. It carries no author, no link to the sample text, no screenshots and no raw data, and one passage is not a measurement, so none of its numbers appear here. If you want to see how a detector behaves on prose of known human origin, the honest move is to run the experiment on something of your own: put an old graded paper through a detector and read what comes back.
What Happened in the Marley Stevens Case?
Marley Stevens was a junior at the University of North Georgia when a criminal justice paper she submitted in October was flagged by Turnitin as AI-generated. She had used the free version of Grammarly to check her grammar and punctuation. She received a zero, was reported to the university’s Office of Student Integrity, and was placed on academic probation, which Fox 5 Atlanta reported in February 2024 was set to run for a year, until February 2025. Tech Times covered it the same week.
Her name appears here because she chose to make the case public and several outlets reported it with her own account in it. What is not here is a tidy ending, because there is not one on the record. Around two weeks after the first story, Fox 5 followed up with the news that Grammarly had invited her to make educational videos with the company, and quoted Grammarly saying that AI detection “as a whole, as a software entity, is faulty” and that she had used only its proofreading tool rather than anything generative. Weigh who is speaking there: Grammarly had a product interest in that framing, which does not make the statement untrue and does mean it is not an independent finding. The most recent reporting located for this page, Inside Higher Ed in November 2024, still refers back to the probation without saying whether it was ever lifted.
One case is not a rate, and nothing about this one establishes how often the same thing happens. What it does establish is the distance a person can travel between running a grammar check and sitting in front of an integrity office, and how little of that distance is covered by being right. It is also the clearest illustration of what the previous section is for. A provenance record switched on before the paper was written would have given her something to hand over, and it still would not have decided the outcome. Evidence that already existed beats an explanation produced afterwards, every time, which is why the timing matters more than the wording.
Does Turnitin Detect Grammarly?
Turnitin answers this question itself, by name, in its own documentation. Its AI writing detection FAQ poses the exact query people type into search engines and answers it: “If students use Grammarly for grammar checks, does Turnitin detect it and flag it as AI? No. Our detector is not tuned to target Grammarly-generated spelling, grammar, and punctuation modifications to content but rather, other AI content written by LLMs such as GPT-3.5.”
Read the scope of that as carefully as the answer. It covers grammar checks. It says nothing about a paragraph regenerated by a rewriting agent, which by Grammarly’s own account is model output, and it describes what the detector is aimed at rather than promising an outcome for every document. It is still the strongest single sentence available to a student in this position, because it comes from the company whose software produced the flag.
The reliability numbers around it need their dates attached. Turnitin’s claim of a “less than 1% false positive rate, to ensure that students are not falsely accused” comes from a blog post published on 16 March 2023 , which is more than three years old at the time of writing, predates several of the company’s own detector updates, and states no sample size and no test method anywhere on the page. In April 2024, Turnitin’s chief product officer Annie Chechitelli told EdSurge that traditional grammar-checking tools do not set off the company’s alarms, and described a deliberate trade: “We probably let about 15 percent [of bot-written text] go by unflagged. We would rather turn down our accuracy than increase our false-positive rate.” EdSurge reports her repeating the roughly one percent figure from the company’s own testing, as the reporter’s summary rather than as a quoted sentence of hers.
In the same article, Grammarly’s head of education, Jenny Maxwell, made an argument about scale that gets quoted as though it were a rebuttal. It is not one, and her own wording says so: even if a detection system is right 98 percent of the time, “that means it falsely flags, say, 2 percent of papers”, and at a university receiving 50,000 papers a year that would be around 1,000 wrongly called cases of cheating. The two percent is an illustrative round number rather than a measurement of anybody’s product. The point that survives is the arithmetic one: at institutional volume, a small error rate is still a queue of real people.
Two more lines from Turnitin’s own guidance are the most concrete things a flagged student can put in front of an instructor. Its AI writing detection model guide states that “To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range”, so anything inside that band is the vendor declining to make a claim at all. The same guide then addresses the people reading your report directly: “Please be reminded that an AI Writing score should not be used as the sole basis for adverse actions against a student.” Neither line asks anyone to argue about a percentage. Both come from the company that produced it. Our breakdown of what the Turnitin indicator actually measures covers the classifier side if you need more of it.
Can Teachers Tell You Used Grammarly?
An instructor reading your finished essay has no record of which suggestions you accepted, because the document does not carry one. Turnitin’s own FAQ says its detector is aimed at model-written content rather than at Grammarly’s spelling, grammar and punctuation fixes, so a clean grammar pass is not something that software reports to anybody either. What an instructor actually holds is a number produced by whatever tool the institution runs, plus whatever the assignment itself left behind. Folk tells are weaker still: our lab’s 510-passage study Every Model Has an Accent measured the famous em-dash at roughly chance level as an AI signal: it is one model’s house style, not a fingerprint of machine writing, so an instructor eyeballing punctuation is guessing.
The realistic risk sits elsewhere, and one university says so to its own faculty rather than to students. The University at Buffalo’s teaching-support team publishes a page telling instructors that Grammarly has expanded well beyond grammar checking, walking through its newer generative features and advising that “AI suggestions are useful starting points, not final answers”. That is awareness guidance for staff rather than a rule for students, and it is worth knowing that this is the briefing some instructors have read. The suspicion in the room is usually about the generative half, even when the conversation uses the single word Grammarly for both halves.
There is also a due-process point that an academic integrity professional put better than I can. Tricia Bertram Gallant, who directs the academic integrity office at UC San Diego, told EdSurge that “If a faculty member can use a tool, accuse a student and give them a zero and it’s done, that’s a problem.” That is the standard worth asking for, and it is not a plea for leniency: it asks that a score open a conversation rather than close one, which is also what Turnitin’s own guidance says.
If English is not your first language and correction tools carry a lot of weight in your drafting, there is a separate body of research you should know about before you ever need it, and it has nothing to do with Grammarly specifically. We cover it in how second-language writing gets read by detectors. Before any of that, one cheap experiment settles more than an argument does: check what a detector returns on writing you already know is yours, from a year nobody could dispute. Instructors weighing what a flag can and cannot support will find the tool-by-tool version of this in our guide for teachers.
Is Using Grammarly Cheating?
Policy decides this one rather than software, and the policies do not agree. Five institutional documents were read directly for this page, and the striking thing is not that they disagree about Grammarly. It is that all five split on the same line, correction on one side and generation on the other, and that four of them hand the actual decision to the individual instructor rather than settling it centrally.
| Institution | What its own document says | Who decides |
|---|---|---|
| University of North Alabama | Two-option syllabus template; Option 1 permits Grammarly for grammar checks and brainstorming, Option 2 prohibits generating work with it | The instructor picks the option |
| Penn State | Names Grammarly among content-generating AI tools covered by the general policy; no institution-wide ban stated | Per class, per assignment |
| University at Buffalo | Faculty-facing guidance on Grammarly’s newer generative features; not a student-facing rule | Advisory to instructors |
| NYU Steinhardt | Model syllabus language permitting Grammarly for brainstorming and refining, with disclosure | Offered to faculty as a template |
| Notre Dame | Where an instructor prohibits generative AI, that prohibition includes editing tools unless stated otherwise | The instructor, by opting in |
The University of North Alabama’s generative AI policy is a template with two options an instructor chooses between. Option 1 says outright that students “may use AI tools (e.g., ChatGPT, Grammarly) to assist with brainstorming ideas, conducting research, grammar checks, or improving the clarity of writing”. Option 2 says that “Using AI tools (e.g., ChatGPT, Jasper, Grammarly AI) to generate essays, assignments, reports, code, or any other form of academic work is not allowed”. Same university, same document, opposite answers, and the variable is the person teaching your section.
Penn State’s academic integrity page on AI names Grammarly explicitly: “This applies to the use of all content-generating AI tools, including Grammarly, Copilot, and other artificially intelligent tools provided by the University.” Being named there is not a prohibition, though, because the same page states the operating rule as “In some classes, students may use AI tools. In some classes, students may not use AI tools.” Nothing on that page bans the tool campus-wide.
The University at Buffalo’s contribution is the faculty briefing described in the previous section, and it sits in a different genre from the other four: it tells instructors what Grammarly’s newer features now do and asks them to review AI-assisted suggestions rather than trust them. Read it as evidence of what your instructor may have been told, not as a rule you can be measured against.
NYU Steinhardt’s syllabus support materials include a general AI template that permits the tool with a condition attached: “AI tools, such as ChatGPT or Grammarly, may be used for brainstorming or refining ideas, but all work submitted must be your own. If AI assistance is used in drafting, please include a note explaining how the tool was used.” Note what that is, exactly. It is language the school offers its own faculty, not a guarantee that every course there works this way, and its condition is disclosure rather than abstinence.
Notre Dame’s honor code update from Fall 2024 is the strictest of the five and still not a blanket rule: “This means that if your instructor prohibits the use of gen AI on an assignment or in a class, this prohibition includes the use of editing tools, unless explicitly stated otherwise.” The tool is not banned there; a prohibition, where one exists, is read as covering editing tools by default. Inside Higher Ed reported that change in November 2024.
So the answer to the heading is conditional, and the condition is not a technical one. Using Grammarly becomes an integrity violation exactly when the assignment in front of you says it does, and the only reliable place to find that out is the syllabus you were given. If the syllabus is silent, ask before you submit rather than after, and ask in writing. A one-line reply from your instructor saying grammar checks are fine is worth more than any page on the internet, this one included, and it comes with a date on it.
Which brings us to the sentence everybody in this argument is stuck with. A student who used a grammar checker and a student who ran a chatbot both end up typing the same words: I only used Grammarly. The sentence is identical in both mouths, it costs nothing to say, and everyone reading it knows that. That is why saying it truthfully so often lands badly, and it is the real reason honest people lose these conversations. What separates the two is never phrasing. It is timing. Evidence that existed before the accusation is a different category of thing from an explanation produced after it: a provenance record switched on before you started, a version history nobody has touched since, the sources you were reading with dates attached, an earlier draft sitting in an inbox because you sent it to a friend. None of that can be assembled retroactively, which is precisely what makes it worth anything.
The corollary is the instruction people most want to ignore. Do not edit the flagged file to make a number move, and do not run it through anything. Whatever else that accomplishes, it converts the one document you own into an edited one and forfeits the only advantage an honest writer has. If a flag has already landed and you need the full sequence, our step-by-step guide to documenting and appealing a false flag picks up from there.
How this page was researched
| Protocol | HumanizeMy Evidence Protocol v1.0 (how it works) |
|---|---|
| Included sources | 22 |
| What they are | Six Grammarly documents (the AI Detector user guide, the support answer on triggering detectors, the Authorship product page, the About Authorship article, the launch announcement, the AI Detector marketing page), three Turnitin documents (the AI writing FAQ, the 2023 false-positives blog, the detection model guide), seven passages across five pieces of journalism (Tech Times, two Fox 5 Atlanta reports, three separate quotes from one EdSurge article, Inside Higher Ed), five institutional policy pages, and one anonymous vendor-adjacent blog carried for characterisation only |
| Capture date | 10 August 2026 |
Screening flow: identified 117, screened 117, excluded 95, included 22.
Exclusions by recorded reason: off-topic 47, duplicate 30, affiliate content with no disclosed method 14, unverifiable on the day 4, coordinated or paid promotion 0. Read that last zero as an absence of signal and nothing more: vendor support documents and university policy pages do not attract the kind of seeding this protocol was built to catch, so there was little there to find.
Limitations, stated rather than buried. Coding was a self-check. Every source was coded once when collected and coded again from its captured quote at the end of the run by the same reader, because no second independent coder was available, so the agreement figure describes internal consistency and not two people agreeing. The count of 117 items was reconstructed in good faith across roughly 34 searches and 30 fetches, counting each address once at first appearance. Three Turnitin pages refused automated requests and were read through a text-rendering proxy pointed at the original addresses, with the raw output searched by hand for the exact wording rather than summarised. The New York Post’s coverage of the Stevens case could not be retrieved at all across four attempts, so it is cited nowhere above and its details appear only where another outlet independently carried them. Two real, identifiable peer-reviewed papers that other pages on this topic cite sit behind a bot check and an institutional login respectively; neither could be read this run, so no figure from either is repeated here and no claim is made about what they contain. Several Grammarly and institutional pages carry no publication date of their own and are dated by the day they were read.
Reviewed every month. When Grammarly changes what its features do, when Turnitin revises its guidance, or when one of the schools quoted above rewrites its policy, the change lands here and the date at the top moves with it. Fırat Mıhcı works in computational linguistics and natural-language processing; the published detection and second-language work is on ResearchGate. Every source linked above was read on 10 August 2026.