Guide
How to tell if something is written by AI
AI writing tools have become so capable that distinguishing human-written text from AI-generated text is genuinely difficult. ChatGPT, Claude, Gemini, and dozens of other models can produce fluent, grammatically correct prose on virtually any topic. But AI-generated text is not identical to human writing. It carries statistical patterns, stylistic tendencies, and structural characteristics that reveal its machine origins. This guide covers both the manual cues you can look for yourself and the forensic analysis methods that provide deeper, more reliable detection.
Why AI-written text is everywhere
The explosion of AI content in 2026
The scale of AI text production is staggering. Millions of people use AI writing assistants daily for emails, reports, articles, social media posts, academic papers, marketing copy, and creative writing. Estimates suggest that AI-assisted or fully AI-generated text accounts for a significant and growing share of content published online. Some industry analyses place the figure at 10-15% of all new web content in 2026, up from negligible levels just three years earlier.
Where AI writing appears (and where you might not expect it)
AI-generated text appears in contexts you would expect: blog posts, product descriptions, social media content, and student assignments. But it also appears in less obvious places: customer service responses, legal document drafts, medical chart notes, news summaries, political communications, dating profile bios, and professional correspondence. The ubiquity of AI writing tools means that virtually any text you read could potentially include AI-generated content, whether fully generated or used as a starting draft that was later edited.
Why it matters to identify AI-generated text
Identifying AI-generated text matters for different reasons in different contexts. In academic settings, it is an integrity question: is the student demonstrating their own learning? In journalism, it is a trust question: was this article researched and written by someone with expertise and accountability? In legal proceedings, it is an evidentiary question: who actually authored this document? In content publishing, it is a quality question: does this content reflect genuine expertise and unique perspective? The answer to whether AI detection matters depends entirely on context, but the need for detection capability spans many important domains.
Manual detection - what to look for
Overly smooth and consistent writing style
Human writing has texture. Individual writers have strong sentences and weak sentences, clear sections and muddled sections, areas of confidence and areas of uncertainty. AI-generated text tends toward a consistent level of polish throughout, with uniform sentence quality, even paragraph density, and a smoothness that lacks the natural variation of human composition. If a piece of writing feels like every paragraph was written with exactly the same skill level and attention, that uniformity itself may be a signal.
Hedging language and lack of strong opinions
AI models are trained to be helpful and balanced, which produces a characteristic hedging style. Look for phrases like "it is worth noting that," "there are many perspectives on," "this is a complex topic," "while some argue," and similar constructions that qualify or soften claims without adding substance. Human writers with genuine expertise tend to make definitive statements in their area of knowledge. AI text tends to maintain a cautious, "on the other hand" approach even on topics where a knowledgeable human would be direct.
List dependency and formulaic structure
AI-generated content frequently relies on numbered lists, bullet points, and formulaic paragraph structures (claim-evidence-transition, repeated three times). While human writers certainly use lists, AI text tends toward a mechanical pattern where every topic is addressed through the same structural template. Look for articles where every section follows an identical pattern: topic sentence, three supporting points, transition to next section. Human writing is more structurally varied, with some ideas developed in depth and others mentioned briefly.
Generic examples and lack of specific detail
One of the most telling characteristics of AI writing is the use of generic rather than specific examples. An AI writing about restaurant management might mention "challenges like staffing and supply costs." A human restaurant manager would mention "the Tuesday night when three servers called in sick during restaurant week and we had a 90-minute wait." AI-generated text tends to remain at a level of abstraction that sounds informed but lacks the concrete, specific, sometimes messy details that come from actual experience.
Unusual vocabulary patterns
AI models favor certain words and phrases with higher probability than human writers typically do. In English-language AI text, watch for overuse of words like "delve," "landscape," "robust," "leveraging," "streamline," "nuanced," "multifaceted," "pivotal," and "foster." No single word is diagnostic (humans use all these words too), but a high concentration of these AI-favored terms in a single piece of text is a statistical signal. The specific overused vocabulary shifts across models and over time, but the tendency toward particular high-probability words is consistent.
Perfect grammar with no natural errors
Human writing contains natural imperfections: occasional grammar deviations, informal constructions, sentence fragments used for emphasis, and style choices that a grammar checker would flag. AI-generated text tends toward grammatical perfection, following all standard rules consistently. Text that is 100% grammatically perfect across a long document is actually statistically unusual for human writing, particularly in informal or semi-formal contexts like blog posts, emails, and social media.
Important caveat: None of these manual signals are conclusive on their own. Skilled human writers can produce smooth, hedging, list-heavy text. Professional editors can remove all natural errors. Manual detection provides useful initial signals, but reliable determination requires forensic analysis that examines statistical properties inaccessible to casual reading.
Statistical markers of AI writing
Perplexity - measuring predictability
Perplexity is the most widely discussed statistical marker for AI detection. It measures how predictable a sequence of words is according to a language model. AI-generated text tends toward low perplexity because the generation process selects high-probability tokens at each position. Human text shows higher and more variable perplexity because humans make word choices based on personal style, emphasis, and knowledge that often diverges from the statistically most likely option.
The practical interpretation: if you ran the text through a language model and it found virtually every word highly predictable, the text is more likely AI-generated. If the model found many unexpected or unusual word choices, the text is more likely human-written. Perplexity analysis is most reliable for longer texts (500+ words) and becomes increasingly uncertain for shorter samples.
Burstiness - sentence length variation (or lack thereof)
Burstiness measures the variation in sentence length throughout a text. Human writing shows high burstiness: a mix of very short sentences (sometimes just a few words for emphasis) and long, complex sentences. The variation follows the writer's rhetorical intent and the rhythm of their thinking. AI-generated text typically shows lower burstiness, with sentences clustering around a narrower length range. This reflects the model's optimization for fluent, consistent output rather than the deliberate rhythm variation of human composition.
Entropy analysis and information density
Information entropy measures how much new information each word adds to the text. Human writing tends to vary in information density, with some passages dense with facts and specifics and others providing context, transition, or rhetorical emphasis. AI-generated text often maintains a more consistent information density, distributing content evenly across paragraphs rather than concentrating it where the writer's knowledge is deepest. This evenness produces an entropy profile that differs from the natural peaks and valleys of human-authored text.
N-gram frequency distributions
N-gram analysis examines the frequency of specific word sequences (pairs, triples, and longer phrases) and compares them to expected distributions for human and AI text. AI models tend to produce certain word combinations more frequently than human writers do, and they avoid certain combinations that humans use naturally. These distributional differences are subtle at the individual n-gram level but become statistically significant when analyzed across the full text.
Type-token ratio and vocabulary richness
The type-token ratio (unique words divided by total words) measures vocabulary diversity. AI-generated text tends to show specific type-token ratio patterns that reflect the model's vocabulary selection process. For most topics, AI text falls within a narrower type-token ratio range than human writing, which varies more depending on the individual writer's vocabulary, the formality of the context, and the specificity of the subject matter. Vocabulary richness analysis can also detect humanized AI text, which often overshoots in vocabulary diversity.
Forensic text analysis methods
Stylometric fingerprinting
Stylometry quantifies writing style through measurable features: function word frequencies (the, of, and, in), punctuation patterns, sentence complexity distributions, paragraph structure, and hundreds of additional characteristics. Every human author produces a distinctive stylometric profile, while AI-generated text produces profiles that are characteristic of the model rather than any individual. Forensic stylometric analysis can evaluate whether a text matches the expected profile of human authorship or shows the statistical signature of model-generated output.
Multi-feature ensemble classifiers
The most effective forensic approach combines many independent features into an ensemble classifier. Rather than relying on perplexity alone (which can be manipulated) or burstiness alone (which varies by genre), ensemble methods combine perplexity, burstiness, entropy, n-gram distributions, type-token ratios, function word frequencies, sentence complexity metrics, and structural features into a single determination. This multi-feature approach is substantially more robust than any single-feature method because it is difficult for an AI model (or a humanizer tool) to simultaneously normalize all features.
Cross-model detection approaches
An important practical challenge is detecting text from models the detector has never been trained on. Cross-model detection methods focus on features that characterize AI generation generally rather than any specific model. Self-supervised approaches, which learn representations from large collections of text without model-specific labels, show promise for generalization to novel models. The goal is a detector that works on text from next year's model without needing to be retrained.
AFIP forensic text analysis
AFIP's text analysis combines multi-feature ensemble classification with evidence-based reporting. The system evaluates text across dozens of independent features at the word, sentence, paragraph, and document levels, synthesizes the results through evidence weighting, and produces a confidence score accompanied by specific forensic findings. The reporting identifies which features contributed most to the determination, providing transparency that supports the result's use in academic, journalistic, or legal contexts.
AI detection tools - how they work and their limits
How major detection tools approach the problem
Most consumer AI detection tools use some combination of perplexity analysis, neural network classification, and feature-based scoring. The tools differ in which features they prioritize, how large their training datasets are, how frequently they are updated with output from new models, and how they communicate results. Some provide a simple "human" or "AI" label, others provide a percentage score, and forensic tools provide detailed evidence reports.
Accuracy rates and false positive risks
| Content type | Typical detection accuracy | False positive rate |
|---|---|---|
| Full AI-generated article (1,000+ words) | 85-95% | 2-5% |
| AI with light human editing | 70-85% | 5-10% |
| AI with heavy human editing | 55-70% | 8-15% |
| Humanizer-processed text | 65-85% (forensic) | 5-8% |
| Short text (<200 words) | 55-70% | 10-20% |
| Non-English text | 60-80% | 10-15% |
The edited-text and mixed-content challenge
One of the most common real-world scenarios is text that was initially drafted by AI and then edited by a human, or text where human-written sections are interspersed with AI-generated sections. This mixed-content scenario is challenging for detection because the human edits introduce genuine human writing characteristics while the AI-generated foundation retains some machine characteristics. Paragraph-level analysis (evaluating each paragraph independently rather than the whole document) can sometimes identify which portions are AI-generated, but accuracy decreases as human editing becomes more extensive.
Short text limitations (under 200 words)
Detection accuracy drops significantly for short text samples. Emails, social media posts, comments, and brief messages often contain too few words for reliable statistical analysis. Most detection methods need at least 200-300 words for reasonable accuracy, and 500+ words for high confidence. For very short text, manual evaluation of the content's specificity, voice, and contextual fit may be more informative than automated detection tools.
Common mistakes in AI detection
Over-relying on a single tool
Different detection tools have different strengths and biases. A text that one tool flags as AI-generated may be classified as human by another. Using a single tool as the definitive arbiter of authorship is methodologically unsound. Best practice involves checking with multiple independent tools and considering the convergence of their results. When tools agree, confidence increases. When they disagree, further investigation is warranted.
Confusing confident writing with AI writing
Some writers naturally produce clear, well-organized, grammatically polished prose. This style can trigger AI detection tools that associate polish with machine generation. Professional writers, experienced academics, and trained journalists may produce text that scores higher on AI detection tools than their less polished peers, not because the text is AI-generated, but because the detection signals overlap with skilled human writing. This is why detection should never rely solely on surface-level style characteristics.
Ignoring context and source credibility
Detection tools provide statistical analysis, not definitive proof. The output should be interpreted in context. Text submitted by a known expert author who has published extensively on the topic carries different prior probability than text from an anonymous source. The cost of a false accusation (wrongly accusing a human author of using AI) can be significant and should be weighed against the detection tool's confidence level.
The false positive problem - accusing humans
False positive rates for AI text detection are not zero. Every detection tool sometimes flags human-written text as AI-generated. The consequences of a false positive can be severe: academic penalties, reputational damage, lost publishing opportunities. Detection results should never be used as the sole basis for punitive action. They should be treated as one data point within a broader evaluation that includes author interviews, draft history, source verification, and contextual assessment.
A practical verification process
Step 1 - Read critically first
Before using any tool, read the text carefully. Does it feel generic or specific? Does the writing voice remain constant throughout, or does it vary as you would expect from a human author working through a complex topic? Are the examples concrete and detailed, or abstract and interchangeable? Your initial impression provides useful context for interpreting tool results later.
Step 2 - Check for statistical markers
Look for the manual signals described above: hedging language, formulaic structure, vocabulary clustering, and stylistic uniformity. Note any sections that feel particularly generic or unusually polished. These observations help you identify which portions of the text may warrant closer analysis.
Step 3 - Run forensic analysis
Submit the text to detection tools, ideally more than one. Compare the confidence scores and findings across tools. If multiple independent tools flag the same text with high confidence, the convergent evidence strengthens the determination. If results are mixed or low-confidence, treat the result as inconclusive and gather additional evidence.
Step 4 - Consider context and evidence holistically
Evaluate the detection results alongside contextual information. Who is the claimed author? Do they have a history of similar writing? Is the writing consistent with their demonstrated knowledge and expertise? Is there a draft history or evidence of the writing process? A 75% AI detection score for an anonymous blog post carries different implications than the same score for a signed article by an established journalist.
Step 5 - Verify with AFIP
For high-stakes determinations where accuracy matters, submit the text to AFIP's forensic text analysis. The multi-feature ensemble approach provides the most robust detection available, and the evidence-based reporting identifies specific forensic signals that support the finding. AFIP's confidence scoring communicates genuine uncertainty honestly, helping you make informed decisions rather than acting on potentially unreliable binary verdicts.
The future of AI text detection
AI text detection is in a period of rapid evolution driven by several converging developments. Foundation model detectors that learn general representations of human versus AI writing are improving cross-model generalization. Watermarking approaches embedded in the generation process (like the methods described by Kirchenbauer et al.) provide a complementary signal for text produced by cooperating platforms. Multi-modal analysis that considers text alongside images, formatting, publishing context, and author metadata enables richer verification than text-alone analysis.
The fundamental challenge will persist: as language models improve, the statistical difference between AI-generated and human-written text narrows. But it does not disappear. Human writing reflects individual experience, idiosyncratic knowledge, emotional state, and communicative intent in ways that statistical generation approximates but does not replicate. Forensic methods that capture these deeper characteristics will continue to provide detection capability even as surface-level distinctions diminish.
Check any text with AFIP forensic analysis
Upload text for multi-method forensic detection with confidence scoring and evidence-based findings.
Analyze text nowFrequently asked questions
Can teachers tell if students use AI to write essays?
Teachers can sometimes spot AI-generated essays through manual reading (generic examples, lack of personal voice, suspiciously consistent quality). Automated detection tools provide additional capability, though accuracy varies. The most reliable approach combines detection tools with process-based assessment: comparing the submitted work with the student's demonstrated in-class writing ability, draft history, and ability to discuss the content in depth. No detection tool is accurate enough to use as the sole basis for academic misconduct charges.
How accurate is AI text detection?
Accuracy depends on the specific tool, the length of text, the generation model, and whether the text was edited or processed through humanizer tools. For full-length, unedited AI text from known models, top forensic tools achieve 85-95% accuracy. For edited, humanized, or short text, accuracy drops to 55-85%. All detection tools produce some false positives (incorrectly flagging human text as AI). Forensic ensemble methods generally outperform single-feature tools, particularly on challenging cases.
Can AI detection tools tell which AI model wrote the text?
Some tools can provide model attribution (estimating whether text came from ChatGPT, Claude, Gemini, etc.) alongside detection. Attribution accuracy is generally lower than detection accuracy because the statistical differences between models are subtler than the differences between AI-generated and human-written text. Attribution works best for longer texts and for models well-represented in the detector's training data.
Does editing AI-generated text make it undetectable?
Editing reduces but does not eliminate detection capability. Light editing (fixing a few phrases, adding personal anecdotes) leaves most AI-generated statistical patterns intact. Heavy editing (substantially rewriting paragraphs, reorganizing structure, injecting personal expertise) reduces AI detection scores significantly but rarely eliminates all forensic signals in longer texts. The more extensively a human rewrites AI-generated content, the more the text genuinely becomes a human-AI collaboration, which is a legitimate use case that many policies distinguish from fully AI-generated submission.
Is it ethical to use AI detection tools?
Using detection tools is ethical when done responsibly. This means treating results as evidence rather than proof, considering false positive risks before taking action, using multiple tools and methods rather than relying on a single verdict, and maintaining transparency about how detection results are used in decision-making. Detection tools become ethically problematic when they are used as sole evidence for punitive action, applied without considering false positive rates, or deployed in ways that create a presumption of guilt.