top of page

LLM Cliché Highlighter AI Writing Cliché Detector Challenges the AI Detection Industry

Updated: Jul 20

Simon Willison released the LLM cliché highlighter AI writing cliché detector after another article overwhelmed him with familiar machine-like phrasing. Instead of assigning an AI probability, his browser tool marks recurring patterns such as “no fluff, no filler, no jargon.” The conflict is immediate. Most AI detectors promise an authorship verdict, while Willison’s tool shows readers the specific language that made a passage feel synthetic.

The distinction matters because conventional detection remains unreliable. OpenAI withdrew its own text classifier in 2023 after reporting weak accuracy, including a 9 percent false-positive rate in one evaluation. Willison takes a narrower route. His tool identifies editorial symptoms without pretending those symptoms prove who, or what, wrote the text.

That makes the project less like an academic integrity system and more like a spellchecker for exhausted prose. It also arrives after Wikipedia editors spent years cataloging recurring signs of AI-assisted writing. The resulting question is not whether software can expose an invisible machine author. It is whether visible clichés provide enough evidence to make writing better.

What the LLM Cliché Highlighter Actually Changed

Willison turned a subjective reaction to AI-flavored prose into an inspectable editing interface.

Willison published the tool on July 17, 2026, through his personal tools site. According to his launch post, frustration supplied the brief. He had encountered another article packed with phrases resembling “no fluff, no filler, no jargon,” then asked Fable 5 to build an application around ten common patterns.

The resulting LLM cliché highlighter accepts pasted text or content loaded from a URL. Analysis runs as the user types. Matching language receives a highlight, while the surrounding sentence gets a separate visual marker that preserves context.

Users can enable or disable individual patterns. They can also review totals, inspect a list of matches, and jump directly to each occurrence. Chains such as “no X, no Y” receive an item count, which helps distinguish a single construction from an extended rhetorical sequence.

The application stores text in the browser through localStorage, the browser feature that preserves data between sessions. Willison says the analysis itself runs entirely on the user’s device. That design reduces the friction of checking drafts while avoiding a required upload to a remote text-analysis service.

The interface can also load a public URL through a text extraction service. That makes it useful for examining published articles, landing pages, newsletters, and documentation. Its primary interaction, however, remains simple: supply text and inspect the highlighted language.

There is a small numerical wrinkle. Willison’s post describes ten common patterns, while the captured interface reports that all 11 patterns are active. The difference likely reflects the way the application groups or counts its rules, but the post does not explain it.

That mismatch does not alter the product’s central purpose. It flags known constructions rather than calculating whether an LLM generated an entire document. A match tells the reader that a selected pattern appears. It does not establish provenance.

This modest claim is the most important feature. The tool does not label a writer dishonest, accuse a student of misconduct, or produce a supposedly scientific percentage. It points to a phrase and gives an editor something concrete to evaluate.

Consider “no fluff, no filler, no jargon.” The construction uses repetition to promise directness, but it often consumes space without adding evidence. The highlighter can surface that weakness whether a person, a model, or a person using a model produced it.

Other patterns shown by Willison include “sit with that,” “you already know,” and constructions built around something being “worth naming.” These phrases are not inherently wrong. Their repeated use can still create a recognizable voice assembled from familiar online rhetoric.

The application therefore changes the unit of analysis. Commercial detectors tend to evaluate the document and return a verdict. Willison evaluates local language and returns editing candidates.

That difference establishes the article’s central tension. The LLM cliché highlighter offers less certainty than an authorship detector, but its limited output is easier to inspect, challenge, and use.

Why AI Writing Detection Needs a Narrower Claim

The tool matters because an observable cliché is not the same thing as proof of AI authorship.

AI text detection has always faced an uncomfortable classification problem. Human and machine writing overlap because language models learn from human text. People also revise machine output, while models can imitate requested tones, formats, and individual styles.

A detector still has to separate those intertwined categories. Its score can appear exact even when the underlying distinction is unstable. Short passages, edited drafts, unfamiliar subject matter, and non-native English can all complicate the result.

OpenAI’s discontinued classifier provides a useful historical reference. The company reported that it identified 26 percent of AI-written samples as likely AI-generated in its English challenge set. It incorrectly classified 9 percent of human text as AI-written.

OpenAI removed the classifier in July 2023 because of its low accuracy. Its archived classifier notice also warns that short text was especially unreliable and edited AI text could evade detection. The company said the system should not serve as a primary decision-making tool.

Those limitations become serious when a score influences grades, employment, publication, or reputation. A false positive does more than produce a bad recommendation. It assigns suspicion to a writer who may struggle to disprove an opaque result.

The LLM cliché highlighter avoids that claim entirely. It does not infer a hidden origin from statistical features. It finds explicit strings or constructions selected by its rules.

That makes its output reproducible in a basic sense. Two readers can see the highlighted phrase, review the sentence, and debate whether the language is tired. They do not have to accept an inaccessible model’s probability as the final answer.

The tool also separates two questions that detector marketing often blends together:

  • Does this text contain habits frequently associated with LLM output?

  • Was this text written by an LLM?

The first question concerns style. The second concerns authorship. A piece can satisfy either condition without satisfying the other.

A human marketing writer might use “no fluff, no filler” because the phrase is common in sales copy. An LLM can avoid it when prompted to use concrete detail. A person can also accept a model’s structural suggestions while rewriting every sentence.

Research on stylometry, the analysis of writing style through measurable features, reinforces this caution. A Computational Linguistics study on the limits of stylometry found that stylistic detection did not separate legitimate language-model applications from machine-generated misinformation. Similar surface features appeared regardless of whether the underlying content was deceptive.

That study addressed a different generation of models and a narrower misinformation problem. Its conceptual warning still applies. Style can help identify patterns, but it cannot reliably reveal intent, truth, or the exact production process.

Willison’s tool accepts that boundary. It treats formulaic prose as an editorial issue instead of converting it into an accusation.

The approach is especially relevant to working editors. Their immediate task is usually not forensic attribution. It is deciding whether a sentence earns its place, sounds specific, and carries the author’s intended meaning.

A phrase-level marker supports that task. An editor can remove the cliché, replace it with evidence, or keep it when the rhythm fits. The output remains advice rather than judgment.

This narrower claim pressures conventional AI detectors in an unexpected way. Willison is not competing on benchmark accuracy. He is questioning whether authorship classification is the right product for ordinary editing.

The Real Opponent Is Verdict-Based AI Detection

The primary contest is between inspectable editorial signals and opaque authorship verdicts.

A verdict-based detector promises a clear answer to an emotionally charged question. Paste the document, wait for analysis, and receive a percentage or label. That simplicity is attractive to teachers, editors, recruiters, and publishers facing more generated text than they can manually review.

The output also hides difficult choices. What threshold counts as AI-written? How was the system calibrated? Which models produced its training data? How does it handle translation, grammar correction, or mixed human-machine drafting?

New models and writing workflows constantly shift the target. A detector trained on yesterday’s unedited chatbot output can struggle with tomorrow’s models. It can also misread text after a person restructures, shortens, or personalizes it.

The LLM cliché highlighter reverses the product logic. It gives the user more interpretation work, not less. Instead of one score, it offers multiple passages that require judgment.

That sounds like a limitation because it is one. It is also the source of the tool’s credibility.

A highlighted “you already know” does not prove anything about authorship. It does reveal an unsupported assumption about the reader. An editor can ask whether the audience truly knows the referenced fact or whether the sentence merely simulates intimacy.

Likewise, “sit with that” can be effective after a difficult or surprising fact. Repeated without such a fact, it asks the reader to supply emotional weight that the argument has not earned. The issue is not machine involvement. The issue is rhetorical substitution.

“No X, no Y” chains present another recognizable mechanism. Parallel negatives create a punchy cadence and advertise what the text supposedly avoids. When several appear together, the copy can spend more effort declaring its directness than being direct.

A conventional detector converts many signals into a conclusion that the user cannot fully inspect. Willison exposes a small set of signals and declines to draw the conclusion. This makes the application weaker as enforcement software but stronger as an editing aid.

Wikipedia’s community-developed guidance occupies similar territory. Editors built a public catalog of recurring language, formatting, citation, and interaction patterns associated with LLM contributions. The guide cautions that its signs are clues rather than proof.

Coverage of the Wikipedia field guide noted that single giveaway words provide thin evidence. Patterns become more useful when reviewers consider them alongside factual problems, citation behavior, and the surrounding editing context.

That distinction matters for professional content teams. Replacing one cliché with another does not repair an article built on unsupported claims. Removing AI-flavored phrasing can even make weak material harder to recognize without improving its substance.

The best use of a cliché highlighter is therefore diagnostic. A cluster of matches tells an editor where to slow down. The editor should then examine the paragraph’s evidence, specificity, and connection to the reader’s actual needs.

This process can expose a deeper problem in AI-assisted publishing. Language models readily produce transitions, summaries, balanced lists, and motivational conclusions. Those strengths help complete a draft, but they can also smooth over missing reporting.

A polished paragraph may contain no original observation. Its sentences can move cleanly while its argument remains stationary. Formulaic expressions often appear where the writer needs a concrete example, a named source, or a defensible claim.

The highlighter cannot identify every such gap. It can act as a visible tripwire around several common ones.

For knowledge workers, this creates a practical review sequence. First, find the repeated rhetorical templates. Second, ask what each template is doing. Third, replace empty emphasis with a fact, example, limitation, or decision.

That workflow fits broader personal knowledge practices. Writers who keep searchable source material can replace generic model language with details drawn from their own work. A structured AI second brain helps preserve those details before the drafting stage begins.

Verdict-based detection offers institutional convenience. Phrase-level review offers individual accountability. Willison’s small tool makes a strong case that accountability produces better prose than an unexplained percentage.

What This AI Writing Cliché Detector Cannot Prove

The LLM cliché highlighter can identify repetition, but it cannot identify an author or measure truth.

Every flagged pattern has human precedents. Language models did not invent parallel negatives, direct address, dramatic pauses, or therapeutic phrasing. They absorbed those devices from writing created by people.

Calling them “LLM clichés” describes their frequency and cultural association, not their origin. A model’s preferred constructions often come from genres already saturated with formulas, including marketing pages, self-help posts, consulting material, and search-oriented articles.

That creates a false-positive risk at the level of interpretation. The application can accurately match a phrase while a reader draws an unjustified conclusion about the writer. Technical accuracy does not guarantee responsible use.

A teacher should not treat a highlighted phrase as evidence of academic misconduct. An editor should not reject a submission because it includes several listed patterns. An employer should not use the interface to evaluate a candidate’s honesty.

The tool’s name makes its subject clear, but its design communicates the safer boundary. It reports matches and flagged sentences. It does not display “human” or “AI” labels.

Users must preserve that restraint. A cliché is a reason to inspect a sentence, not a reason to investigate a person.

Pattern lists also age. Once writers and model developers recognize a disliked construction, prompts can suppress it. A user can request fewer parallel negatives, no therapeutic framing, and more concrete language. The visible signal disappears even if the model still produced most of the draft.

This adaptation creates an arms race for traditional detectors. It creates a maintenance problem for the highlighter too, although the consequences are smaller. An outdated rule set misses newer habits or keeps flagging language that no longer carries the same association.

The current patterns also reflect an English-language reading environment. Rhetorical habits differ across languages, dialects, professional communities, and cultural contexts. A useful pattern in North American technology writing might be ordinary elsewhere.

There is another fairness issue. People who learned English through formal instruction may rely on reusable transitions and balanced sentence structures. Professional templates encourage the same behavior. Flagging those constructions as machine-like can reinforce a narrow idea of authentic voice.

Research on AI-supported writing among Black users has documented tension between assistance and identity. Participants in one 2025 study described benefits from writing tools alongside concern that the resulting language did not represent them. That broader problem cannot be reduced to a list of clichés.

The highlighter may still help users notice when an assistant has flattened their voice. It should not define what a human voice must sound like.

The application also cannot evaluate factual accuracy. A sentence without any listed cliché can contain fabricated statistics, nonexistent sources, or distorted context. A heavily flagged sentence can be completely accurate.

This is where style review must reconnect with source review. Editors need to verify names, dates, quotations, links, and causal claims separately. Removing “it is worth noting” does not validate whatever follows.

Nor can the tool determine whether AI assistance was appropriate. A writer might use a model for brainstorming, grammar correction, translation, or outlining. Another might generate an entire article and conduct no verification. The finished text alone may not reveal that difference.

A more mature publishing policy would focus on responsibility. Who checked the facts? Who owns the argument? What assistance must be disclosed? Can the author explain and defend the finished work?

Cliché highlighting can support such a policy because it encourages close reading. It cannot replace disclosure rules, editorial review, or provenance systems.

The label “detector” therefore requires care. This is an AI writing cliché detector, not a dependable AI authorship detector. Confusing those categories would reproduce the exact overconfidence the tool’s restrained design avoids.

Its limitations are not hidden defects. They define the product’s proper use.

Why Cliché Detection Is Really a Writing Quality Test

The most valuable outcome is not catching AI; it is forcing vague prose to become specific.

Many LLM clichés perform work that evidence should perform. They announce significance, intimacy, clarity, or emotional weight without establishing those qualities. Highlighting them creates a chance to ask what the sentence has not yet earned.

Take a phrase such as “that loss is real, and it is worth naming.” The statement may respond compassionately to a genuine experience. In another context, it can imitate empathy while avoiding the details of who lost what and why.

An editor can improve the second version by naming the consequence. A team missed a deadline. A customer lost access to stored work. A policy removed an expected safeguard. Concrete language gives the reader something to evaluate.

The same principle applies to “you already know.” The phrase claims shared knowledge, which can make prose feel conversational. It can also exclude readers or conceal the absence of explanation.

A revision might state the relevant fact in one sentence. If the fact is unnecessary, the writer can remove the setup entirely. Either decision serves the reader better than simulated familiarity.

“No fluff, no filler, no jargon” invites a more direct repair. Delete the promise and present the useful information. If the following paragraph still feels vague, the slogan was covering a structural weakness.

This is where the LLM cliché highlighter can fit inside an editorial workflow. It works best after the writer has assembled a complete draft and before final copyediting. The goal is not to erase every match automatically.

First, review clusters. Several matches in a short section suggest the prose relies heavily on prefabricated rhythm. One isolated match may be intentional and effective.

Second, examine function. Ask whether the phrase adds meaning, controls pacing, or helps the reader understand the claim. Keep it when it does.

Third, search the writer’s source material for a replacement. Meeting notes, interviews, research papers, product logs, and customer feedback contain details that generic language cannot supply. A searchable knowledge workflow can shorten that step.

Fourth, read the revision aloud. Removing familiar constructions can leave stiff or fragmented prose. The objective is not maximal originality in every sentence. It is a credible voice with enough variation and specificity to sustain attention.

This editing process also reveals why automatic substitution would be a mistake. A tool that replaces every flagged phrase could simply create a new collection of formulas. The next model would learn those alternatives, and readers would tire of them too.

Language changes through repetition. Expressions become clichés because they once worked well enough to imitate. AI accelerates the cycle by producing popular structures at enormous frequency, but humans remain part of that feedback loop.

The better response is not a permanent blacklist. It is active judgment about whether a phrase still carries meaning in its current context.

Content teams can use aggregate matches as a quality signal, but they should resist turning the count into a target. Writers will optimize against any metric. A low match count can coexist with an empty article, just as a high count can appear in an effective personal essay.

The most useful team-level question is qualitative: Which recurring constructions correlate with weak work in this publication? Editors can maintain local guidance based on their audience, format, and subject matter.

A technical documentation team might flag unsupported ease claims such as “simply” or “just.” A newsroom might focus on vague attribution. A research organization might prioritize inflated significance statements and causal leaps.

Willison’s project demonstrates the interface for that approach. Toggleable patterns let users decide which rules matter. Context-aware sentence highlighting prevents the matched phrase from becoming detached from its role.

The browser-only design also supports private drafts. Writers can examine unpublished text without placing it in another hosted AI detection service. However, URL loading involves an external extraction route, so sensitive material belongs in the paste workflow rather than a public URL.

This small distinction reflects the tool’s broader appeal. It does not attempt to become an all-purpose content authority. It offers a focused lens that users can combine with fact-checking, source management, and human editing.

Used well, the LLM cliché highlighter makes writing more accountable. The writer must decide why a phrase exists and what should replace it. That is slower than accepting a detector score, but it produces a document someone has actually reviewed.

Three Signals Will Show Whether the LLM Cliché Highlighter Matters

The next test is whether this experiment becomes an adaptable editing practice or remains a clever snapshot of 2026 prose.

The first signal is pattern maintenance. Willison’s launch post describes ten patterns, while the visible interface shows 11 active rules. Future additions, removals, or refinements would indicate that the list is becoming a maintained vocabulary of recurring LLM habits.

That development would strengthen the tool’s editorial value. A stagnant list would still capture familiar constructions, but it would become less representative as models and prompting practices change.

The important measure is not the total number of rules. A huge catalog would overwhelm writers and increase incidental matches. Useful maintenance should favor recognizable patterns with clear editing implications.

The second signal is adoption by editors, educators, and content teams as a revision aid. Responsible adoption would emphasize discussion of highlighted passages. Punitive adoption would treat matches as evidence of prohibited AI use.

The first outcome would reinforce Willison’s narrow approach. The second would weaken it by converting a transparent style checker into an improvised surveillance tool.

Watch how users describe their results. “This helped me find repetitive language” matches the application’s actual capability. “This proved the author used AI” goes beyond the available evidence.

The third signal is imitation inside mainstream writing products. Grammar checkers, document editors, and publishing systems already identify passive voice, repeated words, tone, and readability issues. They can readily add customizable AI-style pattern checks.

If those products present the feature as editorial guidance, Willison’s framing has traveled. If they package the same signals as an authorship score, the detection industry has absorbed the interface without learning its central lesson.

Model providers will respond indirectly as well. Prompts and system instructions increasingly discourage stock phrases, excessive headings, canned transitions, and exaggerated emphasis. Better defaults can reduce the most obvious symptoms.

That would not make cliché highlighting obsolete. Human writing tools already warn about habits that competent writers can avoid. The value lies in catching them when attention is limited.

The larger uncertainty is whether recognizable AI style remains a stable category. As people consume generated text, they adopt some of its vocabulary and rhythms. Models then train on more text shaped by earlier models. Human and machine habits become harder to separate.

That feedback loop makes authorship detection less dependable. It also makes editing standards more important. Publications need to define the writing they want, not merely the patterns they suspect.

The LLM cliché highlighter points toward that future. It replaces the fantasy of certain detection with a visible conversation about language. Its success should be measured by stronger revisions, not by how many writers it appears to catch.

For readers, the practical action is straightforward. Paste a draft, inspect each match, and ask whether the highlighted phrase contributes evidence, clarity, or rhythm. Keep it when the answer is specific. Rewrite it when the phrase only performs confidence or emotion. Then verify the claims independently, because clean style does not guarantee accurate reporting. The LLM cliché highlighter AI writing cliché detector is most useful when it starts an editorial decision, not when it ends one.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page