Language Naturalness — Weights of the 7 Sub-Metrics Cliché density Concrete vs. abstract Paragraph-opener repetition Syntactic variety Local lexical recycling Connective distribution First-person moderation 18% 18% 15% 15% 12% 12% 10% ~80 phrase-bank hits data / cases vs. hedge words "First / Second / Furthermore" CV + punctuation entropy + long-sentence ratio 100-word sliding-window TTR Shannon entropy parabolic — the middle path wins Weighted average of 7 metrics — real expert writing scores 90+, SEO filler scores below 30.

Why is “Language Naturalness” its own dimension?

Over the past two years, LLMs have mass-produced vast amounts of “written-purely-to-rank” filler content — heavily homogenized, dense with clichés, and lacking specificity. AI providers quickly pushed back: when training their reranking models, they added a naturalness judgment, demoting content that looks like LLM templating.

The blunt conclusion: even if your content is on-topic and structurally correct, once a templated tone gets detected, your odds of being cited drop sharply.

GeoWeb added the “Language Naturalness” dimension in M3-14 (6% weight), using 7 sub-metrics to model the judgment logic of AI providers. Each one is explained below.

Sub-metric 1: Cliché density (18% weight)

What it detects

A predefined “LLM cliché phrase bank” containing roughly 80 common phrases across Chinese and English:

It calculates the number of hits per thousand words.

Scoring logic

Why it carries the highest weight

This is the single strongest signal — someone writing naturally won’t lean on these clichés in every paragraph. A high hit rate = strongly suspected generated content.

Sub-metric 2: Paragraph-opener repetition (15%)

What it detects

Whether consecutive paragraphs open with the same transition word: “First… Second… Furthermore… Finally…” / “Additionally… Moreover… Furthermore… Finally…”

Why it matters

AI-generated content loves this kind of “textbook rigid structure.” When a real expert writes, they mix openers — questions, cases, quotations, definitions, and more. That variety is a signature of a real human.

Scoring standard

The share of the most common opening word:

Sub-metric 3: Syntactic variety (15%)

What it detects

A composite of three independent sub-metrics:

Why it matters

The sentence-length distribution of LLM-generated content often shows “concentration” — most sentences fall within a narrow band (typically 25–35 characters). Human writing mixes extremely short sentences (used for emphasis) with extremely long ones (compound arguments).

Sub-metric 4: Local lexical recycling (12%)

What it detects

Using a 100-word sliding window, it computes the “unique words / total words” ratio (type-token ratio) within each window.

Why it matters

When generating, LLMs tend to repeatedly reach for synonyms within a small range — “provide / boost / improve / optimize / strengthen” might each appear once within the same paragraph. Humans do this far less.

Why a sliding window instead of whole-document statistics

The whole-document TTR is naturally low in long articles (because the total vocabulary gets recycled). The local TTR of a sliding window more precisely detects “in-paragraph repetition,” which is the LLM signature.

Sub-metric 5: Concrete vs. abstract (18% weight, tied for highest with cliché density)

What it detects

It calculates the ratio of “concrete marker density / abstract marker density.”

Why it carries high weight

This is one of the strongest signals for distinguishing “real research / case sharing” from “AI filler”:

Lots of hedge words = ungrounded paraphrasing = strongly suspected generated content.

Sub-metric 6: Connective distribution (12%)

What it detects

The Shannon entropy of the distribution of a predefined list of connectives (“然而” (however), “因此” (therefore), “不過” (yet), “另外” (additionally), “此外” (moreover), “再者” (furthermore)…).

Why it matters

LLMs like to reuse 1–2 connectives over and over (the most common: “however,” “therefore,” “additionally”). Humans naturally rotate through various synonymous alternatives, so the distribution is more even = higher entropy.

Sub-metric 7: First-person moderation (10%)

What it detects

The density of first-person experience markers such as “我們” (we), “In our experience,” “我發現” (I found).

Scoring logic (parabolic, inverted-U curve)

More is not better. The middle path is best.

The combined reading of the 7 sub-metrics

GeoWeb’s “Language Naturalness” score is the weighted average of these 7 metrics. Common combinations:

Why does this dimension carry “only” 6% weight?

LLMs’ anti-slop detection is still evolving — at present, the major AI platforms weight structured readiness more heavily than language naturalness. But this weight will rise over time.

Our advice: content that is genuinely written by a real expert matters more than “deliberately gaming the metrics to look good.” Using prompt engineering to make an LLM article look more human may pass detection in the short term — but the AI providers’ detection models keep updating too. In the long run, only real human writing + real cases + concrete data is safe.

See your language naturalness with a health check

👉 Free GEO Health Check — the “Language Naturalness” dimension lists out all 7 sub-metric scores one by one, and points out which ones are dragging down the overall result.

If your website content needs bulk adjustment to meet naturalness standards (including content-rewrite strategy, writer guidelines, and an automated detection pipeline), we offer a GEO consulting service: [email protected]


GEO Deep-Dive Series #15. Previous article: “The 12 Dimensions, Broken Down: The 5 Checkpoints of FAQ/Q&A Readiness”