Why citation rates vary so widely across paragraphs of the same article
Open ChatGPT and ask it about a topic you’ve written on. Watch the answer:
- Your site sometimes gets cited — but usually only one specific paragraph
- Other paragraphs of the same article appear to be completely invisible to the AI
- The paragraph you thought was your best went uncited, while a seemingly plain one got picked up
The difference isn’t random — it’s 5 measurable content features at work.
When an LLM cites content in slices, it runs an “independent citability assessment” on each candidate paragraph — it isn’t evaluating the whole article, it’s evaluating whether each paragraph is fit to be sliced out and dropped into an answer.
How this relates to Citable Content Design: that post is the “theoretical framework” for the 5 features (why AI prefers these features); this post is the “hands-on writing manual” (good-vs-bad comparisons, rewriting techniques, and the scoring mechanism for each feature). We recommend reading that one first to build the mental model, then coming back here to apply it.
Below we break down the 5 features and the writing techniques.
Feature 1: Answer first (direct answer in the first 30 words)
Why this carries the highest weight
LLMs are extremely sensitive to the “first sentence.” If your paragraph’s opening sentence isn’t a direct answer, the AI will:
- Treat the paragraph as “setup / transition” rather than “core information”
- Skip it during slicing and look for other paragraphs that answer directly
Good vs bad
❌ Buildup-style opening
“Customer experience management is one of the important challenges facing modern enterprises. As technology advances rapidly and the business environment grows ever more complex, companies must actively consider how to use innovative tools to improve overall customer satisfaction. Many companies have rolled out CRM systems, hoping to achieve this goal through technological means. However, the process involves a remarkably broad range of aspects…”
The problem: neither the first nor the second sentence says what CRM is, so the AI skips it during slicing.
✅ Answer-first opening
“CRM (Customer Relationship Management) is a software system for managing customer data, interaction history, and the sales process, with common functions falling into three core modules: contact management, opportunity tracking, and marketing automation.”
The first sentence answers “what is CRM” outright in one breath, so the AI can cite it independently when slicing.
Writing technique
Under every H2/H3 heading, the first sentence of the first paragraph must be:
- “X is Y” (for a definition-type question)
- “There are three main reasons: A, B, C” (for a why-type question)
- “It’s done in four steps: first…” (for a how-to-type question)
- “The conclusion is Z” (for an evaluation-type question)
Delete every transitional opener like “in today’s world,” “in modern society,” and “more and more.”
Feature 2: Quantified and concrete (numbers / ratios / ranges)
Why numbers are a “citation magnet”
LLMs assign a higher confidence score to sentences that carry numbers than to plain-text sentences. The reasons:
- Verifiability: numbers are easy to cross-check against other sources
- Independence: the figure “78%” carries meaning even without context
- Answer precision: when a user asks “how much of X is Y,” a paragraph with the number lands directly on target
Good vs bad
❌ Vague quantifiers
“Many companies saw a significant boost in performance after adopting CRM.”
The problem: how many is “many”? How significant is “significant”? What does “performance” mean? The AI cites vague paragraphs like this very rarely.
✅ Quantified and concrete
“Within 6 months of adopting CRM, [company name]’s customer retention rate rose by [X]%, sales cycle time shortened by [Y]%, and lead conversion rate climbed from [A]% to [B]%.” (The example uses placeholders; when you actually write, fill in the real numbers from your own site.)
Every quantifier has a number, a unit, and a time range. When the AI slices, it can cite this independently as the core answer to the query “the benefits of adopting CRM.”
Writing technique
Find a replacement for every one of these vague quantifiers:
| Vague | Make it concrete (precondition: you actually have the figure) |
|---|---|
| Many companies | [actual ratio]% of mid-sized companies |
| Significant improvement | improved by [actual %] |
| Most | [actual ratio]% |
| Usually | generally [actual ratio] |
| Years of experience | 12 years of industry experience |
| Early on | 2018 |
| Widely used | deployed at over [actual number] companies worldwide |
Important precondition: the right column above is how to write when you have real data. If you don’t have these numbers, it’s better to write “based on industry observation” or “a rough estimate compiled by AI” — LLMs are increasingly good at detecting fabricated precise numbers (a specific percentage with no corresponding public source is a high-frequency AI-padding signal).
If you find yourself writing words like “many,” “most,” or “usually” → force yourself to convert them into concrete numbers or delete the sentence.
Feature 3: Structured (list / table / steps)
Why structure beats prose
When an LLM parses content, structural markup is cheaper than semantic understanding:
- Sees
<table>→ immediately recognizes comparison-type content - Sees
<ol>→ immediately recognizes process-type content - Sees
<ul>→ immediately recognizes list-type content - Sees plain
<p>→ needs deeper semantic analysis before it can slice
When a user asks questions like “how do A and B compare,” “what are the steps to do X,” or “what options are there,” the AI preferentially looks for content blocks that are already structured.
Good vs bad
❌ Prose-narrative comparison
“The advantage of Plan A is its lower price, at $99 per month, but it lacks advanced features like automation and an API. Plan B costs $299 per month and includes full automation, but requires a technician to set up. Plan C, at $599, includes an API and dedicated support, making it suitable for large enterprises…”
✅ Structured as a table
| Plan | Monthly fee | Automation | API | Best for |
|---|---|---|---|---|
| A | $99 | ❌ | ❌ | Small teams |
| B | $299 | ✅ | ❌ | Mid-sized companies |
| C | $599 | ✅ | ✅ | Large enterprises |
When the AI slices, it grabs the entire <table> directly, and the citation rate is 4–6x higher than the prose version.
Writing technique
Every time you finish a paragraph, ask yourself:
- Do I have ≥ 3 parallel items? → Convert to a
<ul>list - Do I have a “first A, then B, then C” sequence? → Convert to an
<ol>numbered list - Do I have a “Plan X vs Plan Y” comparison? → Convert to a
<table>
Chinese writing favors “conveying meaning through eloquence,” but in a GEO scenario structure beats literary flair.
Feature 4: Source attribution (per X / per Y)
Why attributed sentences earn bonus points
LLMs give a bonus to sentences where “the paragraph itself cites another source,” because:
- It shows the author did research rather than making things up
- It gives the AI an anchor for cross-verification
- Even if the original source can’t be verified, the “has a citation” format is itself a quality signal
Good vs bad
❌ Unsourced claim
“Studies show that companies see a 30% lift in performance after adopting CRM.”
The problem: which study? When? Who conducted it? The AI assigns low weight to unsourced numbers.
✅ Sourced claim
“According to [institution]’s [year] report《[report title]》, among Taiwanese mid-sized companies that have run CRM for two or more years, [X]% reported a positive impact on performance, with an average improvement of [Y]%–[Z]%.” (Precondition: the report genuinely exists and the numbers are genuinely verifiable.)
Even though readers won’t go check the details, the AI clearly scores this sentence with higher confidence — but only on the precondition that the cited source genuinely exists. Fabricating a study or report name gets detected by the LLM (it knows the contents of certain well-known studies) and actually costs you points.
Writing technique
Attribute a source for every piece of content that carries a number:
- “According to [institution]’s [year] [report / survey]…”
- “[name] noted in [book / talk / interview]…”
- “[company]’s [product documentation / white paper] reveals…”
Cite even industry common knowledge — the AI has no concept of “industry common knowledge”; all it sees is whether the sentence has a citation or not.
One small trick
If you genuinely can’t find a specific research source (it’s purely your own experience), you can use:
- “Across the 50+ clients I’ve worked with…” (personal evidence) — precondition: you really have worked with that scale
- “Across the projects I handled between 2018 and 2024…” (time range) — precondition: the time range and the projects are real
Personal evidence that supplies “scale + timeframe” still counts as a source — far stronger than a vague “in my experience.” But if you don’t have that scale, write honestly: “based on industry observation” or “compiled by AI” — the AI is increasingly good at detecting invented claims of personal experience.
Feature 5: Self-contained (readable without context)
Why this one is so often overlooked
When an LLM slices, what it slices is an individual paragraph, not the whole article. If your paragraph is written like this:
“As mentioned earlier, the benefits of this approach include: first… second… third…”
What does “earlier” refer to? When the AI slices down to this paragraph, the context is severed, and at citation time it either hallucinates false information or simply doesn’t cite it.
Good vs bad
❌ Context-dependent
“As mentioned earlier, this strategy’s core advantages are threefold: first, it reduces cost; second, it improves speed; third, it improves quality.”
✅ Self-contained
“The core advantages of the lean development strategy are threefold: first, it reduces cost (cutting inventory waste by 30%); second, it improves speed (shortening time-to-market by 40%); third, it improves quality (lowering the defect rate by 25%).”
The second version is comprehensible even when sliced out on its own — it tells you which strategy is being discussed and what the advantages concretely are.
Writing technique
Every time you finish a paragraph, run the “slice test”:
- Cut the paragraph out and paste it into a blank document
- Read it through yourself — can you understand it fully?
- If not, add an opening line that says “what this paragraph is about”
Common phrasings that “need to be removed”:
| Context-dependent | Self-contained rewrite |
|---|---|
| As mentioned earlier | Restate the prior conclusion directly |
| The X mentioned above | “The X in lean development” |
| This approach | “This [specific method name]” |
| Next, we’ll explain | Start explaining directly, no preview needed |
| How should this be handled | “The method for handling [specific scenario]” |
A combined score across the 5 features
Every time you finish a paragraph, run this 5-item self-check:
| Feature | Compliance check |
|---|---|
| 1. Answer first | Does the first sentence answer the topic directly? |
| 2. Quantified and concrete | Is there ≥ 1 number / range / ratio? |
| 3. Structured | Is there a list / table / numbering? |
| 4. Source attribution | Do the numbers / claims have a cited source? |
| 5. Self-contained | Is it independently readable when sliced out? |
| Items met | Tendency for citation-rate lift (rough estimate compiled by AI) |
|---|---|
| All 5 | Significant lift (several-fold) |
| 4 met | Clear lift |
| 3 met | Moderate lift |
| 2 met | Slight lift |
| ≤ 1 | Close to baseline (effectively no gain) |
The exact multiple varies a great deal by industry / content type / competitive intensity — the point isn’t to chase a precise multiple, but to rewrite your important paragraphs up to 4 or more items (the opening paragraph, the conclusion, and paragraphs that answer key questions get top priority).
A hands-on example: before-and-after rewrite
Before (baseline)
“Customer experience management is an important challenge facing modern enterprises. As market competition grows fiercer, companies need to manage customer relationships more efficiently. Many companies have adopted CRM systems, hoping to improve operational efficiency through technological means. However, adopting CRM can’t be done overnight; it requires considering many aspects, including cost, staff training, and process integration.”
5-item score: 1/5 (only “self-contained” barely passes; everything else fails)
After (all 5 met)
“CRM (Customer Relationship Management) is a software system for managing customer data, interaction history, and the sales process. Based on industry observation, common outcomes within 2 years of CRM adoption at Taiwanese mid-sized companies include: both customer retention and conversion rates trending up, and a shorter sales cycle. The adoption cycle typically takes 6–9 months, with the main bottlenecks in three areas: staff training, data migration from legacy systems, and process redesign (these three together account for most of the time).”
5-item score: 4/5 (short on concrete numbers, because if you don’t have real data, you shouldn’t make it up)
Nearly the same word count, yet citability is still clearly higher — the key is that all four features (definition sentence, structure, source attribution, self-containment) are preserved. Fill in concrete numbers if you have them; if not, honestly write that it’s based on industry observation.
Step one: run a health check to see which dimensions you’re on
👉 Free GEO health check — the health check’s “content citability” dimension examines paragraph length / structure / statistics / citation phrasing, and its “AEO readiness” dimension checks answer-first ordering / concise answers right after headings / step structure. These two dimensions most directly reflect how well you’ve hit this article’s 5 features.
If you want to run content-team training (teaching your writers to rewrite according to this article’s 5 features, including the before-and-after scoring mechanism and workflow integration), that falls within the scope of our GEO consulting service: [email protected]
GEO advanced series. Previous post: 「From AI Recommendation to Signed Contract — The 5 Touchpoints in the B2B Customer Decision Journey」