How to Optimize Content for LLMs: Writing for Machines That Don't Read Like Google
How to optimize content for LLMs: how AI pipelines chunk and embed pages, why entity definitions beat keyword stuffing, and 5 fixes for B2B articles.
September 28, 2026 · 18 min read
On this page10 sections
- Key Takeaways
- How LLM Pipelines Read a Page
- How Vector Embeddings Evaluate Topical Authority
- Why Clear Entity Definitions Beat Keyword Stuffing
- The Chunk-Ready Five: 5 Ways to Optimize Content for LLMs
- What Surfer and Clearscope Measure Now
- Worked Example: How a Chunker Cuts a Tallyfold Guide
- How to Check a Page Before You Optimize Content for LLMs
- Where Rankbox Fits
- Frequently Asked Questions
The short answer
To optimize content for LLMs, write so a machine can fetch the page without running JavaScript, cut it into chunks without losing the subject, and match each chunk to a question by meaning. That takes five formatting habits and one writing habit: define each product, feature and term by its full name where you use it, instead of repeating a keyword. Google's classic results rank whole pages. An LLM pipeline retrieves and quotes slices of them.
The slices are small. OpenAI's file search, for example, splits every file into 800-token chunks by default, about 600 words each. Perplexity says its search API treats "the individual sections and spans of documents as first-class units" in its architecture write-up. A chunk that says "it" instead of your product's name, or hides its numbers in an image, arrives at the model half empty.
Meaning beats repetition inside those chunks. In Rankbox's keyword density experiment, per point of density, a definition sentence moved similarity about 3.2 times as much as a keyword mention (median of 8 open embedding models), and repeats stopped helping after the second mention.
This guide shows how to optimize content for LLMs from the pipeline up: how the machines read a page, how embeddings see topical authority, why entity definitions win, and five formatting fixes for technical and B2B articles. For a full worked rewrite of one section, see our before-and-after rewrite of content for LLMs. For passage-by-passage editing, see how to optimize content for AI search.
Key Takeaways
- LLM pipelines fetch, parse, chunk, embed, retrieve and then write. Each step can drop part of your page, so you optimize content for LLMs step by step.
- Most AI crawlers didn't run JavaScript in Vercel and MERJ's December 2024 study. Text that loads in tabs or widgets, or sits in images, may never reach the model.
- Under OpenAI's default chunking (800 tokens, 400 overlap), any section of about 300 words or less always lands whole inside one chunk. A section over about 600 words never does.
- In embedding terms, topical authority works like coverage: how many of a topic's sub-questions find a close chunk from your site. That's a model; no engine publishes an embedding-based authority score.
- Entity definitions beat keyword stuffing. Per point of density, a definition moved similarity 3.2 times as much as a keyword mention in Rankbox's test, and repeats stopped paying after two.
- Surfer and Clearscope still hand writers term lists with usage ranges, and both have added AI search features. Use their lists as a coverage checklist when you optimize content for LLMs, not as a count target.
- Google says its own AI features need no special writing. These fixes target the pipelines that fetch and chunk pages themselves, and they read better for people too.
How LLM Pipelines Read a Page
An LLM answer engine reads your page in six steps, and each step is a place to optimize content for LLMs or to lose it. Vendors publish some steps and keep others private, so treat this table as a model built from their docs, not any one product's spec.
| Step | What the machine does | What vendors document | What your page can lose |
|---|---|---|---|
| 1. Fetch | Downloads the raw HTML | Most AI crawlers don't run JavaScript (Vercel and MERJ, 2024) | Anything rendered in the browser |
| 2. Parse | Turns HTML into text blocks | Layout parsers detect headings, lists and tables (Google Cloud) | Structure, if headings are just bold text |
| 3. Chunk | Cuts the text into pieces | 800-token chunks with 400 overlap by default (OpenAI) | Sections split mid-thought |
| 4. Embed | Turns each chunk into a vector | Chunks lose context from the rest of the page (Anthropic) | The subject, if the chunk says "it" |
| 5. Retrieve | Finds chunks near the question | Keyword and meaning search, then rerankers (Perplexity) | Chunks that match neither words nor meaning |
| 6. Write | Builds the answer from a few chunks | Fetched text can be cut at a token limit (Anthropic) | Facts placed after the cut |
Fetch and parse: the page the machine gets
The fetch step decides what exists. When Vercel and MERJ studied crawler traffic in December 2024, they found "none of the major AI crawlers currently render JavaScript", including OpenAI's OAI-SearchBot, ChatGPT-User and GPTBot, and Anthropic's ClaudeBot. Gemini is different, because it uses Googlebot's rendering. Anthropic's own web fetch tool "does not support websites dynamically rendered with JavaScript."
Parsing decides what survives as structure. Google Cloud's Agent Search docs describe two parsers. The basic one "detects text blocks, but not document elements such as tables, lists, and headings." The layout parser detects all of them in HTML. Perplexity writes parsing rules per site, and notes that a site "heavy on well-structured data in list or table form" can use more formulaic rules.
Chunk, embed and retrieve: the page in pieces
OpenAI's default chunker is a sliding window: 800 tokens per chunk, each overlapping the last by 400. A token is about three-quarters of a word, per OpenAI's token guide. Anthropic calls chunks of "no more than a few hundred tokens" typical.
Each chunk then becomes an embedding, a list of numbers that places its meaning in space (see vector embeddings). Perplexity searches its index "via both modalities," keyword and meaning, then lets cross-encoder rerankers make the final cut. Some search tools then pass the model only short excerpts of each page; our Reverse Prompt playbook works through what that budget means.
What Google does differently
Google's AI Overviews and AI Mode retrieve pages through Google's core ranking systems. Its AI optimization guide says "You don't need to write in a specific way just for generative AI search," and that there's "no requirement to break your content into tiny pieces." Take that at face value for Google. The effort to optimize content for LLMs pays off in the pipelines that fetch, chunk and embed pages themselves, and every fix below also reads better for people.
How Vector Embeddings Evaluate Topical Authority
Embeddings don't score authority directly. No AI search engine publishes a topical authority score built from embeddings, and Google has confirmed a "topic authority" system only for news, as our topical authority entry explains. What embeddings do is decide which of your chunks sit close to which questions.
Authority as coverage
AI engines split one question into several searches, a step called query fan-out. Each sub-query has its own nearest chunks. So in retrieval terms, topical authority looks like coverage: the share of a topic's sub-questions for which your site holds a chunk that is close in meaning and specific enough to quote. This is a model, not a documented ranking factor. It explains, in retrieval terms, why covering a subject's follow-up questions gives a site more chances to be picked.
Here is an illustrative coverage map for one buyer question, "best invoicing software for agencies that bill in several currencies." Tallyfold is a made-up invoicing and payments app for agencies, and Brindlework is a made-up rival.
| Likely sub-question | Tallyfold section that answers it | Brindlework section that answers it |
|---|---|---|
| Which apps invoice in several currencies? | Multi-currency invoicing | None (one line in a features list) |
| Which exchange rate is used? | How Tallyfold sets the rate | None |
| What does conversion cost? | Fees table | Pricing page |
| How fast are payouts? | Payouts | None |
| How is VAT shown on foreign invoices? | Tax on foreign invoices | None |
| Can clients pay by card and bank transfer? | Payment methods | Payment methods |
| Does it handle retainers? | Retainer billing | None |
| Does it sync with accounting tools? | Integrations | Integrations |
| Coverage | 8 of 8 | 3 of 8 |
Brindlework may be the better product. Tallyfold simply has more chunks near more of the questions an engine is likely to ask. To optimize content for LLMs across a topic, map the sub-questions first, then give each one its own section.
Why coverage matters more for small brands
Models know little about rare entities from training alone. Kandpal and colleagues showed that a model's ability to answer a factual question "relates to how many documents associated with that question were seen during pre-training," and that retrieval reduces that dependence. A B2B product with a few hundred mentions on the web is a rare entity. For it, retrieved pages are where the model learns what the product is, so your own pages have to carry the definitions.
Why Clear Entity Definitions Beat Keyword Stuffing
An entity is a specific thing with a name: a product, a feature, a company, a standard. Google framed this shift in 2012 when it launched the Knowledge Graph as "things, not strings." Embedding models make the same move at the level of each chunk. When you optimize content for LLMs, define the thing; don't repeat the string.
What the evidence shows
- Definitions move meaning. In Rankbox's test of 263 paragraphs, per point of density, a definition sentence moved similarity 3.2 times as much as a keyword mention (median of 8 open embedding models). Repeats stopped helping after the second mention. On prompts that described the need without the keyword, each extra mention lowered similarity in all 8 models.
- Self-contained facts retrieve better. The Dense X Retrieval study by Chen and colleagues indexed text as "propositions," short statements that each carry one fact and replace pronouns with full names. That beat indexing by passage, with a 25% relative gain in top-5 recall on a set of entity questions for a weaker retriever, and most of the gain came from questions about rare entities.
- New product names are hard for embeddings. Google Cloud's hybrid search docs say "brand new product names that were added recently" don't work with semantic search alone, because the embedding model never saw them in training.
Together, those two findings set the job. The exact name gives keyword search something to match, and the definition around it gives the embedding a meaning to place.
The entity definition sentence
Write one sentence per entity in this shape: [full name] is a [category] that [does what] for [whom], [with one number or limit]. Put it at the start of the section where the entity first matters.
| Keyword-first sentence | Entity-first sentence |
|---|---|
| "Our invoicing software for agencies makes invoicing software for agencies easy." | "Tallyfold is an invoicing and payments app for agencies that bills clients in 24 currencies." |
| "Retainers are simple with our tool." | "Tallyfold's Retainer Billing sends the same invoice each month and tracks hours used against the retainer." |
| "Integrates with your stack." | "Tallyfold syncs each paid invoice to your accounting software within 15 minutes." |
Every entity-first sentence can be lifted into an answer as it stands. That's the practical meaning of "optimize content for LLMs": each claim carries its own subject. For how many times to name the keyword, use the Twice-Four Rule from the keyword density study: name it twice per passage of about 100 words, and give four defining facts.
Keywords still count, a little
The claim that LLMs "completely ignore" keywords is too strong. In Rankbox's data, the second mention still gave a lift. OpenAI's retrieval API lets developers blend "semantic embedding matches vs. sparse keyword matches," and Perplexity runs both. So when you optimize content for LLMs, keep the exact product, feature and category names. Just stop at two per passage and spend the rest of the words on facts.
The Chunk-Ready Five: 5 Ways to Optimize Content for LLMs
Technical and B2B articles often hide their best facts where machines read worst: spec tables drawn as images, pricing in tabs, setup steps in PDFs. The Chunk-Ready Five is Rankbox's formatting checklist for these articles. Each fix protects one step of the pipeline.
| # | Fix | Pipeline step it protects | Quick check |
|---|---|---|---|
| 1 | Serve every fact as HTML text | Fetch and parse | The fact appears in the raw HTML source |
| 2 | Size each section to fit one chunk | Chunk | Each section runs about 120 to 300 words |
| 3 | Re-anchor each section with full names | Embed | Every section's first sentence names its subject |
| 4 | Put specs, prices and comparisons in HTML tables | Parse and write | Tables use real table tags with headers and units |
| 5 | Write steps as numbered lists and code as code blocks | Parse and retrieve | No screenshots of code, commands or settings |
1. Serve every fact as HTML text
Put every fact a buyer might ask about in the HTML your server sends. Render it on the server (see server-side rendering), and keep answers out of tabs and accordions that load on click. Microsoft's Bing team puts it plainly in an October 2025 guide: "AI systems may not render hidden content, so key details can be skipped."
The same guide warns against keeping core information only in PDFs, which "often lack the structured signals" of HTML, or only in images. That covers most B2B hiding places: spec sheets, rate cards, security whitepapers and architecture diagrams. Publish each as an HTML page and keep the PDF as a download. It's the cheapest way to optimize content for LLMs on an existing site, because nothing needs rewriting.
2. Size each section to fit one chunk
Use a heading wherever the topic changes, and keep each section to about 120 to 300 words. Bing's guide says headings "act like chapter titles that define clear content slices." Layout-aware chunkers follow them: in Google Cloud's version, "all text in a chunk will come from the same layout entity, such as headings, subheadings, and lists."
Fixed-size chunkers ignore headings, and the arithmetic explains the 300-word ceiling. With OpenAI's defaults, a new chunk starts every 400 tokens and runs for 800. Any passage of 400 tokens or fewer (about 300 words) therefore sits whole inside at least one chunk, wherever it starts. A passage longer than 800 tokens (about 600 words) never does. Between those sizes, it depends on luck. The lower bound comes from citation data: our passage-level guide shows why very short sections get cited less.
3. Re-anchor each section with full names
Open every section by naming its subject in full, as if the reader skipped straight to it, because a chunker may deliver it that way. Anthropic's contextual retrieval research exists because chunks lose this context: adding 50 to 100 tokens of context to each chunk cut failed retrievals by 35% on its own. You can do the same in the text itself.
Three habits cover it:
- Start the section with "Tallyfold's payout schedule," not "Payouts."
- Spell out each acronym once per section: "foreign exchange (FX)."
- Use one name per thing. Don't rotate "Tallyfold," "the platform," "our solution" and "the app" for variety.
4. Put specs, prices and comparisons in HTML tables
Use real table markup, with a header row, units in the headers and one fact per cell. In the Table Meets LLM study (Sui and colleagues, WSDM 2024), tables given to the model as HTML beat plain text with separators by 6.76%. Google Cloud's layout parser detects tables in HTML and can attach a written description to each one, so the table can be found by what it's about.
Add one sentence above each table that says what it shows and when it was checked. That sentence is what a retriever matches when the question is "how much does Tallyfold charge to convert currencies?"
5. Write steps as numbered lists and code as code blocks
Write procedures as numbered steps, one action per step, using the exact button and field names from the product. Put code, API calls, config and error messages in code blocks as text, never as screenshots. Exact strings such as invoice.currency, error codes and version numbers are the kind of "out of domain" terms that keyword search can match and embeddings may miss, so they belong in the text word for word. For API guides and developer docs, this fix does most of the work to optimize content for LLMs.
What Surfer and Clearscope Measure Now
The content plan behind this post claimed that SEO writing tools like Surfer and Clearscope "still obsess over keyword density and TF-IDF counts that large language models completely ignore." Half of that holds up as of 28 September 2026. Both tools still build guidance from term lists drawn from Google's top results, with suggested counts. But none of the Surfer or Clearscope pages checked for this post describes the method as TF-IDF, and both tools have added AI search features.
| Surfer | Clearscope | |
|---|---|---|
| Term guidance | "Terms to use" list; its Audit glossary says suggested counts use "the average density on competitors' pages" | Importance from 1 to 10 and a "typical uses" range per term, per its grading guide |
| Score | Content Score, split since May 2026 into SEO Score and AI Search Score | Content Grade, F to A++ |
| AI search features | AI Search Score rates Facts Coverage and Upfront Intent Alignment; AI Tracker covers five AI engines | AI Term Presence shows which terms appear in GPT and Gemini answers; Query Tracking and fan-out data on every plan |
| Entry price, as listed on 28 September 2026 | Standard, $99 a month billed yearly, with 25 AI prompts tracked weekly | Essentials, $129 a month |
Surfer's own Content Score guide now says "how those terms are used in context matters more than how often they appear," and advises writers to "avoid over-optimizing any single suggested term." Clearscope calls its typical-use range "a guide rather than a rule." Clearscope's pages disagree on engines, too: its pricing page lists ChatGPT and Gemini for Query Tracking, and its homepage adds Claude.
So state the gap this way. Count targets were built for lexical ranking, and LLM pipelines stop rewarding repeats early. Both tools' term lists are still useful when you optimize content for LLMs, as a checklist of entities and facts to cover. Treat the suggested counts as a ceiling, not a target.
Worked Example: How a Chunker Cuts a Tallyfold Guide
Here is the arithmetic behind fix 2, on a made-up 2,400-word Tallyfold guide to multi-currency invoicing. Token counts use OpenAI's rule of thumb of 100 tokens per 75 words, and the chunker uses OpenAI's file search defaults.
| Section | Words | Tokens | Position in tokens | Whole inside one chunk? |
|---|---|---|---|---|
| Introduction | 150 | 200 | 0–200 | Yes, always |
| What multi-currency invoicing is | 250 | 333 | 200–533 | Yes, always (under 400 tokens) |
| How exchange rates are set | 700 | 933 | 533–1,467 | Never (over 800 tokens) |
| Fees | 200 | 267 | 1,467–1,733 | Yes, always |
| Tax on foreign invoices | 450 | 600 | 1,733–2,333 | Yes, by luck (chunk 1,600–2,400) |
| Payouts | 300 | 400 | 2,333–2,733 | Yes, always |
| FAQ | 350 | 467 | 2,733–3,200 | Yes, by luck (chunk 2,400–3,200) |
The guide makes seven chunks, starting at token 0, 400, 800 and so on up to 2,400. The exchange-rate section is spread across four of them, so no single chunk holds the whole answer to "what rate does Tallyfold use?"
Now add 100 words to the introduction. Every later section moves down 133 tokens. The tax section now runs from about 1,867 to 2,467, which fits neither the chunk from 1,600 to 2,400 nor the one from 2,000 to 2,800. A paragraph added at the top of the page just split a section near the bottom.
The fix is to split the long sections at natural headings. Break the exchange-rate section into three H3s of about 230 words (when the rate is locked, which rate is used, what happens if it moves before payment), and the tax section into two of 225. Every section is then under 400 tokens, so each one lands whole in some chunk no matter what changes above it.
Two caveats keep this honest. This is OpenAI's default for developer file search, not a documented setting of ChatGPT search. And layout-aware chunkers cut at headings instead, which the same fix also serves. Either way, section length is part of how you optimize content for LLMs, not a style choice.
How to Check a Page Before You Optimize Content for LLMs
Five checks show where a page loses facts, in about 20 minutes per page and with no paid tools.
- Fetch it raw. Run
curl -s https://tallyfold.example/guide | grep -i "0.9%"with a fact from your page. If the fact isn't in the raw HTML, a crawler that doesn't run JavaScript won't see it. - Outline the headings. Paste the page's HTML into our heading structure checker. Every topic change should have its own H2 or H3.
- Count words per section. Flag any section over 300 words, and split anything over 600.
- Read each section's first sentence alone. It should name the product or feature in full.
- Count exact-keyword mentions per passage. Our keyword density checker counts them. Trim anything past two per 100 words.
For the basics around the content, the AI search readiness check scans a URL for 12 signals, including AI crawler access, a single H1 and structured data.
Where Rankbox Fits
If you'd rather not optimize content for LLMs by hand on every new article, Rankbox's Citation-Ready Writer researches the live web and writes 2,000–3,500-word, source-backed articles with headings, definitions and tables. Brand Voice applies the tone, audience, style rules and product details you give it, and mentions your product where it fits. Articles reach your site through Rankbox's API on the Business plan, $49.50 a month with a 7-day trial.
Rankbox doesn't invent your spec table, though. It doesn't import proprietary data or run benchmarks, so your real limits, fees and integration details have to come from you, through the product details in Brand Voice. It also doesn't track AI citations today or publish straight to a CMS. See plans and pricing.
Frequently Asked Questions
How do you optimize content for LLMs?
To optimize content for LLMs, make every fact available as HTML text, split pages into sections of about 120 to 300 words under clear headings, and open each section by naming its subject in full. Define products and features instead of repeating keywords, and put specs, prices and comparisons in real tables.
Do LLMs care about keyword density?
A little, up to a point. In Rankbox's September 2026 test with eight open embedding models, the second mention of a keyword raised similarity, but later mentions didn't help, and definitions moved similarity 3.2 times as much per point of density. Many pipelines also run keyword search, so keep exact names.
What is chunking in LLM search?
Chunking is cutting a page's text into short pieces before turning each piece into an embedding. OpenAI's file search uses 800-token chunks that overlap by 400 tokens by default, which is why section length matters when you optimize content for LLMs. Our content chunking entry has more.
How long should a section be to optimize content for LLMs?
About 120 to 300 words works well. Under OpenAI's default chunking, a section of about 300 words or less always fits whole inside one chunk, while one over about 600 words is always split. Google says pages don't need to be cut into tiny pieces, so don't go much shorter.
Can LLM crawlers read JavaScript content?
Mostly no. Vercel and MERJ found in December 2024 that none of the major AI crawlers, including OpenAI's and Anthropic's, rendered JavaScript. Google's Gemini is different because it uses Googlebot's rendering. Put important content in the HTML your server sends.
Do Surfer and Clearscope help you optimize content for LLMs?
Partly. As of September 2026, Surfer scores "AI Search" on facts coverage and early answers, and Clearscope shows which terms appear in GPT and Gemini answers. Both still suggest term counts from Google's top results, so use their lists to check coverage and treat the counts as a ceiling.
References
- 1.Retrieval guide (vector stores and chunking), OpenAIdevelopers.openai.com ↗
- 2.What are tokens and how to count them?, OpenAI Help Centerhelp.openai.com ↗
- 3.Architecting and evaluating an AI-first search API, Perplexityperplexity.ai ↗
- 4.The rise of the AI crawler, Vercel and MERJvercel.com ↗
- 5.Web fetch tool, Claude Developer Platformplatform.claude.com ↗
- 6.Introducing Contextual Retrieval, Anthropicanthropic.com ↗
- 7.Parse and chunk documents, Google Cloud Agent Searchcloud.google.com ↗
- 8.About hybrid search, Google Clouddocs.cloud.google.com ↗
- 9.Optimizing your website for generative AI features on Google Search, Google Search Centraldevelopers.google.com ↗
- 10.Optimizing your content for inclusion in AI search answers, Microsoft Advertisingabout.ads.microsoft.com ↗
- 11.Introducing the Knowledge Graph: things, not strings, Googleblog.google ↗
- 12.Dense X Retrieval: What Retrieval Granularity Should We Use? (Chen et al.), arXivarxiv.org ↗
- 13.Large Language Models Struggle to Learn Long-Tail Knowledge (Kandpal et al., ICML 2023), arXivarxiv.org ↗
- 14.Table Meets LLM (Sui et al., WSDM 2024), arXivarxiv.org ↗
- 15.Content Score in the Editor explained, Surferdocs.surferseo.com ↗
- 16.Audit glossary, Surferdocs.surferseo.com ↗
- 17.Pricing, Surfersurferseo.com ↗
- 18.How does Clearscope grade your content?, Clearscopeclearscope.io ↗
- 19.Pricing, Clearscopeclearscope.io ↗
- 20.AI Tracker, Surfersurferseo.com ↗
