On this page10 sections
The short answer
To get cited by Perplexity, allow PerplexityBot and Perplexity-User, serve your content as server-rendered HTML, and keep it visibly current. Perplexity searches its own index of over 200 billion URLs, splits each page into self-contained passages, and reranks those passages for every question — so it cites pages whose individual sections answer a question directly, with names, numbers and a date.
Key takeaways
- Perplexity isn’t a Google wrapper: it runs its own index of 200B+ URLs on Vespa, with its own crawler,
PerplexityBot. - It ranks passages, not pages — documents are split into “self-contained spans” and reranked by cross-encoder models.
- Stale content is filtered out before ranking, and citations skew about 250 days fresher than Google’s results.
PerplexityBothonors robots.txt;Perplexity-User, which fetches live for users, “generally ignores” it.- Independent tests found Perplexity fails on client-side-rendered pages. Server-render everything you want cited.
- Perplexity sells no ads or placement. Comet Plus pays partner publishers about 80% of its revenue.
How Perplexity crawls, indexes and ranks#
Perplexity is unusually open about its search stack. Its architecture write-up, published with the Search API in September 2025, explains that it started on third-party search APIs and left them over cost, staleness, latency and “document-level granularity.” The replacement — built on Vespa — “tracks over 200 billion unique URLs” and runs “tens of thousands of indexing operations per second.” Every stage below is from that document.
- 1
A model decides when to crawl each URL
An ML model predicts whether a URL needs indexing and when, calibrated to “the importance and likely update frequency of the specific URL.” Documents from “authoritative domains” and “undercovered topics” are kept hot.
Your lever: Earn links from authoritative sites, update pages substantively, and cover the gaps others don’t.
- 2
Pages are parsed into self-contained spans
A content-understanding module splits each document into spans that are “individually retrieved and ranked at query time.” List- and table-heavy sites get “more formulaic parsing.”
Your lever: Write sections that make sense on their own, and put specs and comparisons in real HTML tables.
- 3
Keyword and semantic retrieval run together
Lexical and embedding retrieval run in parallel and merge into one candidate set, so both exact terms and meaning count.
Your lever: Use the words buyers use — product names, units, category terms — alongside natural explanation.
- 4
Stale and non-responsive content is filtered out
Prefilters remove “clearly non-responsive or stale content” before scoring. The index stores publish and last-updated dates per page — the Sonar API even filters on them.
Your lever: Show a real published and updated date, and keep them honest.
- 5
Cross-encoders rerank passages
Fast lexical and embedding scorers narrow the set; cross-encoder rerankers make the final cut, scoring at document and sub-document level. Rankers keep training on signals from “millions of user requests… each hour.”
Your lever: Answer the question in the first sentence of the section that covers it.
- 6
The answer cites its best passages
A standard answer cites 4–8 sources, shown as numbered cards above the text. When the index isn’t enough,
Perplexity-Userfetches pages live. Research mode “performs dozens of searches, reads hundreds of sources.”
200B+
unique URLs tracked by Perplexity's own index
28.6%
of Perplexity's citations rank in Google's top 10 — the highest of any assistant (Bing: 16.6%)
~250 days
fresher, on average, than the pages in Google's organic results
PerplexityBot, Perplexity-User and robots.txt#
Perplexity documents two bots, with IP ranges published for each. Changes to robots.txt “may take up to 24 hours” to apply.
PerplexityBotAllowSearch indexHonors robots.txtSurfaces and links websites in Perplexity’s search results; “not used to crawl content for AI foundation models.” Honors robots.txt rate limits and backs off when a site struggles. IPs:
perplexity.com/perplexitybot.json.Perplexity-UserAllowUser-triggeredIgnores robots.txtFetches a page live when a user’s question needs it; not used for crawling or training. IPs:
perplexity.com/perplexity-user.json. Perplexity: "this fetcher generally ignores robots.txt rules."
# Eligible for Perplexity search and live fetchesUser-agent: PerplexityBotAllow: /Disallow: /account/Disallow: /cart/ User-agent: Perplexity-UserAllow: / Sitemap: https://example.com/sitemap.xmlIt’s worth knowing the history. In August 2025 Cloudflare reported that when its declared bots were blocked, Perplexity fetched pages with an undeclared Chrome user agent from IPs outside its published ranges, and removed it from its verified-bot list. Perplexity responded that user-requested fetches aren’t crawling and that Cloudflare had misattributed traffic from a third-party browser service. Practically: if you want Perplexity’s traffic, don’t rely on Cloudflare’s default AI-bot rules to get the policy right — set it explicitly.
Technical requirements#
Perplexity has no webmaster console, doesn’t take URL submissions and isn’t an IndexNow participant. Everything rests on its crawler finding, fetching and parsing your pages cleanly — which puts the weight on rendering, speed and structure.
- Server-rendered HTMLRequired
- PerplexityBot didn’t render JavaScript in Vercel’s study, and Glenn Gabe found Perplexity “failed at finding the content for every url I tested” on client-rendered pages.
- Bots allowed at the WAFRequired
- Allow both published IP ranges through bot-fight modes, AI-bot managed rules and rate limits. Refresh the IP lists automatically.
- Fast, stable responsesHelps
Perplexity-Userfetches in real time. Timeouts, 403s or 429s mean the answer is built from someone else’s page.- Accurate datesHelps
- A visible published/updated date plus
datePublishedanddateModified. The index stores dates and prefilters stale content — never fake a refresh. - HTML tables & listsHelps
- Perplexity parses list- and table-heavy content “formulaically” — structure is extracted, not guessed.
- Canonicals on syndicated copiesHelps
- Columbia’s Tow Center caught Perplexity citing republished copies instead of originals. Require partners to canonicalize back to you.
- SitemapsUnconfirmed
- Not documented. Keep one with honest
lastmodvalues, referenced from robots.txt, for general discovery. - HTML versions of PDFsUnconfirmed
- Perplexity cites PDFs, but how it handles open-web PDFs isn’t documented. Publish key reports as HTML too, and never as scans.
- IndexNowNo effect
- Perplexity isn’t on IndexNow’s list of participating engines.
- llms.txtNo effect
- Perplexity publishes its own for developers, but has never said PerplexityBot reads other sites’ files or uses them to rank.
curl -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)" \ https://yoursite.com/pricing | grep -c "Plans start at" # 0 = client-rendered text, or a WAF challengeWhat Perplexity cites#
Of all the assistants, Perplexity’s citations overlap most with classic search — so strong SEO carries over here better than anywhere. On top of that, its passage-level design rewards sections that stand alone, and its freshness filtering rewards pages that stay current.
Passage quality
OfficialSpans are retrieved and reranked individually. A section that answers in its first sentence — with names, numbers and units — competes on its own.
Freshness
OfficialStale content is prefiltered, recrawls follow update frequency, and citations are ordered newest-first with an average age well below Google’s results.
Authoritative domains
OfficialDocuments from authoritative domains get crawl and storage priority — they’re kept hot in the index.
Classic rankings
Observed28.6% of Perplexity’s citations rank in Google’s top 10, the highest of any assistant. BrightEdge’s older data put domain overlap near 60%.
Community and video
ObservedReddit is 46.7% of citations among Perplexity’s top-10 domains — but 6.6% of all citations. YouTube follows. Useful as a second path, not a strategy.
The "59 ranking factors" leak
Our readParameter names found by inspecting Perplexity’s web app — rerankers, time decay, curated domains — are unconfirmed and several look like Discover feed settings. Don’t build tactics on them.
Comet Plus, publishers and ads#
Perplexity is the only major answer engine that pays publishers per use. Comet Plus costs $5 a month on its own and is included in Pro and Max; Perplexity says it distributes the revenue “minus a small portion for compute,” reported as 80% to publishers from an initial $42.5M pool. It pays on three kinds of traffic: human visits, search citations and agent actions. Partners include Condé Nast, Fortune, Le Monde, the LA Times and The Washington Post; publishers can apply via publishers@perplexity.ai.
Ads are gone. Sponsored follow-up questions ran from late 2024, new advertisers stopped around October 2025, and in February 2026 executives said Perplexity isn’t pursuing ads. There is no paid placement in answers.
Tracking Perplexity traffic#
Clicks arrive with a perplexity.ai referrer and no UTM parameters. GA4’s new AI Assistant channel doesn’t name Perplexity — Google’s documentation lists ChatGPT, Gemini, DeepSeek, Copilot and Grok — and practitioners report its traffic still lands in Referral. Give it a custom channel, placed above Referral:
# Session source — Perplexity(^|\.)perplexity\.ai$ # All major AI assistantsperplexity\.ai|chatgpt\.com|openai\.com|gemini\.google\.com|copilot\.microsoft\.com|claude\.aiIn your logs, Perplexity-User hits on a URL are the best evidence that the page is being pulled into live answers. Verify the source IP against both JSON files — the user agents are trivial to spoof:
grep -oE "PerplexityBot|Perplexity-User" access.log | sort | uniq -c # Pages Perplexity fetched live for usersgrep "Perplexity-User" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20Myths worth dropping#
Myth
Perplexity is just Google with an LLM on top.
Reality
It runs its own 200B-URL index and crawler. Google rankings correlate — 28.6% top-10 overlap — but they’re not what Perplexity searches.
Myth
Reddit is half of Perplexity's citations.
Reality
Reddit is 46.7% of the share held by the top-10 cited domains, but 6.6% of all citations (Profound).
Myth
Blocking PerplexityBot keeps you out of Perplexity.
Reality
Your domain, headline and a short summary may still appear, Perplexity-User generally ignores robots.txt, and syndicated copies can be cited instead.
Myth
Submit URLs through IndexNow, llms.txt or a webmaster console.
Reality
None of these exist for Perplexity. Discovery is its crawler and links from sites it already trusts.
The action checklist#
Everything above, in the order we’d do it. Tick items off as you go — your progress is saved in this browser.
0 of 15 done
Perplexity SEO: frequently asked questions#
How does Perplexity choose its sources?
It searches its own index with keyword and semantic retrieval together, filters out stale or off-topic pages, then reranks individual passages with cross-encoder models. A standard answer cites the 4–8 passages that best answer the question.
Does Perplexity use Google or Bing?
No. Since 2025 it has run its own crawler and an index of more than 200 billion URLs, supplemented by third-party crawlers that must respect robots.txt. Pages that rank in Google’s top 10 are still cited more often than average — about 29% overlap, per Ahrefs.
How do I get my website cited by Perplexity?
Allow PerplexityBot and Perplexity-User, make sure your content is in the server-rendered HTML, and write sections that answer a question in their first sentence with accurate dates. Mentions on sites Perplexity cites heavily for your topic help too.
If I block PerplexityBot, will Perplexity stop using my content?
It stops indexing your full text, but may still show your domain, headline and a brief summary. Perplexity-User generally ignores robots.txt, so a hard block needs a WAF rule.
Can Perplexity read JavaScript-rendered sites?
Independent tests say no: Vercel’s study found PerplexityBot doesn’t render JavaScript, and Glenn Gabe’s 2025 test found it failed on every client-side-rendered page. Content that only appears after JavaScript runs is effectively invisible.
How do I track Perplexity traffic in GA4?
Clicks arrive from perplexity.ai with no UTM tags and usually land in Referral — Google doesn’t list Perplexity in the AI Assistant channel. Build a custom channel with a perplexity\.ai regex above Referral.
Does freshness matter for Perplexity?
Yes. Perplexity schedules recrawls by update frequency, filters stale content before ranking, and orders citations newest-first. Ahrefs found its citations average about 250 days fresher than Google’s organic results.
Does Perplexity pay publishers or sell ads?
Comet Plus shares about 80% of its revenue with partner publishers, paying on visits, citations and agent actions. Perplexity stopped pursuing ads in February 2026 and sells no placement in answers.
Sources
- 1.Perplexity crawlersPerplexity Docs · docs.perplexity.ai ↗
- 2.How does Perplexity follow robots.txt?Perplexity Help Center · perplexity.ai ↗
- 3.Architecting and evaluating an AI-first search APIPerplexity Research · research.perplexity.ai ↗
- 4.Sonar search filtersPerplexity Docs · docs.perplexity.ai ↗
- 5.Perplexity builds AI search at scale on Vespa.aiVespa · blog.vespa.ai ↗
- 6.Introducing Comet PlusPerplexity · perplexity.ai ↗
- 7.Agents or bots? Making sense of AI on the open webPerplexity · perplexity.ai ↗
- 8.Perplexity is using stealth, undeclared crawlersCloudflare · blog.cloudflare.com ↗
- 9.Perplexity APIs on Samsung GalaxyPerplexity · perplexity.ai ↗
- 10.AI search overlap with Google and BingAhrefs · ahrefs.com ↗
- 11.Do AI assistants prefer to cite fresh content?Ahrefs · ahrefs.com ↗
- 12.AI platform citation patternsProfound · tryprofound.com ↗
- 13.We compared eight AI search engines. They're all bad at citing newsColumbia Journalism Review · cjr.org ↗
- 14.The rise of the AI crawlerVercel · vercel.com ↗
- 15.AI search and JavaScript renderingGSQi (Glenn Gabe) · gsqi.com ↗
- 16.Participating search enginesIndexNow · indexnow.org ↗
- 17.Default channel group (AI Assistant)Google Analytics Help · support.google.com ↗
