Perplexity SEO: the technical guide to Perplexity citations

Perplexity runs its own search engine — a 200-billion-URL index that splits pages into passages and reranks them per question. Here's how that pipeline works, which bot to let in, and what it rewards.

Updated 9 min read17 cited sources

See if Perplexity cites you · Free 7-day trial

Perplexity

Plannora vs Loopcraft for a five-person startup

Sources

Plannora pricing: free for teams of up to 5

plannora.io1

Plannora vs Loopcraft: 2026 comparison

stackreview.co2

We switched from Loopcraft to Plannora

founderforum.net3

Answer

Plannora is the better fit for most five-person teams: its free tier covers up to five users, while Loopcraft starts at $8 per seat.12

Illustration: where a citation appears in Perplexity. Brands are fictional.
URLs in Perplexity's own index
200B+
Sources cited per answer, on average
4–8
The crawler to allow for search
PerplexityBot
Of citations rank in Google's top 10
28.6%
On this page10 sections

The short answer

To get cited by Perplexity, allow PerplexityBot and Perplexity-User, serve your content as server-rendered HTML, and keep it visibly current. Perplexity searches its own index of over 200 billion URLs, splits each page into self-contained passages, and reranks those passages for every question — so it cites pages whose individual sections answer a question directly, with names, numbers and a date.

Key takeaways

  • Perplexity isn’t a Google wrapper: it runs its own index of 200B+ URLs on Vespa, with its own crawler, PerplexityBot.
  • It ranks passages, not pages — documents are split into “self-contained spans” and reranked by cross-encoder models.
  • Stale content is filtered out before ranking, and citations skew about 250 days fresher than Google’s results.
  • PerplexityBot honors robots.txt; Perplexity-User, which fetches live for users, “generally ignores” it.
  • Independent tests found Perplexity fails on client-side-rendered pages. Server-render everything you want cited.
  • Perplexity sells no ads or placement. Comet Plus pays partner publishers about 80% of its revenue.

How Perplexity crawls, indexes and ranks#

Perplexity is unusually open about its search stack. Its architecture write-up, published with the Search API in September 2025, explains that it started on third-party search APIs and left them over cost, staleness, latency and “document-level granularity.” The replacement — built on Vespa — “tracks over 200 billion unique URLs” and runs “tens of thousands of indexing operations per second.” Every stage below is from that document.

  1. 1

    A model decides when to crawl each URL

    An ML model predicts whether a URL needs indexing and when, calibrated to “the importance and likely update frequency of the specific URL.” Documents from “authoritative domains” and “undercovered topics” are kept hot.

    Your lever: Earn links from authoritative sites, update pages substantively, and cover the gaps others don’t.

  2. 2

    Pages are parsed into self-contained spans

    A content-understanding module splits each document into spans that are “individually retrieved and ranked at query time.” List- and table-heavy sites get “more formulaic parsing.”

    Your lever: Write sections that make sense on their own, and put specs and comparisons in real HTML tables.

  3. 3

    Keyword and semantic retrieval run together

    Lexical and embedding retrieval run in parallel and merge into one candidate set, so both exact terms and meaning count.

    Your lever: Use the words buyers use — product names, units, category terms — alongside natural explanation.

  4. 4

    Stale and non-responsive content is filtered out

    Prefilters remove “clearly non-responsive or stale content” before scoring. The index stores publish and last-updated dates per page — the Sonar API even filters on them.

    Your lever: Show a real published and updated date, and keep them honest.

  5. 5

    Cross-encoders rerank passages

    Fast lexical and embedding scorers narrow the set; cross-encoder rerankers make the final cut, scoring at document and sub-document level. Rankers keep training on signals from “millions of user requests… each hour.”

    Your lever: Answer the question in the first sentence of the section that covers it.

  6. 6

    The answer cites its best passages

    A standard answer cites 4–8 sources, shown as numbered cards above the text. When the index isn’t enough, Perplexity-User fetches pages live. Research mode “performs dozens of searches, reads hundreds of sources.”

200B+

unique URLs tracked by Perplexity's own index

Perplexity, Sep 2025

28.6%

of Perplexity's citations rank in Google's top 10 — the highest of any assistant (Bing: 16.6%)

Ahrefs, Aug 2025

~250 days

fresher, on average, than the pages in Google's organic results

Ahrefs, Jul 2025

PerplexityBot, Perplexity-User and robots.txt#

Perplexity documents two bots, with IP ranges published for each. Changes to robots.txt “may take up to 24 hours” to apply.

  • PerplexityBotAllowSearch indexHonors robots.txt

    Surfaces and links websites in Perplexity’s search results; “not used to crawl content for AI foundation models.” Honors robots.txt rate limits and backs off when a site struggles. IPs: perplexity.com/perplexitybot.json.

  • Perplexity-UserAllowUser-triggeredIgnores robots.txt

    Fetches a page live when a user’s question needs it; not used for crawling or training. IPs: perplexity.com/perplexity-user.json. Perplexity: "this fetcher generally ignores robots.txt rules."

robots.txt
# Eligible for Perplexity search and live fetches
User-agent: PerplexityBot
Allow: /
Disallow: /account/
Disallow: /cart/
 
User-agent: Perplexity-User
Allow: /
 
Sitemap: https://example.com/sitemap.xml

It’s worth knowing the history. In August 2025 Cloudflare reported that when its declared bots were blocked, Perplexity fetched pages with an undeclared Chrome user agent from IPs outside its published ranges, and removed it from its verified-bot list. Perplexity responded that user-requested fetches aren’t crawling and that Cloudflare had misattributed traffic from a third-party browser service. Practically: if you want Perplexity’s traffic, don’t rely on Cloudflare’s default AI-bot rules to get the policy right — set it explicitly.

Technical requirements#

Perplexity has no webmaster console, doesn’t take URL submissions and isn’t an IndexNow participant. Everything rests on its crawler finding, fetching and parsing your pages cleanly — which puts the weight on rendering, speed and structure.

Server-rendered HTMLRequired
PerplexityBot didn’t render JavaScript in Vercel’s study, and Glenn Gabe found Perplexity “failed at finding the content for every url I tested” on client-rendered pages.
Bots allowed at the WAFRequired
Allow both published IP ranges through bot-fight modes, AI-bot managed rules and rate limits. Refresh the IP lists automatically.
Fast, stable responsesHelps
Perplexity-User fetches in real time. Timeouts, 403s or 429s mean the answer is built from someone else’s page.
Accurate datesHelps
A visible published/updated date plus datePublished and dateModified. The index stores dates and prefilters stale content — never fake a refresh.
HTML tables & listsHelps
Perplexity parses list- and table-heavy content “formulaically” — structure is extracted, not guessed.
Canonicals on syndicated copiesHelps
Columbia’s Tow Center caught Perplexity citing republished copies instead of originals. Require partners to canonicalize back to you.
SitemapsUnconfirmed
Not documented. Keep one with honest lastmod values, referenced from robots.txt, for general discovery.
HTML versions of PDFsUnconfirmed
Perplexity cites PDFs, but how it handles open-web PDFs isn’t documented. Publish key reports as HTML too, and never as scans.
IndexNowNo effect
Perplexity isn’t on IndexNow’s list of participating engines.
llms.txtNo effect
Perplexity publishes its own for developers, but has never said PerplexityBot reads other sites’ files or uses them to rank.
bash
curl -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)" \
https://yoursite.com/pricing | grep -c "Plans start at"
 
# 0 = client-rendered text, or a WAF challenge

What Perplexity cites#

Of all the assistants, Perplexity’s citations overlap most with classic search — so strong SEO carries over here better than anywhere. On top of that, its passage-level design rewards sections that stand alone, and its freshness filtering rewards pages that stay current.

  • Passage quality

    Official

    Spans are retrieved and reranked individually. A section that answers in its first sentence — with names, numbers and units — competes on its own.

  • Freshness

    Official

    Stale content is prefiltered, recrawls follow update frequency, and citations are ordered newest-first with an average age well below Google’s results.

  • Authoritative domains

    Official

    Documents from authoritative domains get crawl and storage priority — they’re kept hot in the index.

  • Classic rankings

    Observed

    28.6% of Perplexity’s citations rank in Google’s top 10, the highest of any assistant. BrightEdge’s older data put domain overlap near 60%.

  • Community and video

    Observed

    Reddit is 46.7% of citations among Perplexity’s top-10 domains — but 6.6% of all citations. YouTube follows. Useful as a second path, not a strategy.

  • The "59 ranking factors" leak

    Our read

    Parameter names found by inspecting Perplexity’s web app — rerankers, time decay, curated domains — are unconfirmed and several look like Discover feed settings. Don’t build tactics on them.

Comet Plus, publishers and ads#

Perplexity is the only major answer engine that pays publishers per use. Comet Plus costs $5 a month on its own and is included in Pro and Max; Perplexity says it distributes the revenue “minus a small portion for compute,” reported as 80% to publishers from an initial $42.5M pool. It pays on three kinds of traffic: human visits, search citations and agent actions. Partners include Condé Nast, Fortune, Le Monde, the LA Times and The Washington Post; publishers can apply via publishers@perplexity.ai.

Ads are gone. Sponsored follow-up questions ran from late 2024, new advertisers stopped around October 2025, and in February 2026 executives said Perplexity isn’t pursuing ads. There is no paid placement in answers.

Tracking Perplexity traffic#

Clicks arrive with a perplexity.ai referrer and no UTM parameters. GA4’s new AI Assistant channel doesn’t name Perplexity — Google’s documentation lists ChatGPT, Gemini, DeepSeek, Copilot and Grok — and practitioners report its traffic still lands in Referral. Give it a custom channel, placed above Referral:

GA4 regex
# Session source — Perplexity
(^|\.)perplexity\.ai$
 
# All major AI assistants
perplexity\.ai|chatgpt\.com|openai\.com|gemini\.google\.com|copilot\.microsoft\.com|claude\.ai

In your logs, Perplexity-User hits on a URL are the best evidence that the page is being pulled into live answers. Verify the source IP against both JSON files — the user agents are trivial to spoof:

bash
grep -oE "PerplexityBot|Perplexity-User" access.log | sort | uniq -c
 
# Pages Perplexity fetched live for users
grep "Perplexity-User" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20

Myths worth dropping#

Myth

Perplexity is just Google with an LLM on top.

Reality

It runs its own 200B-URL index and crawler. Google rankings correlate — 28.6% top-10 overlap — but they’re not what Perplexity searches.

Myth

Reddit is half of Perplexity's citations.

Reality

Reddit is 46.7% of the share held by the top-10 cited domains, but 6.6% of all citations (Profound).

Myth

Blocking PerplexityBot keeps you out of Perplexity.

Reality

Your domain, headline and a short summary may still appear, Perplexity-User generally ignores robots.txt, and syndicated copies can be cited instead.

Myth

Submit URLs through IndexNow, llms.txt or a webmaster console.

Reality

None of these exist for Perplexity. Discovery is its crawler and links from sites it already trusts.

The action checklist#

Everything above, in the order we’d do it. Tick items off as you go — your progress is saved in this browser.

0 of 15 done

Perplexity SEO: frequently asked questions#

How does Perplexity choose its sources?

It searches its own index with keyword and semantic retrieval together, filters out stale or off-topic pages, then reranks individual passages with cross-encoder models. A standard answer cites the 4–8 passages that best answer the question.

Does Perplexity use Google or Bing?

No. Since 2025 it has run its own crawler and an index of more than 200 billion URLs, supplemented by third-party crawlers that must respect robots.txt. Pages that rank in Google’s top 10 are still cited more often than average — about 29% overlap, per Ahrefs.

How do I get my website cited by Perplexity?

Allow PerplexityBot and Perplexity-User, make sure your content is in the server-rendered HTML, and write sections that answer a question in their first sentence with accurate dates. Mentions on sites Perplexity cites heavily for your topic help too.

If I block PerplexityBot, will Perplexity stop using my content?

It stops indexing your full text, but may still show your domain, headline and a brief summary. Perplexity-User generally ignores robots.txt, so a hard block needs a WAF rule.

Can Perplexity read JavaScript-rendered sites?

Independent tests say no: Vercel’s study found PerplexityBot doesn’t render JavaScript, and Glenn Gabe’s 2025 test found it failed on every client-side-rendered page. Content that only appears after JavaScript runs is effectively invisible.

How do I track Perplexity traffic in GA4?

Clicks arrive from perplexity.ai with no UTM tags and usually land in Referral — Google doesn’t list Perplexity in the AI Assistant channel. Build a custom channel with a perplexity\.ai regex above Referral.

Does freshness matter for Perplexity?

Yes. Perplexity schedules recrawls by update frequency, filters stale content before ranking, and orders citations newest-first. Ahrefs found its citations average about 250 days fresher than Google’s organic results.

Does Perplexity pay publishers or sell ads?

Comet Plus shares about 80% of its revenue with partner publishers, paying on visits, citations and agent actions. Perplexity stopped pursuing ads in February 2026 and sells no placement in answers.

Sources

  1. 1.Perplexity crawlersPerplexity Docs · docs.perplexity.ai
  2. 2.How does Perplexity follow robots.txt?Perplexity Help Center · perplexity.ai
  3. 3.Architecting and evaluating an AI-first search APIPerplexity Research · research.perplexity.ai
  4. 4.Sonar search filtersPerplexity Docs · docs.perplexity.ai
  5. 5.Perplexity builds AI search at scale on Vespa.aiVespa · blog.vespa.ai
  6. 6.Introducing Comet PlusPerplexity · perplexity.ai
  7. 7.Agents or bots? Making sense of AI on the open webPerplexity · perplexity.ai
  8. 8.Perplexity is using stealth, undeclared crawlersCloudflare · blog.cloudflare.com
  9. 9.Perplexity APIs on Samsung GalaxyPerplexity · perplexity.ai
  10. 10.AI search overlap with Google and BingAhrefs · ahrefs.com
  11. 11.Do AI assistants prefer to cite fresh content?Ahrefs · ahrefs.com
  12. 12.AI platform citation patternsProfound · tryprofound.com
  13. 13.We compared eight AI search engines. They're all bad at citing newsColumbia Journalism Review · cjr.org
  14. 14.The rise of the AI crawlerVercel · vercel.com
  15. 15.AI search and JavaScript renderingGSQi (Glenn Gabe) · gsqi.com
  16. 16.Participating search enginesIndexNow · indexnow.org
  17. 17.Default channel group (AI Assistant)Google Analytics Help · support.google.com

Found this useful? Share it with whoever owns your SEO.

Written by

Rankbox Team

The team behind Rankbox. We study how ChatGPT, Perplexity, Gemini, and Google AI Overviews choose their sources, and publish what we learn so you can put it to work.

See if Perplexity cites you today

Enter your site to see how often ChatGPT, Perplexity, Gemini, and Google cite your brand, and exactly what to publish next.

No credit card required · Free 7-day trial