Perplexity always crawls live. Here is what that means for citations.
Perplexity's live-fetch model rewards recency and clean structure over domain authority, reshaping the citation calculus for institutional publishers.
Key takeaways
- Perplexity runs a live web crawl for every query, so no citation pool is fixed in advance.
- Pages published or updated within 72 hours appear in citations at disproportionately high rates.
- Structured content with clear headings and discrete sections is extracted more often than flowing narrative.
- Video transcripts can influence Perplexity answers even when the video never appears as a visible citation.
- Slow page load times and poor heading hierarchy cost institutional publishers citations regardless of content quality.
Perplexity does not maintain a static index. Every query triggers a live crawl, which means every query is, in principle, winnable by any source that exists at the moment the question is asked. That single architectural fact separates Perplexity from Google AI Overviews and ChatGPT's retrieval layer, and Search Engine Journal's analysis of the platform's raw answer stream makes clear how consequential the difference is.
The investigation, by Suganthan Mohanadasan for Search Engine Journal, goes one level below the finished answer. Rather than inspecting which sources appear in the citation panel, it reads the intermediate stream Perplexity generates before composing its response. That stream shows which URLs the system fetches, in what order, and how it weights them. The findings are specific enough to act on.
What the stream shows
Perplexity fetches between three and ten sources per query, almost always including at least one from the top five organic Google rankings for the equivalent search. Domain authority is not irrelevant, but it is not determinative either. Freshness exerts a stronger gravitational pull than most practitioners assume: pages updated or published within the preceding 72 hours appear with disproportionate frequency, particularly for queries with any news or market-movement dimension. For a B2B brand publishing a quarterly policy brief or an annual sustainability report, the publication timing relative to when the topic is queried matters as much as the brief's intrinsic quality.
The stream also reveals that structured content, numbered lists, clear section headers, and definition-style prose, is pulled into citations at a higher rate than flowing narrative. Perplexity appears to retrieve chunks it can quote or paraphrase cleanly. A page organised as a single unbroken argument is harder for the system to excerpt than one with discrete, labelled sections. This is not new as a general observation about LLM retrieval; the Search Engine Journal analysis is notable for showing it operating specifically inside Perplexity's live-fetch architecture, where the model is deciding in real time which fragments to carry forward.
Video and local results surface in the stream in a way that finished answers conceal. When a query has geographic or instructional intent, Perplexity fetches YouTube transcripts and local business data even if the final rendered answer does not surface them as visible citations. For industrial groups and multilateral institutions whose content includes technical how-to material, this is relevant: a video with a good transcript can influence Perplexity's answer without the institution ever appearing in the citation panel.
The implication for brands that get cited on AI platforms
For financial services firms, multilaterals, and large industrial groups, the live-crawl model creates an opening that a cached-index system forecloses. A report published by CGAP on financial inclusion or a policy brief from UNDRR on disaster-risk metrics can, in principle, enter Perplexity's citation pool the day it is published, provided it is indexed and structured correctly. The constraint is not waiting for a model update or a ranking cycle. The constraint is whether the page is crawlable, whether it loads fast enough for a live fetch under query conditions, and whether its text is chunked in a way that makes clean extraction possible.
That last requirement is where most institutional publishers fail quietly. Organisations in the UN system and major foundations tend to publish long PDFs or dense single-scroll web pages that are technically accessible but structurally opaque to a retrieval model trying to extract a two-sentence answer about a specific sub-question. The competitive disadvantage is not one of authority or credibility; it is one of format.
Perplexity's architecture also penalises inconsistency in a way that slower-moving systems do not. A brand that publishes a definitive piece and then lets the page go stale, no updates, no internal links from newer content, no signals of continued editorial attention, will find that live-crawl recency weighting gradually routes around it, even if the underlying information remains accurate. Freshness is not just a news-site concern.
The practical consequence is that Perplexity citation strategy requires closer alignment between content and technical teams than most large institutions currently have. A policy team can produce an excellent brief; if the CMS renders it without clear heading hierarchy and the server response time exceeds two seconds under load, the brief loses to a thinner source that is faster and better structured. That trade-off is happening every time someone asks Perplexity a question your organisation should own.