Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms


Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Photon now handles retrieval and ranking for all production traffic. It also powers a new Fast Search mode in the Perplexity Search API. Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95.

Is it deployable? Yes, as a hosted API. Set search_type: "fast" on POST /search and pay $1 per 1,000 requests. Photon itself is not open source, so the engine cannot be self-hosted.

Why Perplexity Replaced its Old Engine

The old engine hit 3 limits as the index grew:

  • Tail latency: Production p99 sat near 800 ms. The dataset exceeded RAM, so mlock was not an option. Cold reads triggered major page faults that stalled queries.
  • Merge spikes: During disk index fusion, p99 climbed to about 1.2 s for 10 to 15 minutes.
  • Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also raised the share of partial responses.

Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.

How Photon Works

A load balancer routes each request to a Photon broker. The broker fans out to a shard group and watches for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and fetches key document fields.

  • Adaptive posting lists: Short lists sit inline within a single page. Longer lists split into blocks of fixed document ID ranges. Sparse blocks store sorted offset arrays and use galloping search. Dense blocks use bitmaps, so membership becomes a single bit lookup.
  • Budgeted traversal: A WAND-like algorithm splits lists into driving lists and probe lists. Cheap presence checks bound each candidate’s maximum score first. Exact term frequencies are read only when a candidate can clear the threshold.
  • Docblob records: Each document gets a compact record of frequencies, field masks, and positions. Terms use Elias-Fano encoding, so ranking decodes only the matched terms. Ranking a candidate needs just 1 lookup per document.
  • Batched async reads: Record offsets are known upfront, so disk reads go out in batches through io_uring. The cache checks the whole batch first. Readers take no locks, and eviction uses CLOCK instead of a shared LRU list.
  • Separate build and serve: Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes. A controller rotates serving groups one at a time and warms caches with replayed search-log queries.

A full web index now builds in a single-digit number of hours.

Interactive Explainer: Inside Photon

‘;
resize();return;
}
var nb=Math.min(12,Math.max(3,Math.round(Math.log10(n)*2)));
var g=document.createElement(‘div’);g.className=”plist”;
for(var b=0;b=.5;
var el=document.createElement(‘div’);el.className=”blk “+(dense?’dense’:’sparse’);el.style.animationDelay=(b*45)+’ms’;
var h=”

Block “+(b+1)+’: ‘+(dense?’bitmap’:’offset array’)+’

‘;
if(dense){h+=’

‘;for(var k=0;k<32;k++)h+=’‘;h+=’

‘;}
else{var c=Math.max(2,Math.round(bd*12)),o=0,arr=[];for(var k3=0;k3[‘+arr.join(‘,’)+’]



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *