AI News Feed
Market watch
Products & Applications

Perplexity Releases Photon, a Rust Retrieval Engine That Cuts p99 Latency to 65 ms

Perplexity has released Photon, an in-house Rust retrieval and ranking engine that now handles all production search traffic and powers a new Fast Search mode. Perplexity reports p99 retrieval and ranking latency fell from about 800 ms to about 65 ms, while Fast Search costs about 68% less than the default preset but scores slightly lower.

Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95. The company says p99 retrieval and ranking latency fell from about 800 ms to about 65 ms, covering Photon's stages only. Photon is available as a hosted API; developers set search_type: 'fast' on POST /search and pay $1 per 1,000 requests. Photon itself is not open source, so it cannot be self-hosted.

The company replaced its old engine because the system hit three limits as the index grew. Production p99 latency sat near 800 ms. The dataset exceeded RAM, so mlock was not an option, and cold reads triggered major page faults that stalled queries. During disk index fusion, p99 climbed to about 1.2 seconds for 10 to 15 minutes. Deploying and syncing an extra cluster could take more than a week, and recovery raised the share of partial responses. Perplexity concluded that building from scratch was simpler and cheaper than maintaining its fork.

Photon's architecture starts with a load balancer that routes each request to a Photon broker. The broker fans out to a shard group and watches for timeouts. Each shard runs retrieval, initial ranking and second-stage ranking. The broker then merges candidates and fetches key document fields. Adaptive posting lists keep short lists inline within a single page. Longer lists split into blocks of fixed document ID ranges. Sparse blocks store sorted offset arrays and use galloping search; dense blocks use bitmaps, so membership becomes a single bit lookup.

A WAND-like budgeted traversal algorithm splits lists into driving lists and probe lists. Cheap presence checks bound each candidate's maximum score first, and exact term frequencies are read only when a candidate can clear the threshold. Each document gets a compact docblob record of frequencies, field masks and positions. Terms use Elias-Fano encoding, so ranking decodes only the matched terms, and ranking a candidate needs just one lookup per document.

Record offsets are known upfront, so disk reads go out in batches through io_uring. The cache checks the whole batch first. Readers take no locks, and eviction uses CLOCK instead of a shared LRU list. Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes, separating build and serve. A controller rotates serving groups one at a time and warms caches with replayed search-log queries. A full web index now builds in a single-digit number of hours.

In production, Perplexity says Photon runs on about 20% fewer serving machines than the old content nodes. It stores about 2.5 times as much data per document, which Perplexity used to improve ranking quality. Pinning the same dataset with mlock would need an estimated 4.6 times the resident memory Photon uses today. Index version switches no longer cause latency spikes.

Fast Search pairs Photon with lighter ranking tuned for agentic workflows. Perplexity tested it on six benchmarks: WideSearch, BrowseComp, DSQA, FRAMES, SEAL-0 and SEAL-Hard. Across 3,554 tasks, Fast scored 64.3% at $59.73 in estimated model-plus-search cost. The default preset scored 64.0% at $187.60, making Fast about 68% cheaper. The trade-off appears in broader search quality. On internal long-tail benchmarks, relevance measured by DCG fell from 2.45 to 2.21, and answer availability dropped from 0.596 to 0.567, a loss of 2.9 percentage points. Perplexity recommends Fast for day-to-day agent loops and the default for hard, ambiguous queries.

The API can be called with a query, search_type set to fast and max_results, among other parameters. On Python SDK 0.43.4 and 0.43.5, users pass extra_body with search_type fast. Perplexity lists Fast Search against Exa Instant, Parallel Search Turbo and Tavily ultra-fast. The request parameters are search_type: 'fast', type: 'instant', mode: 'turbo' and search_depth: 'ultra-fast', respectively. Perplexity reports 160 ms p50 and 230 ms p95 latency; Exa lists about 250 ms typical and sub-200 ms at launch; Parallel lists about 200 ms; Tavily publishes no figure and describes ultra-fast as its lowest-latency depth. List price per 1,000 requests is $1 for Perplexity Fast Search, $4 for up to 10 results from Exa, $1 for Parallel Search Turbo, and one credit for Tavily ultra-fast, which is $8 pay-as-you-go and $5 to $7.50 on plans. Results per request are 1 to 20 for Perplexity, 10 in Exa's base price with $1 per 1,000 per extra result, and not specified for Parallel or Tavily. Perplexity notes lower relevance than its default preset; Exa bills extra results separately; Parallel limits English and Japanese queries; Tavily says lower relevance than other depths. Launch dates listed are Sep 24, 2026 for Perplexity Fast Search, Feb 12, 2026 for Exa Instant, Jul 13, 2026 for Parallel Search Turbo and Jan 5, 2026 for Tavily ultra-fast.