Perplexity Releases pplx-embed-v2-context-9b-preview for RAG Retrieval and Evidence
Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a self-hosted contextual embedding model for RAG pipelines that is trained to retrieve answers and their supporting evidence.
Deployment details are narrow. Loading requires transformers>=5.4.0 with trust_remote_code=True, and the model is not yet on the Perplexity API. The model card warns that weights and interface may change without backward compatibility.
The training signal is the central change. RAG systems split long documents into chunks, and a chunk often depends on an entity, heading, or definition stated elsewhere. Contextual models address that with late chunking: the document is encoded in one pass, then pooled per chunk. Training, however, usually marks one gold chunk per query. Every other chunk becomes a negative, including the sentences that make the answer checkable. Perplexity lists three further problems: binary labels give a coarse signal, LLM annotation cost grows linearly with dataset size, and labels are tied to one chunking strategy.
To replace the single gold passage, Perplexity uses a query-aware context compression model as teacher. The teacher reads the query and document together and scores every token. Chunk relevance is the mean of the top n token scores inside each chunk. A temperature-scaled softmax over chunks in the positive document produces a soft target; chunks in other documents get zero. The student is trained with a forward KL divergence distillation loss. A document loss uses InfoNCE, where a document scores as its best chunk, inspired by ColBERT’s MaxSim. Each batch samples a random chunking strategy. Chunks are separated by a learned <|chunk_sep|> token and mean-pooled. The teacher runs only during training, so inference adds no latency or storage, Perplexity says.
The model starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048 dimensions, and Matryoshka training also supports 1024 dimensions. Quantization-aware training enables native int8 embeddings. The release is a model soup of several checkpoints. Training used roughly 430 datasets covering more than 50 languages, with no ConTEB data.
Perplexity’s interactive explainer uses three lease files that share the sentence “Monthly rent is …”; only one belongs to 5 Park Avenue. The query asks when 5 Park Avenue’s lease ends and what the current rent is. The explainer contrasts isolated sentences with contextual late chunking and shows that the same token scores can re-aggregate when chunk boundaries change. The scores in the example are for explanation only and are not model outputs.
On context-bench, which contains 2,099 queries, 38,894 documents and 2,458,072 sentence chunks with exhaustive ranking, the reported numbers use K=10. The MarkTechPost explainer derived Voyage-context-4 values by subtracting Perplexity’s stated gaps of 14.4 and 5.0 points from the reported Perplexity figures. Other Voyage metrics appeared only in Perplexity’s chart. Perplexity also reports that 1024-dimensional int8 embeddings, at 1 KB per vector, slightly exceed voyage-context-4 at 2048-dimensional float32, at 8 KB per vector, on its chunk-retrieval suite. Chunk-size sensitivity from 64 to 512 tokens ranged from 81.0% to 79.9%, measured as mean nDCG@10 across 74 MTEB tasks, according to Perplexity.
Contextual embeddings store one vector per chunk, the same as a normal chunk index, so storage cost depends on vector size. Perplexity presents the release as a self-hosted preview rather than an API product for now.