Perplexity Releases pplx-embed-v2-late Embedding Models in 0.6B and 9B Sizes
Perplexity released pplx-embed-v2-late, ColBERT-style multimodal embedding models in 0.6B and 9B sizes, with MIT-licensed weights on Hugging Face. The 9B model scored 92.4% on MADQA, while the 0.6B model is designed for edge query encoding.
Perplexity describes the 0.6B model as a lightweight query encoder that can run on a laptop, edge device or small GPU. It has about 240 million active text parameters and about 340 million active image parameters, and Perplexity says it stays close to 8B rivals. The 9B model has 7.4 billion active parameters and is listed as 8B on Hugging Face; Perplexity positions it for datacenter or high-memory GPU use, mainly for building the document index. Estimated weight memory is about 1.2 GB for the 0.6B model in bf16 and about 16 GB to 18 GB for the 9B model in bf16. The published checkpoints are stored in F32, which doubles the download size. The 0.6B model uses Qwen3.5-0.8B pruned to 12 text layers, while the 9B model uses Qwen3.5. Both output 128 dimensions per token. They require sentence-transformers 6.0.0 or newer and transformers 5.4.0 or newer, and the model cards show CUDA GPU usage.
According to Perplexity's announcement, the 9B model scored 92.4% on MADQA, the best result cited. The 0.6B model scored 90.1%, beating Mixedbread's retriever at 88.9% but trailing Mixedbread Agentic Search at 93.4%. On domain-specific text across 72 tasks measured by nDCG@10, the 0.6B model scored 78.0% and the 9B model scored 81.3%; Perplexity says the 9B model leads all tested models by 1.6 percentage points, while the 0.6B model is 0.3 points behind gemini-embedding-2. On Q2D-Web Recall@1000, the models scored 73.6% and 74.8%, both above the previous best of 69.3%. On ViDoRe v3 image retrieval nDCG@10, they scored 62.3% and 65.2%; the 0.6B model is within 1.2 points of nemotron-colembed-v2-8b, but Tencent's EVIE scores higher. On ViDoRe v3 Markdown nDCG@10, they scored 61.2% and 64.7%, with the 0.6B model's score the weakest result Perplexity reported even though it was the second-best score on that benchmark. On BrowseComp+, the 9B model scored 64.0%, the lowest figure Perplexity listed for that model but still 4.9 points above the next ColBERT model; no 0.6B score was given. Perplexity says the biggest margin is on BrowseComp+, at 8.7 points over the best dense model. It also says image search is the real gap: Gemini Embedding 2 beats the 9B model on MIRACL-Vision and by 2 points on PPLX-Q2I.
Perplexity highlights mixed-size use: a 9B index queried by the 0.6B model scored 63.5% on ViDoRe v3 image retrieval, above the 62.3% scored when 0.6B is used on both sides, at the same query cost. For text, a 9B index searched with 0.6B queries recovers about half the 9B quality gap at 0.6B query cost. The models' 128-dim token vectors are 16 to 32 times narrower than rivals at 2,048 to 4,096 dimensions.
The models work differently from dense embedding models, which compress a document into one vector. pplx-embed-v2-late keeps a 128-dim vector for every token and scores with MaxSim: each query token finds its best document token and those maxima are summed. Pages are encoded as images, so no OCR step is needed. Perplexity distilled both models from an 18B teacher using LEAF-style token-level training, which it says creates the shared space. Perplexity suggests the models for visual document search over PDFs, slides and scanned reports; low-latency search where the index is built in the cloud with the 9B model and queries run on-device with the 0.6B model; and agentic RAG over large PDF or web collections.
Perplexity lists several caveats. The models store one vector per token, so index size grows with document length. The models are not No. 1 on ViDoRe v3 image retrieval, where Tencent's EVIE scores higher. A single input cannot mix text and images. All scores are self-reported, and the technical report is not out yet. The weights are MIT-licensed with commercial use allowed, and a hosted Perplexity API endpoint is planned but not live.
Editor's Summary
Perplexity's pplx-embed-v2-late adds two self-hosted multimodal embedding models, with the 0.6B version aimed at edge query encoding and the 9B version at high-quality document indexing. Perplexity reported a top MADQA score of 92.4% for the 9B model, but the image retrieval benchmark still trails Tencent's EVIE and the scores are self-reported. The MIT-licensed weights are on Hugging Face, while a hosted API is not yet live.