Cohere Releases Embed 5 With Pro and Fast Tiers for Enterprise Search and RAG
Cohere released Embed 5, a two-tier embedding family for enterprise search, RAG and agentic retrieval, sharing one embedding space.
According to Cohere's model documentation cited by MarkTechPost, the API model IDs are embed-v5.0-pro and embed-v5.0-fast. Both output 2048, 1536, 1024, 768, 512 or 256 dimensions, with 2048 as the default, and return embeddings as float, int8 or binary values. Pro costs $0.12 per 1M text tokens and Fast costs $0.08, while image inputs cost $0.40 per 1M tokens on both tiers. Both models accept text, images and fused text-plus-image inputs, cover more than 100 languages and read up to 128K tokens. MarkTechPost reported that the tiers share one embedding space, so a user can index with one and query with the other.
On ViDoRe V3, Embed 5 Pro averages 85.8, an 8.8-point gain over Embed 4, while Fast averages 84.5, according to MarkTechPost. Voyage 4 Large scores 83.7, Gemini Embedding 2 scores 83.2 and OpenAI text-embedding-3-large scores 75.5 in the same comparison. On Cohere's parsed-PDF suite, Pro leads at 84.8 against Voyage 4 Large at 83.6. Finance is the strongest area for Pro: it ranks first on FinanceBench at 80.1, FinQA at 90.0 and ViDoRe V3 Finance at 85.0, with Fast ranking second on all three. Multilingual results are mixed. Pro leads the European-language average at 77, but Gemini Embedding 2 beats Pro on nine of 10 further languages in Cohere's own results table, including Japanese, Arabic, Hindi and Telugu. MarkTechPost noted that most of these numbers use RCP-nDCG@10, a new Cohere metric that reorders a fixed candidate set, measuring reranking quality more than first-stage retrieval; Cohere published the evaluation code, but independent replication is still pending.
Cohere tested every corpus and query pairing across 40 development datasets. Normalized to Pro plus Pro at 100, a Pro index queried with Fast scored 98.4, and an all-Fast setup scored 96.6. Cohere recommends indexing with Pro and querying with Fast, provided both sides use the same output dimension. The split is aimed at agentic workloads, in which an agent may issue dozens of searches per task and query latency compounds. Cohere's team reported that Fast processed 377.3 documents per second, compared with 159.7 for Pro.
Embed 5 uses Matryoshka representation learning and lower-precision outputs to reduce storage costs. A 2048-dimension float32 vector needs 8 KB, a 1024-dimension int8 vector needs 1 KB and a 256-dimension binary vector needs 32 bytes. Across 100 million chunks, raw storage drops from about 819 GB to 3.2 GB, according to MarkTechPost. Cohere recommends 1024-dimension int8 as the default, citing near-full-precision quality.