Cohere ships Embed 5 with shared retrieval space across tiers
Cohere released Embed 5, a two-tier embedding family whose Pro variant scored 85.8 on ViDoRe V3 — 10.3 points above OpenAI's text-embedding-3-large. Both tiers share one embedding space.
Updated
Why it matters
- Embed 5 Pro scored 85.8 on ViDoRe V3, 10.3 points above OpenAI's text-embedding-3-large at 75.5 and 2.1 points above Voyage 4 Large at 83.7
- Pro and Fast share one embedding space; Pro-indexed corpus queried with Fast scored 98.4 versus 100 for Pro-plus-Pro across 40 datasets
- Pro costs $0.12 per 1M text tokens, Fast costs $0.08, and image inputs cost $0.40 per 1M tokens on both tiers
- Fast processes 377.3 documents per second versus 159.7 for Pro, per Cohere's measurements
- Most benchmark numbers rely on Cohere's new RCP-nDCG@10 metric, with code published but independent replication still pending
Cohere released Embed 5 on September 30, 2026, a two-tier embedding model family whose Pro variant scored 85.8 on the company's ViDoRe V3 retrieval benchmark — 8.8 points above Embed 4, 2.1 points above Voyage 4 Large, and 10.3 points above OpenAI's text-embedding-3-large.
The launch lands in a market where embedding quality increasingly decides how well enterprise search, retrieval-augmented generation, and agentic systems perform on real workloads. Cohere is positioning the family around a structural choice: the two tiers share one embedding space, letting customers index a corpus with one model and query it with the other.
What did Cohere actually ship?
The family includes Embed 5 Pro, tuned for maximum retrieval quality, and Embed 5 Fast, tuned for latency and cost on the live query path. Both accept text, images, and fused text-plus-image inputs. Both cover 100-plus languages and read up to 128,000 tokens per call. API model IDs are embed-v5.0-pro and embed-v5.0-fast, per Cohere's model docs.
Output dimensions come in 2048, 1536, 1024, 768, 512, or 256, with 2048 as the default. Embeddings return as float, int8, or binary. Pro costs $0.12 per 1 million text tokens. Fast costs $0.08. Image inputs cost $0.40 per 1 million tokens on both tiers.
The image support matters for scanned pages, slide decks, schematics, and charts where text extraction drops information. Embed 5 can embed a page image directly or fuse an image with its metadata into a single vector, the company said.
Why split retrieval into two tiers that share one space?
Cohere tested every index-and-query pairing across 40 development datasets. Normalized to a Pro-plus-Pro baseline of 100, a Pro index queried with Fast scored 98.4. An all-Fast setup scored 96.6. Cohere's recommended pattern is to index with Pro and query with Fast. The only constraint: both sides must use the same output dimension.
The split targets agentic workloads. An agent may issue dozens of searches per task, and query latency compounds across a session. Cohere's team reports Fast processed 377.3 documents per second, versus 159.7 for Pro on the same setup.
The architectural bet is that enterprises want retrieval quality at index time but cheap inference at query time. Sharing one embedding space makes that hybrid configuration work without rebuilding the index when teams upgrade models or rebalance traffic.
How does Embed 5 compare to Voyage 4 Large, Gemini, and OpenAI?
On ViDoRe V3, Embed 5 Pro averaged 85.8 and Embed 5 Fast averaged 84.5, per Cohere. Voyage 4 Large scored 83.7. Gemini Embedding 2 scored 83.2. OpenAI's text-embedding-3-large scored 75.5. On Cohere's parsed-PDF suite, Pro led at 84.8 against Voyage 4 Large at 83.6.
Finance is the strongest showing for Pro. It ranked first on FinanceBench (80.1), FinQA (90.0), and ViDoRe V3 Finance (85.0). Fast ranked second on all three. The pattern suggests Cohere trained heavily on financial documents, a vertical where retrieval precision maps directly to research and compliance outcomes.
Multilingual results are mixed. Pro led the European-language average at 77. However, Gemini Embedding 2 beat Pro on 9 of 10 further languages in Cohere's own results table. Those include Japanese, Arabic, Hindi, and Telugu.
What does the new benchmark metric actually measure?
Most published numbers use RCP-nDCG@10, a metric Cohere introduced alongside the model. It reorders a fixed candidate set, so it measures reranking quality more than first-stage retrieval recall. Cohere published the evaluation code at github.com/cohere-ai/rcp-ndcg.
Independent replication is still pending. Treat Cohere's benchmark numbers as vendor claims until outside teams reproduce them on shared corpora.
How much does storage shrink at scale?
Embed 5 uses Matryoshka representation learning plus lower-precision outputs. A 2048-dim float32 vector needs 8 KB. A 1024-dim int8 vector needs 1 KB. A 256-dim binary vector needs 32 bytes.
Across 100 million chunks, raw storage drops from roughly 819 GB to 3.2 GB when switching from 2048-dim float32 to 256-dim binary. Cohere recommends 1024-dim int8 as the default, citing near-full-precision quality at one-eighth the storage. Binary trades some accuracy and suits a first-pass search before reranking.
For vector databases running on commodity hardware, those bytes matter. Storage often dominates the bill for large enterprise corpora, and the option to drop precision without rebuilding pipelines changes the cost equation.
Where can enterprises run it today?
Both tiers are generally available. Routes include the Cohere API and Model Vault, Microsoft Foundry on Azure, and Amazon SageMaker on the AWS Marketplace. Private VPC or on-prem serving runs through vLLM, the open-source inference engine. Cohere North, the company's retrieval product, also integrates Embed 5 internally.
The breadth of deployment options reflects enterprise demand for hosting choice. Customers in regulated industries usually require VPC or on-prem deployments to keep data inside their own infrastructure. The vLLM path opens that route without locking customers to Cohere's serving stack.
What comes next?
The launch marks Cohere's first embedding release since Embed 4 and tightens competition with Voyage AI, Google, and OpenAI on retrieval benchmarks. The shared-space design is the architectural bet: it requires customers to trust that Pro and Fast will stay in sync as Cohere updates either tier.
Independent replications of RCP-nDCG@10 will shape how seriously enterprise buyers treat the ViDoRe V3 lead over Voyage 4 Large. Cohere's own published code lowers the bar for outside validation. Until outside teams reproduce those numbers on shared corpora, the comparison remains a vendor claim rather than a settled fact.
Original: aws.amazon.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
230 articles
Related articles
- Perplexity's New Embedding Model Retrieves Answers With Their Evidence
- OpenAI Ships o1 to Developers With 60% Cheaper Audio
- OpenAI Ships gpt-image-1 to API After 700M-Image Debut
- OpenAI Adds Remote MCP Support and New Built-In Tools to Responses API
- Google and Kaggle launch FACTS Benchmark Suite for LLM factuality