Models

Perplexity's New Embedding Model Retrieves Answers With Their Evidence

Perplexity and turbopuffer released pplx-embed-v2-context-9b-preview, a contextual embedding model that retrieves answers plus their supporting evidence, trained via distillation instead of gold-passage labels.

By Marcus Bennett4 min read

Updated

Why it matters

  • pplx-embed-v2-context-9b-preview is available on Hugging Face under the MIT license, requiring transformers>=5.4.0 with trust_remote_code=True; it is not yet on the Perplexity API.
  • Training distills from Perplexity's query-aware context compression teacher using forward KL divergence and InfoNCE document loss, replacing single gold-passage labels with soft targets.
  • Perplexity reports beating voyage-context-4 by 14.4 and 5.0 points on context-bench (2,099 queries, 38,894 documents) at K=10, with 1024-dim int8 embeddings (1 KB/vector) slightly exceeding voyage-context-4's 2048-dim float32 (8 KB/vector).

Perplexity Research, working with turbopuffer, has released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines that trains on a fundamentally different signal: instead of learning to retrieve one "gold passage," it learns to retrieve the answer together with the context needed to verify it.

The weights are available now. You can self-host the preview from Hugging Face under the MIT license. Loading requires transformers>=5.4.0 with trust_remote_code=True. The model is not yet on the Perplexity API, and the model card warns that weights and the interface may change without backward compatibility.

Why the gold passage falls short

RAG systems split long documents into chunks. A chunk often depends on an entity, a heading, or a definition stated elsewhere in the document. Contextual embedding models address part of this problem with late chunking, the technique from a 2024 paper (arXiv:2409.04701): the document is encoded in one pass, then pooled per chunk.

Training remains the bottleneck. Standard pipelines mark one gold chunk per query. Every other chunk becomes a negative — including the sentences that make the answer checkable. Perplexity identifies three additional problems with this setup. Binary labels give a coarse signal. LLM annotation cost grows linearly with dataset size. And labels are tied to one chunking strategy, so changing how you split documents invalidates the annotation.

The stakes are practical. Retrieval quality caps the quality of every downstream RAG answer, and re-annotating training data for each chunking scheme does not scale. Perplexity's answer is to stop hand-labeling chunks altogether.

How the training works

The teacher is Perplexity's query-aware context compression model, previously described in a separate blog post. It reads the query and the document together and scores every token.

From those token scores, the training pipeline builds four components:

  • Chunk relevance: the mean of the top n token scores inside each chunk.
  • Soft target: a temperature-scaled softmax over chunks in the positive document. Chunks in other documents get zero.
  • Distillation loss: forward KL divergence between teacher and student distributions.
  • Document loss: InfoNCE, where a document scores as its best chunk — an approach inspired by ColBERT's MaxSim (arXiv:2004.12832).

Each training batch samples a random chunking strategy. Chunks are separated by a learned <|chunk_sep|> token and mean-pooled. Because the teacher runs only during training, inference adds no latency and no storage overhead.

The architecture itself starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048-dimensional embeddings. Matryoshka training (arXiv:2205.13147) also supports 1024 dimensions, and quantization-aware training enables native int8 embeddings. The released model is a soup of several checkpoints. Training used roughly 430 datasets covering more than 50 languages, with no ConTEB data included.

Reported results

Perplexity evaluated the model on context-bench, which comprises 2,099 queries, 38,894 documents, and 2,458,072 sentence chunks, with exhaustive ranking. At K = 10, Perplexity reports the preview beating voyage-context-4 by 14.4 points on one metric and 5.0 points on another (Voyage's baseline values can be derived from those stated gaps). Other Voyage metrics appear only in Perplexity's chart.

The storage math matters as much as the leaderboard. Contextual embeddings store one vector per chunk, the same as a conventional chunk index, so cost depends on vector size. Perplexity reports that a 1024-dimension int8 configuration — 1 KB per vector — slightly exceeds voyage-context-4 at 2048-dimension float32, which costs 8 KB per vector, on its chunk-retrieval suite. At context-bench scale, that is roughly 2.4 GB versus 19.6 GB for vectors alone, before index overhead.

The model is also robust to chunk size. Across chunk sizes from 64 to 512 tokens, mean nDCG@10 across 74 MTEB tasks varies only from 81.0% to 79.9%, as reported by Perplexity.

What to watch

The "preview" label is doing real work here: weights and interfaces may change without backward compatibility, and API availability has not been announced. But the training recipe — teacher distillation from token-level scores, random chunking per batch, no manual annotation — is a concrete attempt to break the link between retrieval training data and one fixed chunking strategy. If the approach holds up under independent evaluation, the cost of building and maintaining embedding models for production RAG could drop substantially.

Original: perplexity.ai

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

132 articles

Related articles

  1. Google Releases DiffusionGemma, a 26B Model That Generates Text Four Times Faster
  2. Google Converts Gemma 2 Into Encoder-Decoder Models With T5Gemma
  3. OpenAI Ships o1 to Developers With 60% Cheaper Audio
  4. OpenAI Launches Model Distillation Suite in Its API
  5. OpenAI's GPT-5.1-Codex-Max runs coding tasks for 24 hours straight

« Previous article