top of page

Perplexity pplx-embed-v2-late Splits Multimodal Search Between 9B Indexing and 0.6B Queries

13 hours ago
13 min read

Perplexity released two pplx-embed-v2-late models with a notable split: a 9B model builds richer indexes, while a 0.6B model handles faster queries. Both models search text, images, and rendered document pages in the same embedding space.

That pairing matters more than the parameter counts. Retrieval teams usually choose one embedding model and accept its quality, latency, and infrastructure costs everywhere. Perplexity instead proposes spending more compute when documents enter the index, then using a smaller encoder on the request path.

The models also challenge the standard pipeline for searching visually complex PDFs. Instead of extracting text through OCR, a developer can encode a rendered page and retrieve it with a text query. However, the approach replaces some parsing costs with larger indexes and more expensive scoring.

Perplexity Released Two Models That Work as One Retrieval System

The central release is not simply a pair of checkpoints. It is an asymmetric retrieval design built around a shared embedding space.

Perplexity published 0.6B and 9B versions of pplx-embed-v2-late under an MIT license. The weights are available through separate Hugging Face model repositories, including the 0.6B model and its larger 9B counterpart.

Both models are multimodal late-interaction retrievers. Late interaction means that documents and queries are encoded separately, but their individual token vectors interact during scoring. This differs from dense retrieval, which usually reduces each input to one vector.

Each model outputs one 128-dimensional vector per token. MaxSim scoring then finds the strongest document-token match for every query token and sums those maximum similarities. Different query terms can therefore match different regions of a page.

Perplexity built the family on Qwen3.5 backbones with bidirectional attention. According to the model card, both released models were distilled from an internal 18B ColBERT teacher. ColBERT is a retrieval architecture that preserves token-level representations for late comparison.

The company fully fine-tuned the smaller model. For the 9B version, it fully tuned the final eight transformer layers while adapting the remaining layers and vision encoder with LoRA.

The advertised sizes also require context. The smaller checkpoint contains about 594 million total parameters, but Perplexity reports 340 million active parameters. Text encoding activates roughly 240 million parameters, while image encoding activates about 340 million.

Perplexity produced the smaller text tower by pruning a 24-layer Qwen3.5-0.8B tower to 12 layers. Its token-embedding table accounts for another 254 million parameters, according to the company.

That construction targets query-time economics. Indexing can run offline, across parallel hardware, and only when documents change. Query encoding sits on the live request path, where every additional millisecond affects the user experience.

The shared space connects those two workloads. A company can encode documents with the 9B model, then search the resulting index with queries encoded by the 0.6B model. Replacing the query encoder does not require rebuilding that index.

Perplexity also describes a local-cloud configuration. A device could encode a private query or local document with the smaller model, then compare that representation with results from a cloud-hosted 9B index.

This flexibility remains a technical proposal rather than a managed product promise. The model card says the checkpoints work with recent Sentence Transformers and Transformers releases. It also says no inference provider currently serves the smaller checkpoint.

Perplexity says late-interaction, dense, and contextual embeddings will reach its API platform progressively. Until that happens, teams evaluating pplx-embed-v2-late should assume they must operate the models and retrieval infrastructure themselves.

The Perplexity pplx-embed-v2-late Mechanism Preserves Page Detail

Perplexity is betting that token-level matching can retain evidence that a single document vector often compresses away.

A dense embedding model represents a query and document with one vector each. Retrieval becomes an efficient nearest-neighbor search, which works well across very large collections. Yet the vector must summarize every potentially relevant detail.

That compression becomes harder as documents grow longer or contain unrelated sections. It becomes harder again when pages include charts, tables, diagrams, captions, and layout-dependent meaning. A single representation has limited room for all those signals.

Chunking reduces the amount of information placed into each vector. However, it can separate a table from its label, a chart from its legend, or a clause from an important qualification. Parsing rules also vary across document formats.

A cross-encoder addresses part of the problem by processing a query and candidate together. That joint attention supports detailed comparisons, but the model must run again for every query-candidate pair. It is usually practical only for reranking a short candidate list.

Late interaction occupies the middle ground. Documents still receive their representations before a query arrives. The retrieval system then performs multiple token-level comparisons instead of calculating one inner product per candidate.

Perplexity’s technical explanation illustrates the method with MaxSim. Each query token selects its best document-token match, and the system adds those similarities into a document score.

Consider a query about a statute, deadline, and exception. A single query vector blends those concepts. MaxSim can match each concept against a separate passage, label, or visual region within the same page.

The same mechanism applies to images. A rendered PDF page enters the vision encoder as an image instead of passing through OCR first. A text query can then retrieve the visual representation directly.

This does not mean the model “reads” a PDF file without preparation. The application must render each relevant page into an image and encode that image. The distinction concerns how the searchable representation is produced.

Skipping OCR can preserve layout and visual relationships that text extraction loses. A financial table’s row and column positions may carry essential meaning. A diagram may communicate relationships that its caption only partially describes.

OCR-free retrieval can also avoid recognition errors in scans, unusual fonts, and complicated page structures. However, it does not automatically provide extracted text for highlighting, quotation, access controls, or downstream language-model context.

Many applications will therefore retain parsing alongside visual retrieval. Visual embeddings can identify a promising page, while OCR or native PDF text supplies exact passages afterward. The techniques can complement each other.

Perplexity trained the two models on 186 million query-document pairs drawn from 594 datasets and 46 languages. It reports that 88.3% were text-to-text pairs, 8.3% text-to-image, and 3.4% text-to-visual-document.

The sampling mixture increased the relative presence of visual data. Perplexity says its final sampling weights produced 56.5% text-to-text, 30.9% text-to-image, and 12.6% text-to-visual-document examples.

Those details matter because “multimodal” covers several different problems. Retrieving a photograph is not identical to finding evidence inside a dense annual-report page. Training balance influences which use cases receive the strongest representation.

The released model card also specifies an implementation constraint. Text-only and image-only items require separate encoding calls, and mixed text-plus-image inputs are not supported in one item. Applications must design ingestion accordingly.

Shared Embeddings Put Pressure on Single-Model Retrieval Pipelines

The competitive pressure falls on retrieval systems that use one encoder size for both offline indexing and latency-sensitive queries.

Most embedding deployments treat the model as a uniform component. The same checkpoint embeds a corpus and every incoming query. That symmetry simplifies operations, but it ignores the different economics of those jobs.

Document encoding is usually an amortized expense. A company might process a page once, then answer thousands of searches against its stored representation. It can schedule indexing on larger hardware or batch the work.

Query encoding repeats for every search. It affects response time, concurrency, and device feasibility. Running a large vision-language encoder for each request can erase gains obtained during offline indexing.

Perplexity’s shared space separates those choices. The 9B model can spend additional compute capturing document information, while the 0.6B model produces compatible queries. The index retains some benefit from the larger document encoder.

In Perplexity’s evaluation across 72 domain-specific retrieval tasks, the asymmetric setup gained an average of 1.6 percentage points over using 0.6B on both sides. The query encoder remained unchanged.

For ViDoRe v3 image retrieval, the 0.6B query and 9B document configuration scored 63.5% nDCG@10. The symmetric 0.6B configuration scored 62.3%, a difference of 1.2 points.

Using the 9B model for both queries and documents still produced the strongest reported domain average, at 81.3%. Perplexity says the asymmetric configuration recovered roughly half the text-quality gap without larger query-time encoding.

That is the release’s most practical argument. The smaller model does not need to match every 9B result by itself. It only needs to make a high-quality 9B index useful under tighter serving constraints.

The approach pressures standard dense models, but it also competes with other multi-vector retrievers. Perplexity compares its models with Qwen3-VL-Embedding, EVIE, TopK Embed, and Nvidia’s Nemotron ColEmbed family.

On the public image portion of ViDoRe v3, Perplexity reports 65.2% nDCG@10 for the 9B model and 62.3% for the 0.6B model. The corresponding markdown scores were 64.7% and 61.2%.

Perplexity says the 0.6B model came within 1.2 points of Nemotron ColEmbed V2 8B on image retrieval. It also emphasizes that its outputs use 128 dimensions per token.

That dimension comparison speaks directly to index feasibility. Perplexity lists output dimensions of 2,048 for EVIE-4.5B and 4,096 for larger EVIE and Nemotron models. Fewer dimensions can reduce each stored token vector.

Dimensions alone do not determine production cost. The number of retained tokens, numeric precision, compression method, index structure, and candidate-generation strategy also matter. Perplexity has not published a complete storage calculation for representative corpora.

The alternative is not disappearing. Dense retrieval remains easier to index and search at enormous scale. Cross-encoders remain attractive for reranking. Hybrid lexical search still protects exact identifiers, names, and rare technical terms.

Pplx-embed-v2-late is therefore more likely to become one stage in a retrieval stack than a universal replacement. Perplexity itself describes late interaction as either richer first-stage retrieval or a later stage in web-scale systems.

For teams building a searchable knowledge base, the design question becomes more specific. They must decide which documents justify visual, multi-vector indexing and which remain efficient with text retrieval.

The Benchmark Results Are Strong, but They Remain Company-Reported

The published scores support serious testing, yet they do not settle real-world latency, storage, or retrieval quality.

Perplexity reports a 64.0% answer accuracy score for the 9B model on BrowseComp+. That result exceeded the next ColBERT model by 4.9 percentage points and the next dense model by 8.7 points.

BrowseComp+ uses a fixed corpus instead of live web search. Its benchmark design includes 830 difficult queries and roughly 100,000 curated web documents with human-verified supporting evidence.

A fixed collection improves reproducibility. Researchers can separate retrieval quality from changes in commercial search engines or the open web. It also makes the benchmark narrower than operating a live, continuously changing web index.

Perplexity paired its retriever with GPT-OSS-120B at high effort. Another language model judged whether the generated answer matched the reference. The reported 64.0% therefore measures an agent-retriever system, not an isolated embedding score.

The company says the 0.6B model also exceeded every model outside the pplx-embed-v2-late family. However, the announcement does not provide every underlying score as searchable text. The complete technical report is scheduled for later publication.

On MADQA, Perplexity reports 92.4% answer accuracy for its 9B retriever and 90.1% for the 0.6B model. Both were paired with Gemini 3.5 Flash.

MADQA evaluates agentic search across heterogeneous PDFs. The underlying MADQA paper describes 2,250 human-authored questions grounded in 800 documents, while the reported evaluation uses a 500-question subset.

Perplexity says that subset spans more than 18,000 pages. The questions cannot be answered from general knowledge, so the agent must retrieve evidence from the document collection.

The 9B result exceeded a standard Mixedbread retriever with the same agent by 3.5 points. Mixedbread Agentic Search reached 93.4%, which Perplexity says fell within its confidence interval.

That distinction is important. Mixedbread Agentic Search includes a search sub-agent that can plan and execute several searches per outer call. Pplx-embed-v2-late functions as a retriever inside the agent rather than an entire agentic search service.

MADQA’s authors also identify a wider limitation in document agents. Their study found that strong systems can approach human accuracy while succeeding on different questions. Agents often compensate for weak strategy through repeated searches.

A better retriever can reduce that waste, but it cannot guarantee good search planning or evidence synthesis. Retrieval accuracy, answer accuracy, page-level evidence quality, latency, and tool-call count should all be evaluated separately.

Perplexity also reports strong ViDoRe v3 results. That public benchmark gives developers a more direct view of page retrieval than answer-generation tests. Still, benchmark collections cannot reproduce every enterprise document format.

Real corpora contain duplicates, access restrictions, revisions, handwritten annotations, low-resolution scans, and pages with nearly identical layouts. They also contain domain-specific abbreviations that may be absent from general training mixtures.

The internal PPLX-Q2I benchmark introduces another verification gap. Perplexity built it from production image-search logs and evaluated 10,000 queries against 100,000 images. Outside researchers cannot yet reproduce that private test.

Perplexity says both models beat Qwen3-VL-Embedding-8B by more than nine points on PPLX-Q2I. It also says the 9B model trailed Gemini Embedding 2 by about two points. Those findings should remain company-attributed.

The release deserves attention because the weights and public benchmark checkpoints permit independent testing. It does not deserve automatic acceptance as the best option for every corpus.

Developers should construct an evaluation set from their own documents and real queries. It should include exact lookup, cross-page evidence, visual tables, obscure terminology, and intentionally difficult negatives.

They should also compare equal end-to-end systems. One setup should not receive better OCR, more search rounds, or a stronger reranker unless those differences represent the intended production design.

Multi-Vector Search Shifts Costs Rather Than Removing Them

Pplx-embed-v2-late avoids one compression bottleneck by accepting larger representations and more involved candidate scoring.

Dense retrieval stores one vector for each document or chunk. Late interaction retains multiple vectors, often one for every unpruned token. A long page can therefore generate many searchable representations.

Even at 128 dimensions, those vectors accumulate. Index storage depends on the token count, numeric format, compression scheme, metadata, and retrieval engine. Page-level visual representations can further change the calculation.

MaxSim also requires more work than one query-document inner product. Each query token must locate its strongest match among document tokens. Efficient serving needs specialized indexing, pruning, or staged retrieval.

Perplexity acknowledges this tradeoff in its announcement. It says late interaction requires different indexing and serving choices from single-vector approximate nearest-neighbor retrieval. Longer documents increase the cost.

The shared model space helps with query encoding, but it does not erase candidate-scoring costs. A lightweight query encoder can still produce a request that is expensive to match against millions of token vectors.

Teams should therefore measure four separate latency components: query encoding, candidate generation, MaxSim scoring, and downstream reranking or generation. Reporting only model inference time hides much of the user experience.

Memory use also needs careful measurement. The 9B checkpoint may be an offline component, but indexing a frequently changing corpus can still demand persistent GPU capacity. Re-encoding document revisions adds operational work.

A visual-document pipeline requires page rendering before model inference. Large PDFs need pagination, image normalization, failure handling, metadata mapping, and deletion workflows. OCR may disappear from retrieval, but ingestion remains a system problem.

OCR also retains advantages. Extracted text supports keyword search, highlighting, citations, compliance review, and direct language-model context. A rendered-page embedding cannot reproduce those functions on its own.

The likely production design is hybrid. A system can index native text for exact retrieval, keep visual embeddings for layout-sensitive pages, and use a reranker on a bounded candidate set.

Access control deserves equal attention. Retrieval indexes must filter unauthorized material before results reach an agent. A shared local-cloud embedding space does not provide document permissions or privacy guarantees automatically.

The proposed on-device query path raises further questions. Perplexity calls the 0.6B model suitable for edge devices, but device classes vary widely. Memory limits, acceleration support, quantization, and battery use will shape actual feasibility.

The published checkpoint uses F32 tensors on its Hugging Face page. Developers will likely test lower-precision or platform-specific variants, but those conversions require quality checks. Quantization can alter retrieval rankings.

Compatibility is another early-stage concern. The model card requires Sentence Transformers 6.0 or newer and Transformers 5.4 or newer. It also warns that PyLate inserts query and document markers in a different position.

The export uses native Sentence Transformers modules and does not require custom Python code. That lowers integration friction, but it does not provide a complete production index or managed endpoint.

Licensing is comparatively straightforward. The MIT license permits broad use and modification. Still, adopters must review model dependencies, training-data considerations, and their own handling of sensitive documents.

“Open source” can also obscure important distinctions. Perplexity released open weights and implementation instructions, but it has not released the full training corpus or internal 18B teacher.

The company says it excluded datasets related to evaluated benchmarks from training. That is a useful methodological claim, but independent researchers need the promised technical report to examine contamination controls and evaluation details.

The prudent conclusion is not that late interaction costs too much. It is that the expense moves. Teams trade OCR dependence and single-vector compression for richer indexes, token-level scoring, and more specialized infrastructure.

Three Signals Will Show Whether the Design Travels Beyond Benchmarks

The next test is whether independent deployments can reproduce the quality gains without unacceptable storage, latency, or operational complexity.

The first signal is independent evaluation of the released weights. Researchers and retrieval vendors can now compare both models on public visual-document tasks and private industry corpora.

Reproduction should cover the asymmetric configuration, not only symmetric 0.6B and 9B tests. The central claim depends on small-model queries retaining value against a large-model index.

If independent results preserve the reported gains across legal, financial, technical, and scanned documents, Perplexity’s design becomes a credible deployment pattern. Large quality drops would weaken the shared-space argument.

The second signal is a complete technical report with index measurements. Perplexity says that report will arrive later this year. It should disclose retrieval settings, compression, hardware, latency, and storage per document token.

The report should also explain confidence intervals, training-data filtering, and benchmark configuration. Those details will show whether the reported accuracy gains survive comparable resource limits.

Storage is especially important because output dimension is only one variable. A 128-dimensional token vector sounds compact beside a 4,096-dimensional alternative, but total index size depends on retained token counts.

The third signal is product support. Perplexity says it will progressively add late-interaction, dense, and contextual embeddings to its API platform. A managed endpoint would reveal how the company packages indexing and serving tradeoffs.

API support would also broaden testing beyond teams that can operate custom GPU infrastructure. Adoption will remain narrower if users must assemble rendering, indexing, MaxSim search, and scaling themselves.

The rollout should clarify whether customers can mix 9B document indexes with 0.6B queries through one managed service. That configuration is the release’s strongest operational idea.

Pricing is not yet a useful comparison, and no figures should be inferred from Perplexity’s earlier embedding services. Multi-vector storage and scoring differ substantially from single-vector text embeddings.

Competitor responses also matter, but they are supporting evidence rather than the main test. Qwen, Nvidia, Google, Mixedbread, and other retrieval providers can improve quality, reduce dimensions, or offer easier managed systems.

Perplexity’s advantage will not rest on one leaderboard snapshot. It will rest on whether the shared space lowers live-query costs while preserving enough of the larger indexer’s retrieval quality.

For developers, the immediate action is a bounded evaluation. Build a representative corpus, render the visually complex pages, and preserve a text baseline. Then compare symmetric and asymmetric configurations under the same retrieval budget.

Measure answer accuracy, evidence-page recall, index size, ingestion throughput, query latency, and failure cases. Include OCR and hybrid pipelines, since visual retrieval does not eliminate every reason to parse text.

Perplexity pplx-embed-v2-late presents a clear hypothesis: document encoding and query encoding should not share the same compute budget. The open weights make that hypothesis testable.

The remaining question is operational, not conceptual. Can a 9B index and 0.6B query path outperform simpler retrieval after storage, scoring, updates, permissions, and downstream evidence extraction are counted?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page