Average nDCG@10 versus pooling factor at the full 2048-dimensional width, with image and Markdown panels for native and crosslingual queries.
Average nDCG@10 versus pooling factor at the full 2048-dimensional width, with image and Markdown panels for native and crosslingual queries. Open source material ↗

Evidence at a glance

token 128 。Evidence
18B ColBERT。Evidence
0.6B parameters 594M; parametersApprox. 340M。Evidence
7.4B。9B parameters
9B 92.4%。MADQA
9B 64.0%。BrowseComp+

The mechanism in one line

InputReduce the input to a workable scale

Compress the visual or contextual input before the main reasoning path.

MechanismSpend compute where it matters

Route or verify the expensive step instead of repeating the full path.

OutcomeEnd with a measurable workflow result

Translate the mechanism into a bounded deployment or evaluation check.

One Retrieval System, Two Different Cost Centers

Perplexity has released pplx-embed-v2-late, a family of multimodal embedding models for retrieving text, images, and rendered PDF pages. It comes in 0.6B and 9B sizes and addresses a practical tension: retrieval needs sufficiently fine-grained matches between a question and document content, but using a larger model for every live query raises compute and latency costs.

The design choice worth examining is that Perplexity does not equate a larger model with running a larger model on every query. It proposes building a high-quality index with the 9B model and encoding live queries with the 0.6B model. The two models share an embedding space, making that division possible in principle. The tradeoff is that some cost moves from query-time inference to index creation and storage.

Token-Level Matching Gives Up the Compactness of a Single Vector

A conventional dense embedding compresses an entire document into one vector. That keeps indexes compact and retrieval relatively straightforward, but compression can erase local correspondences. pplx-embed-v2-late instead produces a 128-dimensional vector for each token and uses MaxSim at retrieval time: each query token finds its closest document token, and the resulting match scores are aggregated. A concept in a question can therefore match a local expression in a document without requiring the whole document and query to be similar as single vectors.

That extra detail is not free. Every token leaves a representation in the index, so longer documents generally produce more stored vectors, and scoring candidates means processing more tokens. The 128 dimensions are much narrower than the 2,048 to 4,096 dimensions cited for competing models, but that alone does not make the index smaller. A single-vector model and a token-level model store different units of information. Teams need to estimate total cost from corpus length, chunking, index implementation, and candidate volume together.

Rendered Pages Avoid an OCR Step, but Input Boundaries Remain

The models can retrieve pages after they are rendered as images, targeting not only selectable text but also scanned documents, PDF pages, and slides. Compared with parsing a page or running OCR before retrieval, this route may preserve information in layout, charts, and visual page structure. Perplexity presents it as a direction for multimodal retrieval, but the supplied evidence does not independently establish how much it helps across different document types.

There is also a clear input constraint: a single input cannot mix text and images. In other words, the ability to retrieve a page as an image does not mean an application can freely combine text and image inputs in one request. Before integration, teams should establish how documents are rendered, how text queries and image pages are encoded, and whether the application must coordinate separate input paths. A multimodal label alone does not imply an unrestricted joint-query interface.

A 92.4% Result Is Strong Evidence, Not a Universal Quality Claim

Perplexity reports 92.4% accuracy for the 9B model on MADQA, compared with 90.1% for the 0.6B model. The company describes the evaluation as covering 800 PDFs, more than 18,000 pages, and 500 human-written questions, with a Gemini 3.5 Flash agent searching through the retriever under test. In the same comparison, Gemini 3.5 Flash paired with a Mixedbread retriever scored 88.9%, while the Mixedbread Agentic Search setup scored 93.4%. The result is strong for this PDF question-answering evaluation, but it does not establish a lead across retrieval tasks in general.

Other metrics produce a more mixed picture. The model card lists ViDoRe v3 image-retrieval nDCG@10 scores of 62.3% for 0.6B and 65.2% for 9B. On ViDoRe v3 Markdown, the scores are 61.2% and 64.7%, with 0.6B ranked second in the comparison described in the supplied material. The material also notes stronger competitors in image retrieval and different rankings on BrowseComp+ and domain-specific text tasks. These numbers come from Perplexity's announcement or model cards, and a technical report is not yet available. They should be read as publisher-reported benchmark results, not as independently reproduced conclusions.

Mixed Sizes Change the System Split, Not the Cost of Retrieval

A shared embedding space offers a practical option: use different model sizes for indexing and querying. One example in the supplied material uses a 9B-built index with 0.6B queries and reaches 63.5% on ViDoRe v3 image retrieval, above the 62.3% achieved when both sides use 0.6B, while retaining the smaller model at query time. Perplexity also says this arrangement recovers about half the text-quality gap between the 9B model and a lightweight setup, but the material gives no corresponding score. That claim should not be treated as a guarantee across workloads.

Deployment figures also need their stated qualifications. The 0.6B model has about 594 million total parameters, with roughly 240 million active for text and 340 million for images. The 9B model card lists about 7.4B active parameters, while the Hugging Face page describes it as 8B and the release name says 9B. The supplied estimate for bf16 weight memory is about 1.2 GB for the smaller model and 16 to 18 GB for the larger one; these are weight-only estimates, not full runtime requirements. The published F32 weights take roughly twice as much download space. Both models are available under the MIT license for self-hosting, but Perplexity's hosted API is described as planned, not live.

Budget for Long-Document Indexes Before Changing Architecture

Teams working with large collections of PDFs, scanned reports, or slides should consider pplx-embed-v2-late for evaluation, particularly when visual information is hard to preserve through ordinary text parsing. A lightweight query model paired with a higher-quality index also makes it possible to keep online inference local while moving heavier computation to offline indexing. But self-hostable models do not make the full system lightweight: index creation, storage, updates, and candidate scoring remain part of the architecture's cost.

A sound decision process starts with representative documents and real queries. Measure index size, build time, retrieval latency, and task quality, then compare them with the existing dense-retrieval or OCR pipeline. Include long documents and chart-heavy pages, since token-level storage costs vary with content volume. If better local matching and visual-page retrieval justify the indexing overhead, the design has a practical case. If the main constraints are storage, updates, or straightforward text search, a single 92.4% benchmark result is not enough reason to migrate.