Perplexity has released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models. They come in 2 sizes: 0.6B for fast, cheap queries and 9B for maximum quality. Both models retrieve text, images and rendered PDF pages, and they share one embedding space.
Is it deployable? Yes, if you host it yourself. Both models are on Hugging Face under the MIT license. A hosted Perplexity API endpoint is planned but not live.
TL;DR
The best
- The 0.6B model uses about 340M active parameters for images and stays close to 8B rivals.
- A 9B index can be searched with 0.6B queries, recovering about half the 9B quality gap on text at 0.6B query cost.
- Its 128-dim token vectors are 16x to 32x narrower than rivals at 2,048 to 4,096 dims.
- MIT license, with commercial use allowed.
The worst
- It stores 1 vector per token, so index size grows with document length.
- It is not #1 on ViDoRe v3 image retrieval; Tencent’s EVIE scores higher.
- A single input cannot mix text and images.
- All scores are self-reported, and the technical report is not out yet.
Model Size and What It Runs On
| Metrics | pplx-embed-v2-late-0.6b | pplx-embed-v2-late-9b |
|---|---|---|
| Total parameters | 594M | 9B (Hugging Face lists 8B) |
| Active parameters | ~240M text, 340M image | 7.4B |
| Base model | Qwen3.5-0.8B, pruned to 12 text layers | Qwen3.5 |
| Output | 128 dims per token | 128 dims per token |
| Weights in memory (bf16, our estimate) | ~1.2 GB | ~16 to 18 GB |
| Intended machine | Laptop, edge device or small GPU | Datacenter or high-memory GPU |
| Perplexity’s suggested role | Live query encoder, 100% local | Building the document index |
Perplexity designed the 0.6B model as a lightweight query encoder that can also run on edge devices. Both model cards show CUDA GPU usage. They need sentence-transformers >= 6.0.0 and transformers >= 5.4.0. The memory figures are our estimate at 2 bytes per parameter, for weights only. The published checkpoints are stored in F32, which doubles the download size.
How Well It Performs: Best and Worst Scores
All numbers below are from Perplexity’s announcement:
| Benchmark | 0.6B | 9B | Where it stands |
|---|---|---|---|
| MADQA (agentic PDF QA, accuracy) | 90.1% | 92.4% (best) | Beats Mixedbread’s retriever (88.9%); trails Mixedbread Agentic Search (93.4%) |
| Domain-specific text (72 tasks, nDCG@10) | 78.0% | 81.3% | 9B leads all tested models by 1.6pp; 0.6B is 0.3pp behind gemini-embedding-2 |
| Q2D-Web (Recall@1000) | 73.6% | 74.8% | Both beat the previous best of 69.3% |
| ViDoRe v3 image (nDCG@10) | 62.3% | 65.2% | 0.6B is within 1.2pp of nemotron-colembed-v2-8b; EVIE leads |
| ViDoRe v3 Markdown (nDCG@10) | 61.2% (worst) | 64.7% | Both beat every external model tested |
| BrowseComp+ (accuracy) | Not given | 64.0% (lowest) | Still 4.9pp above the next ColBERT model |
The strongest result: 92.4% on MADQA, set by the 9B model. The biggest margin is on BrowseComp+, at 8.7pp over the best dense model.
The weakest result: 61.2% on ViDoRe v3 Markdown, from the 0.6B model. It is still the 2nd-best score on that benchmark. Image search is the real gap: Gemini Embedding 2 beats the 9B model on MIRACL-Vision and by 2pp on PPLX-Q2I.
Mixing sizes: a 9B index queried by the 0.6B model scored 63.5% on ViDoRe v3 image retrieval. That beats 62.3% with 0.6B on both sides, at the same query cost.
How It Works
Dense models compress a document into 1 vector. pplx-embed-v2-late instead keeps a 128-dim vector for every token. It scores with MaxSim: each query token finds its best document token, and those maxima are summed. Pages are encoded as images, so no OCR step is needed. Perplexity distilled both models from an 18B teacher using LEAF-style token-level training. That training is what creates the shared space.
Best Use Cases
- The best fit is visual document search over PDFs, slides and scanned reports.
- Low-latency search: index in the cloud with 9B, then query on-device with 0.6B.
- Agentic RAG over large PDF or web collections.
Interactive Explainer
“;D.forEach(function(d){h+=’
‘+d+’
‘});h+=’
max
‘;
Q.forEach(function(q,i){h+=’
‘+q+’
‘;S[i].forEach(function(v,j){h+=’
‘+v.toFixed(2)+’
‘});h+=’
‘+Math.max.apply(null,S[i]).toFixed(2)+’
‘});
g.innerHTML=h;
function col(v){return ‘rgba(32,128,141,’+(0.12+v*0.88)+’)’}
var running=false,timers=[];
function clearAll(){timers.forEach(clearTimeout);timers=[];$$(‘.cell’).forEach(function(c){c.className=”cell”;c.style.background=”});$$(‘.mx’).forEach(function(m){m.classList.remove(‘show’)});$$(‘.qt’).forEach(function(q){q.classList.remove(‘act’)});setScore(0)}
function setScore(s){$(‘#sval’).textContent=s.toFixed(2);$(‘#sfill’).style.width=(s/3*100)+’%’}
function later(f,t){timers.push(setTimeout(f,t))}
function runRow(i,t0,acc,cb){
var best=S[i].indexOf(Math.max.apply(null,S[i]));
later(function(){R.querySelector(‘[data-q=”‘+i+'”]’).classList.add(‘act’)},t0);
S[i].forEach(function(v,j){later(function(){var c=$(‘#c’+i+’_’+j);c.classList.add(‘show’,’scan’);c.style.background=col(v);if(j>0)$(‘#c’+i+’_’+(j-1)).classList.remove(‘scan’)},t0+120+j*110)});
later(function(){$(‘#c’+i+’_7’).classList.remove(‘scan’);$(‘#c’+i+’_’+best).classList.add(‘max’);$(‘#m’+i).classList.add(‘show’);setScore(acc+S[i][best]);if(cb)cb(acc+S[i][best])},t0+120+8*110+150);
}
$(‘#run’).onclick=function(){clearAll();var acc=0,t=0;Q.forEach(function(q,i){(function(i,t){runRow(i,t,0,null)})(i,t);t+=1250});
var sums=[0];Q.forEach(function(q,i){sums.push(sums[i]+Math.max.apply(null,S[i]))});
Q.forEach(function(q,i){later(function(){setScore(sums[i+1])},i*1250+120+8*110+160)});};
$$(‘.qt’).forEach(function(el){el.onclick=function(){var i=+el.dataset.q;clearAll();runRow(i,0,0,null)}});
$(‘#dbtn’).onclick=function(){var d=$(‘#dense’);var on=d.classList.toggle(‘on’);var sq=$(‘#squash’);sq.classList.remove(‘squashed’);if(on){setTimeout(function(){sq.classList.add(‘squashed’)},250)}this.textContent=on?’Hide dense comparison’:’Compare with a dense vector’;setTimeout(fit,80);setTimeout(fit,1200)};
/* 2. Storage */
var DOCS=[1e3,1e4,1e5,1e6,1e7,1e8],DL=[‘1K’,’10K’,’100K’,’1M’,’10M’,’100M’];
var M=[{n:’pplx-embed-v2-late (128/token)’,d:128,p:1},{n:’topk-embed-v1-small (2,048/token, full)’,d:2048},{n:’nemotron-colembed-vl-8b-v2 (4,096/token)’,d:4096},{n:’Gemini Embedding 2 (1 x 3,072, dense)’,d:3072,single:1}];
function fmt(b){var u=[‘B’,’KB’,’MB’,’GB’,’TB’,’PB’],i=0;while(b>=1000&&i<5){b/=1000;i++}return (b<10?b.toFixed(2):b<100?b.toFixed(1):Math.round(b))+’ ‘+u[i]}
function drawStore(){var tk=+$(‘#tok’).value,dc=DOCS[+$(‘#docs’).value];$(‘#tokv’).textContent=tk;$(‘#docsv’).textContent=DL[+$(‘#docs’).value];
var vals=M.map(function(m){return (m.single?1:tk)*m.d*2*dc});var mx=Math.max.apply(null,vals);
var b=$(‘#bars’);if(!b.children.length){b.innerHTML=M.map(function(m,i){return ”}).join(”)}
vals.forEach(function(v,i){$(‘#f’+i).style.width=Math.max(1.5,v/mx*100)+’%’;$(‘#n’+i).textContent=fmt(v)});}
$(‘#tok’).oninput=drawStore;$(‘#docs’).oninput=drawStore;
/* 3. Shared space */
var MODES=[
{i:’9B’,q:’9B’,wi:’Cloud GPU, amortized’,wq:’Cloud GPU, every request’,r:’Best quality: 81.3% average on 72 domain-specific tasks and 65.2% on ViDoRe v3 image retrieval.’},
{i:’0.6B’,q:’0.6B’,wi:’Local or edge’,wq:’Local or edge’,r:’Cheapest, can run 100% locally: 78.0% domain average and 62.3% on ViDoRe v3, with about 240M active params for text and 340M for images.’},
{i:’9B’,q:’0.6B’,wi:’Cloud GPU, one-off’,wq:’Fast, cheap encoder’,r:’Asymmetric setup: +1.6pp over 0.6B on both sides for text, and 63.5% vs 62.3% on ViDoRe v3, at the same query cost. Perplexity says it recovers roughly half the 9B gap on text.’},
{i:’9B’,q:’0.6B’,wi:’Cloud-hosted index’,wq:’On-device docs and queries’,r:’Encode private local documents or live queries on-device with 0.6B, then compare or merge them with results from a 9B cloud index. No separate benchmark is reported for this mode.’}];
function setMode(k){$$(‘.mode’).forEach(function(m){m.classList.toggle(‘on’,+m.dataset.m===k)});var m=MODES[k];
[[‘#mi’,m.i],[‘#mq’,m.q]].forEach(function(p){var el=$(p[0]);el.textContent=p[1];el.className=”model “+(p[1]===’9B’?’big’:’small’);void el.offsetWidth;el.classList.add(‘pop’)});
$(‘#wi’).textContent=m.wi;$(‘#wq’).textContent=m.wq;$(‘#res’).innerHTML=m.r;setTimeout(fit,60)}
$$(‘.mode’).forEach(function(m){m.onclick=function(){setMode(+m.dataset.m)}});setMode(0);
/* 4. Benchmarks */
var B=[
{k:’ViDoRe v3 (image)’,u:’nDCG@10′,rows:[[‘pplx-embed-v2-late-9b’,65.2,1],[‘nemotron-colembed-v2-8b’,63.5],[‘pplx-embed-v2-late-0.6b’,62.3,1]],n:’Perplexity says its 0.6B model sits within 1.2pp of nemotron-colembed-v2-8b (63.5% implied) while activating 340M params for images. Tencent EVIE leads; it is listed by Perplexity as a vision-only baseline.’},
{k:’ViDoRe v3 (Markdown)’,u:’nDCG@10′,rows:[[‘pplx-embed-v2-late-9b’,64.7,1],[‘pplx-embed-v2-late-0.6b’,61.2,1],[‘Next best external model’,60.6],[‘topk-embed-v1-0.8b’,56.9]],n:’Text retrieval over OCR-parsed pages. topk-embed-v1-0.8b is 4.3pp behind the 0.6B model, per Perplexity.’},
{k:’Domain-specific (72 tasks)’,u:’avg nDCG@10′,rows:[[‘pplx-embed-v2-late-9b’,81.3,1],[‘Next best model’,79.7],[‘gemini-embedding-2’,78.3],[‘pplx-embed-v2-late-0.6b’,78.0,1]],n:’Finance, legal, health, conversation, tech and multilingual tasks excluded from training data. Next best and Gemini values derived from the 1.6pp and 0.3pp gaps Perplexity reports.’},
{k:’Q2D-Web’,u:’Recall@1000′,rows:[[‘pplx-embed-v2-late-9b’,74.8,1],[‘pplx-embed-v2-late-0.6b’,73.6,1],[‘Previous best’,69.3]],n:’Combined judgments on Perplexity’s 70K-query, 190M-document web benchmark.’},
{k:’BrowseComp+’,u:’accuracy %’,rows:[[‘pplx-embed-v2-late-9b’,64.0,1],[‘Next best ColBERT model’,59.1],[‘Next best dense model’,55.3]],n:’GPT-OSS-120B agent. Baseline values derived from the 4.9pp and 8.7pp gaps Perplexity reports.’},
{k:’MADQA’,u:’accuracy %’,rows:[[‘Mixedbread Agentic Search’,93.4],[‘pplx-embed-v2-late-9b’,92.4,1],[‘pplx-embed-v2-late-0.6b’,90.1,1],[‘Mixedbread retriever’,88.9]],n:’Gemini 3.5 Flash agent over 800 PDFs and 18,000+ pages. Mixedbread Agentic Search uses a planning sub-agent; Perplexity says the gap is within its confidence interval.’}];
var curB=0;$(‘#pills’).innerHTML=B.map(function(b,i){return ‘‘+b.k+’‘}).join(”);
function drawBench(k){curB=k;$$(‘.pill’).forEach(function(p){p.classList.toggle(‘on’,+p.dataset.b===k)});var b=B[k],lo=Math.min.apply(null,b.rows.map(function(r){return r[1]}))-8;
$(‘#bbars’).innerHTML=b.rows.map(function(r,i){return ‘
‘}).join(”);
$(‘#bnote’).textContent=b.u+’. ‘+b.n;
requestAnimationFrame(function(){requestAnimationFrame(function(){b.rows.forEach(function(r,i){$(‘#bf’+i).style.width=((r[1]-lo)/(100-lo)*100)+’%’})})});setTimeout(fit,60)}
$$(‘.pill’).forEach(function(p){p.onclick=function(){drawBench(+p.dataset.b)}});
window.addEventListener(‘load’,function(){fit();setTimeout(function(){$(‘#run’).click()},500)});window.addEventListener(‘resize’,fit);setTimeout(fit,300);
})();


