Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Here’s what connects the Cornell alleged rape case to those before: accused men not being held to serious account | Emma Brockes

    October 8, 2026

    Most UK diplomats to leave East Jerusalem consulate, Israel says, as Miliband says ‘vital services’ to remain

    October 8, 2026

    Polanski’s final pitch in Holborn by-election: I’ll keep Burnham honest – POLITICO

    October 8, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Here’s what connects the Cornell alleged rape case to those before: accused men not being held to serious account | Emma Brockes
    • Most UK diplomats to leave East Jerusalem consulate, Israel says, as Miliband says ‘vital services’ to remain
    • Polanski’s final pitch in Holborn by-election: I’ll keep Burnham honest – POLITICO
    • British consulate in East Jerusalem will stay open as UK mission, says Ed Miliband | Foreign policy
    • Interview with Corriere della Sera
    • Tropical Storm Isaias Set to Be First Atlantic Hurricane of 2026 Season
    • Tensorlake npm Package Compromised to Deliver Shai-Hulud Credential-Stealing Worm
    • SEC and CFTC Crypto Rules ‘Fall Short’ of Clarity, Says Rep. French Hill
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Thursday, October 8
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKOctober 8, 2026 Artificial Intelligence No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Perplexity has released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models. They come in 2 sizes: 0.6B for fast, cheap queries and 9B for maximum quality. Both models retrieve text, images and rendered PDF pages, and they share one embedding space.

    Is it deployable? Yes, if you host it yourself. Both models are on Hugging Face under the MIT license. A hosted Perplexity API endpoint is planned but not live.

    TL;DR

    The best

    • The 0.6B model uses about 340M active parameters for images and stays close to 8B rivals.
    • A 9B index can be searched with 0.6B queries, recovering about half the 9B quality gap on text at 0.6B query cost.
    • Its 128-dim token vectors are 16x to 32x narrower than rivals at 2,048 to 4,096 dims.
    • MIT license, with commercial use allowed.

    The worst

    • It stores 1 vector per token, so index size grows with document length.
    • It is not #1 on ViDoRe v3 image retrieval; Tencent’s EVIE scores higher.
    • A single input cannot mix text and images.
    • All scores are self-reported, and the technical report is not out yet.

    Model Size and What It Runs On

    Metrics pplx-embed-v2-late-0.6b pplx-embed-v2-late-9b
    Total parameters 594M 9B (Hugging Face lists 8B)
    Active parameters ~240M text, 340M image 7.4B
    Base model Qwen3.5-0.8B, pruned to 12 text layers Qwen3.5
    Output 128 dims per token 128 dims per token
    Weights in memory (bf16, our estimate) ~1.2 GB ~16 to 18 GB
    Intended machine Laptop, edge device or small GPU Datacenter or high-memory GPU
    Perplexity’s suggested role Live query encoder, 100% local Building the document index

    Perplexity designed the 0.6B model as a lightweight query encoder that can also run on edge devices. Both model cards show CUDA GPU usage. They need sentence-transformers >= 6.0.0 and transformers >= 5.4.0. The memory figures are our estimate at 2 bytes per parameter, for weights only. The published checkpoints are stored in F32, which doubles the download size.

    How Well It Performs: Best and Worst Scores

    All numbers below are from Perplexity’s announcement:

    Benchmark 0.6B 9B Where it stands
    MADQA (agentic PDF QA, accuracy) 90.1% 92.4% (best) Beats Mixedbread’s retriever (88.9%); trails Mixedbread Agentic Search (93.4%)
    Domain-specific text (72 tasks, nDCG@10) 78.0% 81.3% 9B leads all tested models by 1.6pp; 0.6B is 0.3pp behind gemini-embedding-2
    Q2D-Web (Recall@1000) 73.6% 74.8% Both beat the previous best of 69.3%
    ViDoRe v3 image (nDCG@10) 62.3% 65.2% 0.6B is within 1.2pp of nemotron-colembed-v2-8b; EVIE leads
    ViDoRe v3 Markdown (nDCG@10) 61.2% (worst) 64.7% Both beat every external model tested
    BrowseComp+ (accuracy) Not given 64.0% (lowest) Still 4.9pp above the next ColBERT model

    The strongest result: 92.4% on MADQA, set by the 9B model. The biggest margin is on BrowseComp+, at 8.7pp over the best dense model.

    The weakest result: 61.2% on ViDoRe v3 Markdown, from the 0.6B model. It is still the 2nd-best score on that benchmark. Image search is the real gap: Gemini Embedding 2 beats the 9B model on MIRACL-Vision and by 2pp on PPLX-Q2I.

    Mixing sizes: a 9B index queried by the 0.6B model scored 63.5% on ViDoRe v3 image retrieval. That beats 62.3% with 0.6B on both sides, at the same query cost.

    How It Works

    Dense models compress a document into 1 vector. pplx-embed-v2-late instead keeps a 128-dim vector for every token. It scores with MaxSim: each query token finds its best document token, and those maxima are summed. Pages are encoded as images, so no OCR step is needed. Perplexity distilled both models from an 18B teacher using LEAF-style token-level training. That training is what creates the shared space.

    Best Use Cases

    • The best fit is visual document search over PDFs, slides and scanned reports.
    • Low-latency search: index in the cloud with 9B, then query on-device with 0.6B.
    • Agentic RAG over large PDF or web collections.

    Interactive Explainer

    “;D.forEach(function(d){h+=’

    ‘+d+’

    ‘});h+=’

    max

    ‘;
    Q.forEach(function(q,i){h+=’

    ‘+q+’

    ‘;S[i].forEach(function(v,j){h+=’

    ‘+v.toFixed(2)+’

    ‘});h+=’

    ‘+Math.max.apply(null,S[i]).toFixed(2)+’

    ‘});
    g.innerHTML=h;
    function col(v){return ‘rgba(32,128,141,’+(0.12+v*0.88)+’)’}
    var running=false,timers=[];
    function clearAll(){timers.forEach(clearTimeout);timers=[];$$(‘.cell’).forEach(function(c){c.className=”cell”;c.style.background=”});$$(‘.mx’).forEach(function(m){m.classList.remove(‘show’)});$$(‘.qt’).forEach(function(q){q.classList.remove(‘act’)});setScore(0)}
    function setScore(s){$(‘#sval’).textContent=s.toFixed(2);$(‘#sfill’).style.width=(s/3*100)+’%’}
    function later(f,t){timers.push(setTimeout(f,t))}
    function runRow(i,t0,acc,cb){
    var best=S[i].indexOf(Math.max.apply(null,S[i]));
    later(function(){R.querySelector(‘[data-q=”‘+i+'”]’).classList.add(‘act’)},t0);
    S[i].forEach(function(v,j){later(function(){var c=$(‘#c’+i+’_’+j);c.classList.add(‘show’,’scan’);c.style.background=col(v);if(j>0)$(‘#c’+i+’_’+(j-1)).classList.remove(‘scan’)},t0+120+j*110)});
    later(function(){$(‘#c’+i+’_7’).classList.remove(‘scan’);$(‘#c’+i+’_’+best).classList.add(‘max’);$(‘#m’+i).classList.add(‘show’);setScore(acc+S[i][best]);if(cb)cb(acc+S[i][best])},t0+120+8*110+150);
    }
    $(‘#run’).onclick=function(){clearAll();var acc=0,t=0;Q.forEach(function(q,i){(function(i,t){runRow(i,t,0,null)})(i,t);t+=1250});
    var sums=[0];Q.forEach(function(q,i){sums.push(sums[i]+Math.max.apply(null,S[i]))});
    Q.forEach(function(q,i){later(function(){setScore(sums[i+1])},i*1250+120+8*110+160)});};
    $$(‘.qt’).forEach(function(el){el.onclick=function(){var i=+el.dataset.q;clearAll();runRow(i,0,0,null)}});
    $(‘#dbtn’).onclick=function(){var d=$(‘#dense’);var on=d.classList.toggle(‘on’);var sq=$(‘#squash’);sq.classList.remove(‘squashed’);if(on){setTimeout(function(){sq.classList.add(‘squashed’)},250)}this.textContent=on?’Hide dense comparison’:’Compare with a dense vector’;setTimeout(fit,80);setTimeout(fit,1200)};

    /* 2. Storage */
    var DOCS=[1e3,1e4,1e5,1e6,1e7,1e8],DL=[‘1K’,’10K’,’100K’,’1M’,’10M’,’100M’];
    var M=[{n:’pplx-embed-v2-late (128/token)’,d:128,p:1},{n:’topk-embed-v1-small (2,048/token, full)’,d:2048},{n:’nemotron-colembed-vl-8b-v2 (4,096/token)’,d:4096},{n:’Gemini Embedding 2 (1 x 3,072, dense)’,d:3072,single:1}];
    function fmt(b){var u=[‘B’,’KB’,’MB’,’GB’,’TB’,’PB’],i=0;while(b>=1000&&i<5){b/=1000;i++}return (b<10?b.toFixed(2):b<100?b.toFixed(1):Math.round(b))+’ ‘+u[i]}
    function drawStore(){var tk=+$(‘#tok’).value,dc=DOCS[+$(‘#docs’).value];$(‘#tokv’).textContent=tk;$(‘#docsv’).textContent=DL[+$(‘#docs’).value];
    var vals=M.map(function(m){return (m.single?1:tk)*m.d*2*dc});var mx=Math.max.apply(null,vals);
    var b=$(‘#bars’);if(!b.children.length){b.innerHTML=M.map(function(m,i){return ”}).join(”)}
    vals.forEach(function(v,i){$(‘#f’+i).style.width=Math.max(1.5,v/mx*100)+’%’;$(‘#n’+i).textContent=fmt(v)});}
    $(‘#tok’).oninput=drawStore;$(‘#docs’).oninput=drawStore;

    /* 3. Shared space */
    var MODES=[
    {i:’9B’,q:’9B’,wi:’Cloud GPU, amortized’,wq:’Cloud GPU, every request’,r:’Best quality: 81.3% average on 72 domain-specific tasks and 65.2% on ViDoRe v3 image retrieval.’},
    {i:’0.6B’,q:’0.6B’,wi:’Local or edge’,wq:’Local or edge’,r:’Cheapest, can run 100% locally: 78.0% domain average and 62.3% on ViDoRe v3, with about 240M active params for text and 340M for images.’},
    {i:’9B’,q:’0.6B’,wi:’Cloud GPU, one-off’,wq:’Fast, cheap encoder’,r:’Asymmetric setup: +1.6pp over 0.6B on both sides for text, and 63.5% vs 62.3% on ViDoRe v3, at the same query cost. Perplexity says it recovers roughly half the 9B gap on text.’},
    {i:’9B’,q:’0.6B’,wi:’Cloud-hosted index’,wq:’On-device docs and queries’,r:’Encode private local documents or live queries on-device with 0.6B, then compare or merge them with results from a 9B cloud index. No separate benchmark is reported for this mode.’}];
    function setMode(k){$$(‘.mode’).forEach(function(m){m.classList.toggle(‘on’,+m.dataset.m===k)});var m=MODES[k];
    [[‘#mi’,m.i],[‘#mq’,m.q]].forEach(function(p){var el=$(p[0]);el.textContent=p[1];el.className=”model “+(p[1]===’9B’?’big’:’small’);void el.offsetWidth;el.classList.add(‘pop’)});
    $(‘#wi’).textContent=m.wi;$(‘#wq’).textContent=m.wq;$(‘#res’).innerHTML=m.r;setTimeout(fit,60)}
    $$(‘.mode’).forEach(function(m){m.onclick=function(){setMode(+m.dataset.m)}});setMode(0);

    /* 4. Benchmarks */
    var B=[
    {k:’ViDoRe v3 (image)’,u:’nDCG@10′,rows:[[‘pplx-embed-v2-late-9b’,65.2,1],[‘nemotron-colembed-v2-8b’,63.5],[‘pplx-embed-v2-late-0.6b’,62.3,1]],n:’Perplexity says its 0.6B model sits within 1.2pp of nemotron-colembed-v2-8b (63.5% implied) while activating 340M params for images. Tencent EVIE leads; it is listed by Perplexity as a vision-only baseline.’},
    {k:’ViDoRe v3 (Markdown)’,u:’nDCG@10′,rows:[[‘pplx-embed-v2-late-9b’,64.7,1],[‘pplx-embed-v2-late-0.6b’,61.2,1],[‘Next best external model’,60.6],[‘topk-embed-v1-0.8b’,56.9]],n:’Text retrieval over OCR-parsed pages. topk-embed-v1-0.8b is 4.3pp behind the 0.6B model, per Perplexity.’},
    {k:’Domain-specific (72 tasks)’,u:’avg nDCG@10′,rows:[[‘pplx-embed-v2-late-9b’,81.3,1],[‘Next best model’,79.7],[‘gemini-embedding-2’,78.3],[‘pplx-embed-v2-late-0.6b’,78.0,1]],n:’Finance, legal, health, conversation, tech and multilingual tasks excluded from training data. Next best and Gemini values derived from the 1.6pp and 0.3pp gaps Perplexity reports.’},
    {k:’Q2D-Web’,u:’Recall@1000′,rows:[[‘pplx-embed-v2-late-9b’,74.8,1],[‘pplx-embed-v2-late-0.6b’,73.6,1],[‘Previous best’,69.3]],n:’Combined judgments on Perplexity’s 70K-query, 190M-document web benchmark.’},
    {k:’BrowseComp+’,u:’accuracy %’,rows:[[‘pplx-embed-v2-late-9b’,64.0,1],[‘Next best ColBERT model’,59.1],[‘Next best dense model’,55.3]],n:’GPT-OSS-120B agent. Baseline values derived from the 4.9pp and 8.7pp gaps Perplexity reports.’},
    {k:’MADQA’,u:’accuracy %’,rows:[[‘Mixedbread Agentic Search’,93.4],[‘pplx-embed-v2-late-9b’,92.4,1],[‘pplx-embed-v2-late-0.6b’,90.1,1],[‘Mixedbread retriever’,88.9]],n:’Gemini 3.5 Flash agent over 800 PDFs and 18,000+ pages. Mixedbread Agentic Search uses a planning sub-agent; Perplexity says the gap is within its confidence interval.’}];
    var curB=0;$(‘#pills’).innerHTML=B.map(function(b,i){return ‘‘+b.k+’‘}).join(”);
    function drawBench(k){curB=k;$$(‘.pill’).forEach(function(p){p.classList.toggle(‘on’,+p.dataset.b===k)});var b=B[k],lo=Math.min.apply(null,b.rows.map(function(r){return r[1]}))-8;
    $(‘#bbars’).innerHTML=b.rows.map(function(r,i){return ‘

    ‘+r[0]+’‘+r[1].toFixed(1)+’%

    ‘}).join(”);
    $(‘#bnote’).textContent=b.u+’. ‘+b.n;
    requestAnimationFrame(function(){requestAnimationFrame(function(){b.rows.forEach(function(r,i){$(‘#bf’+i).style.width=((r[1]-lo)/(100-lo)*100)+’%’})})});setTimeout(fit,60)}
    $$(‘.pill’).forEach(function(p){p.onclick=function(){drawBench(+p.dataset.b)}});

    window.addEventListener(‘load’,function(){fit();setTimeout(function(){$(‘#run’).click()},500)});window.addEventListener(‘resize’,fit);setTimeout(fit,300);
    })();

    0.6B Edge MADQA model Perplexity pplxembedv2late Releases Scoring
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Anthropic Releases Claude Haiku 5.5: A Small Model With 1M Context Priced at $0.10 per Million Input Tokens

    Anthropic Launches Haiku 5.5: Its Cheapest and Fastest Claude Model Yet

    What Happens When a Trusted Model Repo Changes? Unsloth Studio Re-Checks Before It Runs

    OpenAI Says a Secret AI Model Cracked Hundreds of Open Math Problems in One Prompt—Mathematicians Want Receipts

    Microsoft releases new Nvidia-chip AI PCs with revamped Windows 11

    Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Here’s what connects the Cornell alleged rape case to those before: accused men not being held to serious account | Emma Brockes

    October 8, 2026

    Most UK diplomats to leave East Jerusalem consulate, Israel says, as Miliband says ‘vital services’ to remain

    October 8, 2026

    Polanski’s final pitch in Holborn by-election: I’ll keep Burnham honest – POLITICO

    October 8, 2026

    British consulate in East Jerusalem will stay open as UK mission, says Ed Miliband | Foreign policy

    October 8, 2026
    Latest Posts

    British national shot dead in Kashmir by Pakistani security forces | Kashmir

    August 10, 2026

    Climate change doubled likelihood of Canada’s extreme fire weather, study finds

    August 10, 2026

    Scientists say just 7 days of meditation can rewire your brain

    August 10, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Here’s what connects the Cornell alleged rape case to those before: accused men not being held to serious account | Emma Brockes

    October 8, 2026

    Most UK diplomats to leave East Jerusalem consulate, Israel says, as Miliband says ‘vital services’ to remain

    October 8, 2026

    Polanski’s final pitch in Holborn by-election: I’ll keep Burnham honest – POLITICO

    October 8, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.