Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Magento StyleSmuggler zero-day exploited to deploy Linux backdoor

    September 7, 2026

    Middle East Crypto Activity Triples to $350 Billion Amid Ongoing Conflict, Report Finds

    September 7, 2026

    Story about ‘K9 Valor’ surviving suicide bomber, saving 47 soldiers isn’t what it seems

    September 7, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Magento StyleSmuggler zero-day exploited to deploy Linux backdoor
    • Middle East Crypto Activity Triples to $350 Billion Amid Ongoing Conflict, Report Finds
    • Story about ‘K9 Valor’ surviving suicide bomber, saving 47 soldiers isn’t what it seems
    • Alleged killers of Australian surfers and American friend go on trial in Mexico | Mexico
    • Ireland hails EU tax agreement on carbon imports and electronic waste – POLITICO
    • A Reform UK without Nigel Farage used to be unthinkable. Not any more | Gaby Hinsliff
    • UK begins talks to rejoin EU security missions after Brexit
    • Best Tech Labor Day Sales I’d Shop Myself (2026): Vacuums, Headphones, and More
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Monday, September 7
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKSeptember 6, 2026 Artificial Intelligence No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Retrieval quality in an AI search product is bounded by two things: how good the embedding model is, and how cheaply you can run it across an index. This week, Perplexity Engineering team published Fast Embeddings on GPUs, an under-the-hood account of the second — the serving infrastructure behind pplx-embed and the ranking models used across Perplexity Search, Computer and the API Platform.

    Perplexity team states that embedding inference on the GPU side has largely converged across engines on mature Hopper and Blackwell hardware. The wins sit in the runtime and harness around the model: CUDA graph management, an async result-tracking abstraction, and a Rust request path.

    Two traffic patterns, one engine

    Perplexity frames embedding serving as two workloads. Batch embedding happens when building or re-indexing the vector database, where throughput minimizes cost. Online embedding happens at query time, where a short query must be embedded fast. Scoring sits in between: after vector search, large document batches are ranked, balancing both.

    The key decision is that Perplexity did not build a separate embedding engine. Because embedding models are small Transformers, batch embedding resembles compute-bound prefill and online embedding, often a few tokens, resembles memory-bound decode. So the research team reuses the prefill and decode kernels from its LLM stack.

    Ivy, Tulip and ROSE

    Three services handle a request:

    • Ivy is a Rust HTTP gateway. It does the CPU-side work — JSON parsing, tokenization, input templating, batch splitting — and translates requests into a custom gRPC protocol. It also splits large-batch requests into chunks and load-balances them across replicas, which corrects the load imbalance that arises when production payloads vary in size.
    • Tulip is the inference server interface: a gRPC server built with Rust, tokio and tonic, handling scheduling and batching before dispatching to the engine.
    • ROSE (Runtime-Optimized Serving Engine) implements model inference. It is primarily Python, provides kernels, layers and model definitions, manages CUDA graphs, and exposes a step() function to Tulip.

    Why the scheduler is deliberately simple

    Tulip picks sequences first-come, first-served while requests accumulate. That simplicity is justified by a measurement: for small embedding models at the sequence lengths Perplexity serves, the linear cost of dense layers dominates the quadratic cost of attention. Latency is therefore roughly proportional to token count, not sequence count. Once a batch saturates the GPU, around 512 tokens on a sub-billion-parameter model, packing in more sequences does not improve efficiency.

    CUDA graphs and LazyTensors

    On small batches, CPU-side kernel launching can outweigh GPU execution. Perplexity builds whole-model CUDA graphs for all embedding models, capturing every launch into a single driver call. Because embedding models are small, the inflection point where GPU work exceeds launch cost arrives at batches of thousands of tokens and tens of sequences. Some attention implementations block full-model graphs by depending on dynamic host-side inputs; Perplexity upstreamed changes to FlashInfer to enable capture.

    Graphs must be captured per configuration, so token counts are padded to buckets that are multiples of 64 or 256. That still yields thousands of graphs and multiple minutes of capture per model. The fix is lazy capture: each configuration gets an eager warmup run, then triggers capture and replay on its second hit. This costs p99 latency at startup but spreads minutes of eager work across hours.

    The second piece is the LazyTensor, which tracks a page-locked host buffer plus a cudaMemcpyAsync and a CUDA event. Instead of step() blocking on the device, it returns a LazyTensor, letting a Rust async task wait on batch N while the CPU enqueues N+1.

    details Embedding GPU Ivy Perplexity pplxembed rose serve Stack Tulip
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    MG Ship adds AI route optimisation as logistics returns accelerate

    IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B

    H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

    Meta FAIR Introduces AI Research Preference Models (RPMs): Ranking ML Experiments Before Spending GPU Hours

    UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents

    GitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Magento StyleSmuggler zero-day exploited to deploy Linux backdoor

    September 7, 2026

    Middle East Crypto Activity Triples to $350 Billion Amid Ongoing Conflict, Report Finds

    September 7, 2026

    Story about ‘K9 Valor’ surviving suicide bomber, saving 47 soldiers isn’t what it seems

    September 7, 2026

    Alleged killers of Australian surfers and American friend go on trial in Mexico | Mexico

    September 7, 2026
    Latest Posts

    Book Review: ‘Pure Men’ by Mohamed Mbougar Sarr

    August 1, 2026

    Bitcoin ETFs Post First Monthly Inflow Since April

    August 1, 2026

    Adobe Campaign Classic CVSS 10.0 Flaw Could Run Code Without User Interaction

    August 1, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Magento StyleSmuggler zero-day exploited to deploy Linux backdoor

    September 7, 2026

    Middle East Crypto Activity Triples to $350 Billion Amid Ongoing Conflict, Report Finds

    September 7, 2026

    Story about ‘K9 Valor’ surviving suicide bomber, saving 47 soldiers isn’t what it seems

    September 7, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.