Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Macron to hold first talks with UK’s Burnham after viewing Bayeux Tapestry in London

    September 3, 2026

    How Merkel and Merz lost eastern Germany to the AfD – POLITICO

    September 3, 2026

    MoD withheld information about nuclear test veterans’ records, ex-minister claims

    September 3, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Macron to hold first talks with UK’s Burnham after viewing Bayeux Tapestry in London
    • How Merkel and Merz lost eastern Germany to the AfD – POLITICO
    • MoD withheld information about nuclear test veterans’ records, ex-minister claims
    • Hands-on with Philips Hue’s new Liane 360° Rope Lights
    • Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
    • Researcher Releases FalconFlank PoC Showing Privilege Escalation in CrowdStrike Falcon
    • Thailand puts private wallets and offshore crypto transfers on notice in a major new crypto rule
    • Ontology forces urgent node upgrade after restarting chain hit by malicious activity
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Thursday, September 3
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKSeptember 3, 2026 Artificial Intelligence No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. It is a single-process runtime: a Rust layer loads the checkpoint and drives the generation loop, an OpenAI-compatible chat-completions API streams tokens, and hand-written Metal kernels execute the model. Neither PyTorch nor MLX sits in the execution path. Lily is deliberately narrow with one model, Qwen3.6-35B-A3B, on one hardware family and that narrowness is the performance argument.

    Is it deployable? Yes. A standalone demo is public in the pplx-garden repository. A Rust and Metal inference server offering greedy text generation through a minimal OpenAI-compatible HTTP API. The 4-bit checkpoint is 19.4 GB, so an Apple silicon Mac with 32 GB or more of unified memory is the realistic floor; Perplexity’s shipping Hybrid Compute product lists macOS 15+, 24 GB minimum and 32 GB for best results.

    Why specialize at all?

    The default Mac stack is MLX plus MLX-LM, which already ships a Qwen implementation with grouped expert work, a fused recurrent Metal kernel, and GQA-aware attention. But its operations must stay reusable across architectures. Lily gives that up and puts model structure, execution plans, and kernel selection in one runtime.

    Three workload shapes

    Qwen3.6-35B-A3B stores 35B parameters and activates roughly 3B per token. A router scores 256 experts and picks eight, alongside one shared expert that sees every token. It also mixes 10 full-attention layers using grouped-query attention (16 query heads, two KV heads) with 30 Gated DeltaNet layers. That yields three patterns: uneven expert groups, attention over a growing KV cache, and a fixed-size recurrence.

    Prefill: keep weights packed, keep routing on the GPU

    The checkpoint uses groupwise affine 4-bit quantization, every group of 64 weights sharing a bfloat16 scale and bias, about 70 GB of bfloat16 weights compressed to 19.4 GB. Metal 4 tensor operations consume bfloat16, so weights must be reconstructed first. Lily does that one tile at a time inside the grouped GEMM, holding results in threadgroup memory and accumulating in FP32, so the expanded array never reaches unified memory. In Perplexity’s ablation that fusion raised end-to-end prefill 77.4% at a 512-token prompt.

    Keeping the routing histogram, prefix scan, scatter and block map inside a single GPU command buffer added 89% at 512 tokens by removing CPU synchronization inside each MoE layer. Moving from 16-row to 32-row tiles with four simdgroups added 13.2% at 2K; a register-resident Gated DeltaNet scan added 5.6%. Expert GEMMs are roughly 90% of prefill time. Long prompts run in bounded chunks so temporary activations do not compete with weights and cache for memory.

    Decode: minimize bytes moved per token

    Batch-1 decode has almost no weight reuse, so bandwidth sets the ceiling. One recorded step launched 795 kernels forming 555 sequential stages; Lily records real dependencies in a concurrent Metal pass so independent kernels overlap. The selected token is written straight into the next step’s GPU-resident input slot, removing a per-token CPU round trip, and four kernel chains are fused to keep intermediates in registers.

    Coalesced cache reads lifted key bandwidth from 33.8 to 47.9 GB/s and value bandwidth from 42.0 to 61.8 GB/s. GQA packing, four query heads sharing one threadgroup so each KV row loads once, improved decode 23.8% at 32K. A fixed-block attention layout at 32K and above improved decode 7.7% at 32K, 27.4% at 64K, and 40.2% at 128K.

    Results

    On one 40-core, 128 GB M5 Max at batch 1, loading identical 4-bit checkpoint bytes against MLX-LM’s fastest direct-generation path across ten lengths from 256 to 128K tokens, Lily averaged 4,156 prefill tokens/s versus 3,388 (1.23x) and 170.0 decode tokens/s versus 126.4 (1.35x). At a 4K prompt and 4K context it reached 5,749.9 and 186.6 tokens/s against 4,737.5 and 140.9, and was faster at every recorded point: 1.12–1.42x prefill, 1.31–1.37x decode. A teacher-forced check across 192 positions put Lily’s perplexity 0.04% higher, with the same top-ranked token 96.35% of the time.

    Key Takeaways

    • Lily is a Rust + Metal engine for Qwen3.6-35B-A3B on Apple silicon, with no PyTorch or MLX in the path.
    • Averages 1.23x MLX-LM prefill and 1.35x decode on a 40-core, 128 GB M5 Max.
    • Biggest prefill wins: GPU-resident expert routing (+89%) and dequantization fused into the grouped GEMM (+77.4%).
    • Biggest decode wins: GQA packing (+23.8% at 32K) and fixed-block attention (+40.2% at 128K).

    Check out the Technical details here and the GitHub Repo here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

    Apple engine Inference Lily Metal open Perplexity Qwen3.635BA3B Rust Silicon Sources
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Palo Alto Networks paid $500M for Thrive-backed Console, sources say

    Why MCP servers are becoming AI’s newest attack surface

    Qwen Developers Open-Sources zg (zvec-grep): A Local-First Search Layer Unifying ripgrep, BM25, and Vector Search

    ChatGPT Ads passes $1B run rate in 200 days

    Motional and MIT AI explains self-driving car decisions

    Google DeepMind Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber: One Core Model, Two Access Envelopes

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Macron to hold first talks with UK’s Burnham after viewing Bayeux Tapestry in London

    September 3, 2026

    How Merkel and Merz lost eastern Germany to the AfD – POLITICO

    September 3, 2026

    MoD withheld information about nuclear test veterans’ records, ex-minister claims

    September 3, 2026

    Hands-on with Philips Hue’s new Liane 360° Rope Lights

    September 3, 2026
    Latest Posts

    Australia news live: Reformers member tells hearing he used factional funds to pay for bucks night; Taylor refuses to answer multiple Icac-related questions | Australia news

    July 31, 2026

    Trump administration to end Medicare Part D subsidy program. Will costs increase?

    July 31, 2026

    FP Live: Daniel Yergin on Why Energy Prices Didn’t Soar Higher This Year

    July 31, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Macron to hold first talks with UK’s Burnham after viewing Bayeux Tapestry in London

    September 3, 2026

    How Merkel and Merz lost eastern Germany to the AfD – POLITICO

    September 3, 2026

    MoD withheld information about nuclear test veterans’ records, ex-minister claims

    September 3, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.