Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Restrict sale of pet flea treatments, government urged

    August 5, 2026

    Elon Musk repeatedly one-upped his execs on SpaceX’s first earnings call

    August 5, 2026

    Google Deletes 3 ADK AI Workflows After Malicious GitHub Issue Could Trigger Privileged Agent

    August 5, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Restrict sale of pet flea treatments, government urged
    • Elon Musk repeatedly one-upped his execs on SpaceX’s first earnings call
    • Google Deletes 3 ADK AI Workflows After Malicious GitHub Issue Could Trigger Privileged Agent
    • Self Custody Is Dead. Long Live Self Custody
    • NASA’s PUNCH Sharpens Solar Storm Forecasting in First Test
    • Edward Fishman on the Top Global Chokepoints—and How to Navigate Them
    • Firefighters in Greece gain ground against wildfires near Athens
    • Charity Commission to investigate donations to illegal Israeli settlements | Charities
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Wednesday, August 5
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Pixel-Native RAG: A Practical Guide to Visual Document Indexing

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 4, 2026 Artificial Intelligence No Comments7 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    @dataclass
    class Tile:
       tile_id: str
       doc_id: str
       source: str
       kind: str
       page: int
       seq: int
       y0: int
       y1: int
       path: str
       ocr_text: str = ""
       title: str = ""
    def _doc_id_from_source(src: str) -> str:
       tail = src.rstrip("/").split("/")[-1] or src
       tail = re.sub(r".(html?|pdf|png|jpg)$", "", tail, flags=re.I)
       return re.sub(r"[^A-Za-z0-9_.-()]+", "_", tail)[:80] or hashlib.md5(src.encode()).hexdigest()[:10]
    def _ahash(img, size: int = 8) -> int:
       """64-bit average hash — cheap near-duplicate detection for repeated headers."""
       import numpy as np
       g = img.convert("L").resize((size, size))
       a = np.asarray(g, dtype="float32")
       bits = (a > a.mean()).flatten()
       out = 0
       for b in bits:
           out = (out << 1) | int(b)
       return out
    def _hamming(a: int, b: int) -> int:
       return bin(a ^ b).count("1")
    def _is_informative(img, cfg: Config) -> bool:
       """Reject blank / solid-colour tiles before they ever reach the GPU."""
       import numpy as np
       a = np.asarray(img.convert("L"), dtype="float32")
       return float(a.std()) >= cfg.blank_std_threshold
    def _save_tile(img, out_dir: Path, name: str) -> str:
       out_dir.mkdir(parents=True, exist_ok=True)
       p = out_dir / f"{name}.png"
       img.convert("RGB").save(p, format="PNG", optimize=True)
       return str(p)
    def slice_image_to_tiles(img, cfg: Config, *, doc_id: str, source: str, kind: str,
                            page: int, out_dir: Path, start_seq: int = 0,
                            seen_hashes: Optional[List[int]] = None,
                            title: str = "") -> List[Tile]:
       """Vertical sliding window with overlap. Used for PDFs and text fallback."""
       from PIL import Image
       seen_hashes = seen_hashes if seen_hashes is not None else []
       W, H = img.size
       if W != cfg.tile_width:
           new_h = max(1, int(H * cfg.tile_width / W))
           img = img.resize((cfg.tile_width, new_h))
           W, H = img.size
       step = max(1, cfg.tile_height - cfg.tile_overlap)
       tiles: List[Tile] = []
       y, seq = 0, start_seq
       while y < H and (seq - start_seq) < cfg.max_tiles_per_doc:
           h = min(cfg.tile_height, H - y)
           if h < cfg.min_tile_height and seq > start_seq:
               break
           crop = img.crop((0, y, W, y + h))
           if _is_informative(crop, cfg):
               hsh = _ahash(crop)
               if all(_hamming(hsh, s) > cfg.dedup_hamming for s in seen_hashes):
                   seen_hashes.append(hsh)
                   tid = f"{doc_id}__p{page}__t{seq}"
                   tiles.append(Tile(
                       tile_id=tid, doc_id=doc_id, source=source, kind=kind, page=page,
                       seq=seq, y0=y, y1=y + h, title=title,
                       path=_save_tile(crop, out_dir, tid),
                   ))
                   seq += 1
           y += step
       return tiles
    _JS_AUTOSCROLL = """
    async () => {
     await new Promise((resolve) => {
       let y = 0;
       const timer = setInterval(() => {
         window.scrollBy(0, 800);
         y += 800;
         if (y >= document.body.scrollHeight || y > 40000) {
           clearInterval(timer);
           window.scrollTo(0, 0);
           setTimeout(resolve, 250);
         }
       }, 40);
     });
    }
    """
    _JS_FLATTEN = """
    () => {
     document.querySelectorAll('*').forEach((el) => {
       const s = getComputedStyle(el);
       if (s.position === 'fixed' || s.position === 'sticky') el.style.position = 'absolute';
     });
     document.querySelectorAll('[role="dialog"], .cookie, #cookie-banner, .cc-banner')
       .forEach((el) => el.remove());
    }
    """
    _CSS_CLEANUP = """
    * { animation: none !important; transition: none !important;
       scroll-behavior: auto !important; }
    html { -webkit-font-smoothing: antialiased; }
    video, iframe[src*="youtube"] { visibility: hidden !important; }
    """
    _UA = ("Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) "
          "Chrome/124.0 Safari/537.36 PixelRAG-Tutorial/1.0")
    async def _render_urls_async(urls: List[str], cfg: Config, out_dir: Path) -> List[Tile]:
       from playwright.async_api import async_playwright
       from PIL import Image
       all_tiles: List[Tile] = []
       async with async_playwright() as pw:
           browser = await pw.chromium.launch(headless=True, args=cfg.headless_args)
           ctx = await browser.new_context(
               viewport={"width": cfg.tile_width, "height": cfg.tile_height},
               device_scale_factor=cfg.device_scale,
               user_agent=_UA,
               java_script_enabled=True,
           )
           for url in urls:
               doc_id = _doc_id_from_source(url)
               page = await ctx.new_page()
               try:
                   await page.goto(url, wait_until="domcontentloaded", timeout=cfg.nav_timeout_ms)
                   try:
                       await page.wait_for_load_state("networkidle", timeout=12000)
                   except Exception:
                       pass
                   await page.evaluate(_JS_AUTOSCROLL)
                   await page.add_style_tag(content=_CSS_CLEANUP)
                   await page.evaluate(_JS_FLATTEN)
                   title = (await page.title()) or doc_id
                   height = await page.evaluate(
                       "() => Math.max(document.body.scrollHeight, "
                       "document.documentElement.scrollHeight)")
                   height = int(min(height, cfg.max_page_height))
                   step = max(1, cfg.tile_height - cfg.tile_overlap)
                   seen: List[int] = []
                   y, seq = 0, 0
                   while y < height and seq < cfg.max_tiles_per_doc:
                       h = min(cfg.tile_height, height - y)
                       if h < cfg.min_tile_height and seq > 0:
                           break
                       buf = await page.screenshot(
                           full_page=True, type="png",
                           clip={"x": 0, "y": y, "width": cfg.tile_width, "height": h})
                       img = Image.open(io.BytesIO(buf)).convert("RGB")
                       if img.size[0] != cfg.tile_width:
                           img = img.resize((cfg.tile_width,
                                             max(1, int(img.size[1] * cfg.tile_width / img.size[0]))))
                       if _is_informative(img, cfg):
                           hsh = _ahash(img)
                           if all(_hamming(hsh, s) > cfg.dedup_hamming for s in seen):
                               seen.append(hsh)
                               tid = f"{doc_id}__p0__t{seq}"
                               all_tiles.append(Tile(
                                   tile_id=tid, doc_id=doc_id, source=url, kind="web",
                                   page=0, seq=seq, y0=y, y1=y + h, title=title,
                                   path=_save_tile(img, out_dir, tid)))
                               seq += 1
                       y += step
                   log.info("  rendered %-34s -> %2d tiles (page %dpx)", doc_id, seq, height)
               except Exception as exc:
                   log.warning("  FAILED %s (%s)", url, type(exc).__name__)
               finally:
                   await page.close()
           await ctx.close()
           await browser.close()
       return all_tiles
    def render_urls(urls: List[str], cfg: Config, out_dir: Path) -> List[Tile]:
       """Screenshot every URL into tiles; degrade to the text renderer on failure."""
       try:
           tiles = run_async(_render_urls_async(urls, cfg, out_dir))
           if tiles:
               return tiles
           log.warning("Browser produced no tiles — using text-render fallback.")
       except Exception as exc:
           log.warning("Playwright unavailable (%s: %s) — using text-render fallback.",
                       type(exc).__name__, str(exc)[:160])
       return [t for u in urls for t in render_url_as_text(u, cfg, out_dir)]
    def _strip_html(html: str) -> str:
       html = re.sub(r"(?is)<(script|style|nav|footer|header|noscript).*?1>", " ", html)
       html = re.sub(r"(?s)", " ", html)
       html = re.sub(r"(?i)(p|div|h[1-6]|li|tr|br)>", "n", html)
       text = re.sub(r"(?s)<[^>]+>", " ", html)
       for a, b in [(" ", " "), ("&", "&"), ("<", "<"), (">", ">"), (""", '"')]:
           text = text.replace(a, b)
       text = re.sub(r"[d+]", "", text)
       text = re.sub(r"[ t]+", " ", text)
       return re.sub(r"n{2,}", "n", text).strip()
    def _mono_font(size: int = 20):
       from PIL import ImageFont
       for cand in ("/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf",
                    "/usr/share/fonts/truetype/liberation/LiberationSans-Regular.ttf"):
           if os.path.exists(cand):
               return ImageFont.truetype(cand, size)
       try:
           import matplotlib.font_manager as fm
           return ImageFont.truetype(fm.findfont("DejaVu Sans"), size)
       except Exception:
           return ImageFont.load_default()
    def text_to_image(text: str, cfg: Config, title: str = "") -> Any:
       """Render plain text onto a tall white canvas — a browser-free stand-in."""
       from PIL import Image, ImageDraw
       font, tfont = _mono_font(20), _mono_font(30)
       pad, lh, wrap = 40, 30, max(20, (cfg.tile_width - 80) // 11)
       lines: List[str] = []
       for para in text.split("n"):
           para = para.strip()
           if not para:
               continue
           while len(para) > wrap:
               cut = para.rfind(" ", 0, wrap)
               cut = cut if cut > 0 else wrap
               lines.append(para[:cut])
               para = para[cut:].lstrip()
           lines.append(para)
       lines = lines[:900]
       height = pad * 2 + 60 + lh * len(lines)
       img = Image.new("RGB", (cfg.tile_width, max(cfg.tile_height, height)), "white")
       d = ImageDraw.Draw(img)
       d.text((pad, pad), title[:60], font=tfont, fill=(15, 15, 15))
       for i, ln in enumerate(lines):
           d.text((pad, pad + 60 + i * lh), ln, font=font, fill=(35, 35, 35))
       return img
    def render_url_as_text(url: str, cfg: Config, out_dir: Path) -> List[Tile]:
       import requests
       doc_id = _doc_id_from_source(url)
       try:
           r = requests.get(url, timeout=30, headers={"User-Agent": _UA})
           r.raise_for_status()
           body = _strip_html(r.text)
           m = re.search(r"(?is)(.*?)", r.text)
           title = m.group(1).strip() if m else doc_id
       except Exception as exc:
           log.warning("  fetch failed for %s (%s)", url, type(exc).__name__)
           return []
       img = text_to_image(body, cfg, title=title)
       log.info("  text-rendered %-30s -> canvas %dpx", doc_id, img.size[1])
       return slice_image_to_tiles(img, cfg, doc_id=doc_id, source=url, kind="text",
                                   page=0, out_dir=out_dir, title=title)
    def render_pdf(pdf_path: str, cfg: Config, out_dir: Path, dpi: int = 150) -> List[Tile]:
       import fitz
       from PIL import Image
       doc_id = _doc_id_from_source(pdf_path)
       tiles: List[Tile] = []
       with fitz.open(pdf_path) as doc:
           title = (doc.metadata or {}).get("title") or doc_id
           n_pages = doc.page_count
           for pno in range(n_pages):
               pix = doc[pno].get_pixmap(dpi=dpi)
               img = Image.frombytes("RGB", (pix.width, pix.height), pix.samples)
               tiles += slice_image_to_tiles(img, cfg, doc_id=doc_id, source=pdf_path,
                                             kind="pdf", page=pno, out_dir=out_dir,
                                             title=title)
       log.info("  rendered %-34s -> %2d tiles (%d pages)", doc_id, len(tiles), n_pages)
       return tiles
    def make_synthetic_pdf(path: Path) -> str:
       """A tiny PDF so the tutorial always exercises the PDF path, offline or not."""
       import fitz
       body = [
           ("PixelRAG Internal Note", 22),
           ("", 12),
           ("Why pixel-native retrieval?", 16),
           ("Parsers are per-site glue code. A renderer is one code path for every", 11),
           ("document type: HTML, PDF, scanned fax, spreadsheet export, dashboard.", 11),
           ("", 11),
           ("Tiling policy", 16),
           ("Tiles are 1024x1024 with 128px of vertical overlap. Overlap keeps a", 11),
           ("sentence or table row from being split across two embeddings, which is", 11),
           ("the single biggest source of recall loss in naive screenshot pipelines.", 11),
           ("", 11),
           ("Serving", 16),
           ("FAISS inner-product over L2-normalised vectors equals cosine similarity.", 11),
           ("Tile scores are max-pooled per document so one strong tile can surface", 11),
           ("a long page, mirroring late-interaction retrieval behaviour.", 11),
           ("", 11),
           ("The mitochondria reference is a joke; the overlap advice is not.", 11),
       ]
       doc = fitz.open()
       page = doc.new_page()
       y = 72
       for line, size in body:
           page.insert_text((72, y), line, fontsize=size, fontname="helv")
           y += size + 8
       doc.save(str(path))
       doc.close()
       return str(path)
    
    document Guide Indexing PixelNative Practical RAG Visual
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

    Building an Advanced AI Skill Security Auditing Pipeline with NVIDIA SkillSpector, LangGraph, YARA Rules, SARIF, and CI Policy Gates

    Reflex Open Sources XY: A Rust-Backed Super-Fast Python Charting Library That Keeps 100 Million Point Charts Interactive

    Genspark Open Sources GenOffice: A Free, Ad-Free AI Office Suite for macOS and Windows with Docs, Sheets, Slides, PDF

    Y Combinator Open-Sources QM: An MIT-Licensed Multiplayer Agent Harness That Runs In Slack And The Web

    Evaluating Multimodal Vision Models with Moonshot PerceptionBench Using Robust Data Loading and Automated Judging

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Restrict sale of pet flea treatments, government urged

    August 5, 2026

    Elon Musk repeatedly one-upped his execs on SpaceX’s first earnings call

    August 5, 2026

    Google Deletes 3 ADK AI Workflows After Malicious GitHub Issue Could Trigger Privileged Agent

    August 5, 2026

    Self Custody Is Dead. Long Live Self Custody

    August 4, 2026
    Latest Posts

    Oil prices hit $100 for the first time since May

    July 23, 2026

    Pew Survey: China May Be Liked More, but It Is Celebrating a Race It Never Ran

    July 23, 2026

    Yinson Production and PTSC’s FSO heads off to Southeast Asian oil project

    July 23, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Restrict sale of pet flea treatments, government urged

    August 5, 2026

    Elon Musk repeatedly one-upped his execs on SpaceX’s first earnings call

    August 5, 2026

    Google Deletes 3 ADK AI Workflows After Malicious GitHub Issue Could Trigger Privileged Agent

    August 5, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.