Web Scraping Will Never Be the Same: Introducing PixelRAG
Web Scraping Will Never Be the Same: Introducing PixelRAG
PixelRAG is an open-source retriever framework that uses page images instead of traditional HTML parsing.
PixelRAG is an open-source retriever framework that utilizes page images instead of traditional HTML parsing.
According to the developers, traditional HTML-to-text pipelines can lose over 40% of page content, including tables, graphs, and markup elements. PixelRAG operates on the document as it is rendered for the user.
How the Pipeline Works:
- Renders each document (web pages, PDFs, images) into a set of tiles.
- Creates embeddings using Qwen3-VL-Embedding, fine-tuned via LoRA on screenshots.
- Builds a FAISS index and provides an API for searching.
By replacing the reader model with a more powerful one, accuracy can increase without re-indexing, as the index only retains pixels.
For experiments, the project team created a visual index of the entire Wikipedia - over 30 million screenshots. As a result, even in this format, the system outperforms the best text-based RAG baseline by 18.1% in text-only question-answering tasks.
A plugin for Claude Code has also been introduced, allowing for the analysis of rendered pages through screenshots without working with the DOM.
The entire project is published as open access under the Apache-2.0 license, and the article includes detailed error analyses, ablation studies, and comparisons with more than 25 VLM models.
Why it matters
AnalysisThis new approach to web scraping could significantly enhance data collection accuracy, reducing information loss, which is crucial for developers and researchers.
Discuss in community
Share your questions and insights with developers