RAG benchmark — JDF vs PDF
The question: if the same document enters a RAG pipeline as a PDF or as a JDF, which one lets the retriever find the right passage, and what does each cost to index and query? Everything else is held equal: same content, same embedding models, same lexical retriever, same questions, same hit rule. Everything on this page is generated from bench/results/*.json; nothing is typed in by hand.
git clone https://github.com/uurtech/jdf && cd jdf/bench
pip install -r requirements.txt
python rag_bench.py # accuracy: BM25 + a local embedding model
python rag_bench.py --embedder ollama:nomic-embed-text,st:BAAI/bge-base-en-v1.5
python cost_bench.py # RAG cost: 1,000 PDF files vs 1,000 JDF files
python rag_bench.py --verify # re-run and compare with the published numbers
Each installed PDF parser (PyMuPDF, pdfplumber, pypdf, poppler's pdftotext) becomes a pipeline; missing ones are skipped. Nothing from the corpus leaves your machine.
In 42 seconds
Every number on screen is read from docs/bench.json. Source: demos/rag-benchmark.
Retrieval accuracy
PDF side: pdf → text → LangChain-style fixed chunks (1000/200 and 2000/200 chars) → retriever. JDF side: jdf → jdf chunk --strategy section → retriever — section-aware chunks, tables serialised as Header: value rows, the heading breadcrumb prefixed to the text that gets embedded. Third pipeline, PDF → jdf convert → jdf chunk: the same PDFs converted with the shipped importer (tables rebuilt as real table elements, headings detected from font size), then chunked like native JDF — the path for anyone who only has PDFs today. Same retrievers for all. A retrieved chunk is a hit when it comes from the right document and contains both the answer and its row/subject key.
BM25 — lexical, offline, deterministic
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 100.0% | 100.0% | 89.6% | 100.0% | 0.945 | 100.0% | 798 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 100.0% | 100.0% | 85.4% | 99.5% | 0.925 | 100.0% | 848 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 98.4% | 100.0% | 75.0% | 99.0% | 0.858 | 100.0% | 1,194 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 97.9% | 100.0% | 84.4% | 100.0% | 0.918 | 99.2% | 1,759 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 99.5% | 100.0% | 75.5% | 100.0% | 0.865 | 100.0% | 1,190 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 97.9% | 100.0% | 84.4% | 100.0% | 0.918 | 99.2% | 1,744 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 99.5% | 100.0% | 75.0% | 100.0% | 0.863 | 100.0% | 1,193 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 97.9% | 100.0% | 84.4% | 100.0% | 0.918 | 99.2% | 1,755 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 85.9% | 91.1% | 68.2% | 87.0% | 0.766 | 89.2% | 1,016 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 92.7% | 98.4% | 73.4% | 99.5% | 0.852 | 92.5% | 2,193 |
Embeddings — nomic-embed-text
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 99.0% | 100.0% | 76.6% | 99.0% | 0.864 | 98.3% | 753 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 99.0% | 99.5% | 80.2% | 98.4% | 0.884 | 98.3% | 721 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 81.8% | 98.4% | 41.7% | 91.1% | 0.606 | 80.0% | 1,226 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 77.1% | 99.5% | 63.0% | 99.0% | 0.760 | 82.5% | 1,781 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 80.7% | 99.0% | 42.2% | 91.1% | 0.608 | 79.2% | 1,219 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 77.1% | 99.0% | 63.5% | 97.9% | 0.763 | 82.5% | 1,773 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 83.3% | 97.9% | 41.1% | 92.2% | 0.605 | 81.7% | 1,224 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 77.1% | 99.0% | 63.0% | 97.4% | 0.759 | 82.5% | 1,774 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 71.9% | 86.5% | 35.9% | 73.4% | 0.508 | 70.8% | 1,040 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 62.0% | 89.1% | 46.4% | 93.8% | 0.642 | 60.0% | 2,269 |
Embeddings — BAAI/bge-small-en-v1.5
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 96.9% | 100.0% | 46.9% | 91.7% | 0.646 | 95.0% | 716 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 95.8% | 98.4% | 40.1% | 91.7% | 0.601 | 94.2% | 523 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 74.5% | 94.3% | 34.4% | 83.3% | 0.530 | 71.7% | 1,214 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 64.1% | 92.2% | 50.0% | 90.6% | 0.649 | 67.5% | 1,783 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 75.0% | 93.8% | 34.9% | 84.4% | 0.534 | 72.5% | 1,210 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 66.1% | 92.2% | 50.5% | 91.7% | 0.659 | 69.2% | 1,771 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 75.0% | 94.8% | 31.8% | 84.4% | 0.515 | 72.5% | 1,210 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 66.1% | 90.1% | 52.1% | 88.5% | 0.665 | 69.2% | 1,777 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 66.7% | 82.8% | 34.9% | 71.4% | 0.487 | 61.7% | 1,044 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 48.4% | 87.0% | 38.0% | 95.3% | 0.572 | 49.2% | 2,211 |
Embeddings — sentence-transformers/all-MiniLM-L6-v2
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 97.4% | 100.0% | 40.1% | 91.1% | 0.599 | 95.8% | 693 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 96.4% | 100.0% | 34.4% | 78.1% | 0.537 | 94.2% | 529 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 70.3% | 90.1% | 30.2% | 80.2% | 0.492 | 63.3% | 1,204 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 60.9% | 82.8% | 45.8% | 81.8% | 0.606 | 63.3% | 1,695 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 65.6% | 89.1% | 28.1% | 78.1% | 0.474 | 58.3% | 1,193 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 61.5% | 83.9% | 45.8% | 82.8% | 0.605 | 62.5% | 1,684 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 66.7% | 88.0% | 28.1% | 78.6% | 0.477 | 59.2% | 1,197 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 62.5% | 84.9% | 42.7% | 81.8% | 0.590 | 63.3% | 1,693 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 66.1% | 82.8% | 29.2% | 68.8% | 0.441 | 60.8% | 1,021 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 54.2% | 83.9% | 37.0% | 91.7% | 0.571 | 50.8% | 2,214 |
Embeddings — BAAI/bge-base-en-v1.5
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 96.9% | 100.0% | 63.0% | 94.8% | 0.758 | 95.0% | 728 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 99.5% | 99.5% | 54.2% | 94.8% | 0.694 | 99.2% | 552 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 78.6% | 97.4% | 31.8% | 90.1% | 0.528 | 81.7% | 1,220 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 56.8% | 94.3% | 46.9% | 92.7% | 0.630 | 69.2% | 1,793 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 80.2% | 97.4% | 32.3% | 90.6% | 0.534 | 83.3% | 1,215 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 57.8% | 95.3% | 47.4% | 93.2% | 0.634 | 70.0% | 1,778 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 77.6% | 98.4% | 30.2% | 91.7% | 0.519 | 78.3% | 1,215 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 57.3% | 94.8% | 46.9% | 91.1% | 0.631 | 69.2% | 1,790 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 74.0% | 87.5% | 34.4% | 76.6% | 0.504 | 73.3% | 1,037 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 44.8% | 89.6% | 34.4% | 95.3% | 0.547 | 48.3% | 2,296 |
R@1k tokens = the answer was inside the first 1,000 tokens of retrieved context (chunk-size neutral: a 2,000-character PDF chunk “hits” more often at top-1 simply because it is a quarter of the document, and then costs 3× the tokens). Hit = right document and the chunk contains both the answer and its row/subject key. 24 documents, 120 pages, 192 questions (120 table cells, 48 prose, 24 list). Embeddings run locally (nomic-embed-text via Ollama 0.32.6; BAAI/bge-small-en-v1.5 via sentence-transformers 6.0.1; sentence-transformers/all-MiniLM-L6-v2 via sentence-transformers 6.0.1; BAAI/bge-base-en-v1.5 via sentence-transformers 6.0.1). JDF chunk hashes verified; editing one paragraph re-embeds 1 of 192 chunks. Apple M5, 2026-09-13.
RAG cost — 1,000 PDF files vs 1,000 JDF files
Same pipeline on both sides, only the input format differs: chunks → embeddings → vector store → top-5 context → LLM. The 24-document corpus is cycled to 1,000 files per format; embedding tokens are counted from the chunks each pipeline produces; embedding time is measured throughput; dollar figures are tokens × the published prices in bench/prices.json. The benchmark never calls a paid API.
| Per 1,000 documents | JDF · jdf chunk (section) | PDF · pypdf extract_text() · fixed 1000/200 |
|---|---|---|
| Accuracy · answer in first 1,000 tokens (nomic-embed-text) | 99.0% | 83.3% |
| Accuracy · top-1 hit | 76.6% | 41.1% |
| Chunks | 8,000 | 6,291 |
| Embedding tokens, initial index | 1,305,820 | 1,390,646 |
| Embedding cost · OpenAI text-embedding-3-small | $0.0261 | $0.0278 |
| Local embedding time · bge-small-en-v1.5 (measured throughput) | 43.8 s | 32.4 s |
| Vector-store payload | 5.8 MB | 5.6 MB |
| Re-embed tokens when one paragraph changes in every document | 91,000 | 1,461,000 |
| Re-index cost · OpenAI text-embedding-3-small | $0.0018 | $0.0292 |
| LLM input tokens per 1,000,000 queries (top-5 context) | 753,364,583 | 1,223,625,000 |
| LLM input cost · Claude Sonnet 5 input | $1,506.73 | $2,447.25 |
1,000 files per format = the 24-document corpus cycled. Tokens are counted from the chunks each pipeline produces (ceil(chars/4), both sides); embedding time is measured throughput on Apple M5 (2026-09-13) applied to the totals. Prices from bench/prices.json (as of 2026-09-12); edit it for your provider — the benchmark never calls a paid API. The PDF column is the PDF pipeline that scored best in the accuracy run. Scaling: figures are measured at 1,000 documents and 1,000,000 queries; anything at other volumes (10,000 documents, 10M queries) is a linear estimate, not a measurement — index and re-index costs scale with documents, query cost scales with questions asked. Re-index: JDF re-embeds only chunks whose content hash changed (jdf embed --incremental); a PDF has no chunk identity, so an edit means re-chunking and re-embedding the whole document.
Method
Corpus
24 “Quarterly Operations Report” documents, 4–5 pages each, generated deterministically from a fixed seed. Each has eight sections: two prose sections, five tables (regional performance, headcount, pricing, service levels, vendor spend) and a risk list, plus a repeating page header and footer. The JDF files are the originals; the PDFs were printed from them by the real jdf.js renderer in Google Chrome, so the PDF text layer is clean and complete — the PDF side is not handicapped by scans or exotic fonts. Every document, every question and every retrieved rank is in bench/.
Questions
192, generated with the corpus and stored with ground truth: 120 table-cell lookups (“What was Acme Logistics's churn rate in EMEA for Q3 2025?”), 48 prose facts (on-call lead, capex budget), 24 list items (risk owner). Every question names the document the way a user would.
Why a generated corpus
A retrieval benchmark needs to know, for every question, which passage answers it. Public PDF collections don't ship that, and hand-labelling hundreds of answers is where benchmark bias usually creeps in. Generating the documents makes the ground truth exact and the whole thing auditable, and --verify proves nobody edited a document or a question after the fact by checking SHA-256 hashes against the manifest. The obvious limitation: synthetic reports are cleaner than real ones — real PDFs are usually worse for the PDF side (scans, multi-column layouts, broken tables), so treat these numbers as a floor for the gap. If you run the pipelines on your own corpus, open an issue with the numbers.
Reading the accuracy numbers honestly
Top-1 depends on chunk size: a 2,000-character PDF chunk is a quarter of the document, so it “hits” more often at rank 1 — and then costs three times the tokens per query. With the two smallest embedding models the PDF 2,000-char configuration edges out JDF at top-1; on the chunk-size-neutral metric (answer within the first 1,000 tokens of context) JDF leads with every retriever tested. Both views are in the tables above.