Edit in place
Double-click any paragraph, heading, list item, table cell, image — inline editor opens for that element only. Enter commits. Auto-saves to disk in 150ms.
the AI-native document standard
Built for the AI era — a JSON document format that LLMs read without a translator. PDFs were designed for humans and printers; JDF is designed for the systems that read, write, and reason over documents next. Beautiful pages. Editable in any text editor. Diffable in git. Generated with one line of code. Searchable with grep. Validatable with a JSON Schema. Embeddable on any web page.
Linux, Windows, and Intel Macs: grab a build from the latest release.
brew tap uurtech/jdf<jdf src="...">npx @uurtech/jdf-cli chunk report.jdf<jdf src="hello-world.jdf"> tag, rendered by jdf.js. View source, embed it on your own page in one line.It's JSON. Every consequence below falls out of that.
JSON is the language LLMs already speak. No PDF parsing, no broken columns, no lost tables. Hand a .jdf to GPT or Claude and it sees the structure exactly as you wrote it. More →
cat, grep, jq, VS Code, every linter, every diff tool. No plugin, no learning curve.
git diff is realChange one heading, the diff is one line. Documents review the same way as code.
JSON.stringify(doc). No PDF library, no font dictionaries, no encoding rules.
Autocomplete in your editor. CLI validation in CI. Type-safe generation.
grep "TODO" *.jdf works. jq queries pull every table cell on the page.
Opens the same way today and in 20 years. JSON has no proprietary owner.
A retrieval pipeline that ingests JDF skips most of the work it would do on a PDF. The structure that PDF parsers try to reconstruct is already in the file.
| Pipeline stage | JDF | |
|---|---|---|
| Parse / extract | pdfplumber / pymupdf / unstructured — layout analysis, font heuristics, OCR fallback for image-only pages |
JSON.parse(content) — no layout reconstruction, the structure is already in the file |
| Chunking | Token-windowed splits that frequently slice through tables, lists, footnotes | Each element (text / richtext / table / list / image) is a natural retrieval unit — no chunker config |
| Metadata | Synthesised after the fact (page number, "is this a heading?") and often wrong | First-class on every element: type, heading, page index, position, link target |
| Embedding noise | Repeated page headers / footers / page numbers leak into chunks | header and footer live in their own tree, never in content chunks |
| Re-indexing on edit | Re-parse + re-chunk + re-embed the whole PDF | Diff the JSON, re-embed only the changed elements |
| Tables | Cells smear across columns; multi-row headers collapse | { headers: [...], rows: [[...]] } — every cell at its real coordinate |
| Images / figures | Dropped or stubbed as [image] |
Stored in resources.images with alt text and a stable anchor — a vision step can fetch it at the exact retrieval point |
Chunking and embedding are built in. jdf chunk reads JDF's heading hierarchy for section-aware chunks (tables serialized as Header: value rows, so column meaning survives). jdf embed turns them into vectors — locally via Ollama by default, so no data leaves your machine — and --incremental re-embeds only the chunks whose content hash changed. Edit one paragraph in a 500-page doc → one embedding call, not five hundred.
# deterministic, offline — same doc → byte-identical chunks & stable hashes
$ jdf chunk report.jdf
{"id":"p3e7","text":"…","path":["Report","Pricing"],"page":3,"types":["text","table"],"tokens":142,"hash":"ab12cd"}
# local embeddings; Ollama auto-starts via Docker if needed; skips unchanged chunks
$ jdf embed report.jdf --incremental # → report.embeddings.json
Twenty-four multi-page operations reports, authored as JDF and printed to PDF by a real browser, so both pipelines see identical content. 192 questions with known answers (120 table cells, 48 prose facts, 24 list items). The PDF side runs the usual Python parsers (PyMuPDF, pdfplumber, pypdf, poppler) and LangChain-style fixed chunking at two sizes; the JDF side runs jdf chunk. A third pipeline answers the question everyone with a PDF archive has: the same PDFs put through jdf convert (tables rebuilt as real table elements, headings detected) and then jdf chunk. Same local embedding models, same BM25, same hit rule for both. Every number below is regenerated by python rag_bench.py / python cost_bench.py and checked by --verify. Method and full tables →
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 100.0% | 100.0% | 89.6% | 100.0% | 0.945 | 100.0% | 798 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 100.0% | 100.0% | 85.4% | 99.5% | 0.925 | 100.0% | 848 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 98.4% | 100.0% | 75.0% | 99.0% | 0.858 | 100.0% | 1,194 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 97.9% | 100.0% | 84.4% | 100.0% | 0.918 | 99.2% | 1,759 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 99.5% | 100.0% | 75.5% | 100.0% | 0.865 | 100.0% | 1,190 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 97.9% | 100.0% | 84.4% | 100.0% | 0.918 | 99.2% | 1,744 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 99.5% | 100.0% | 75.0% | 100.0% | 0.863 | 100.0% | 1,193 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 97.9% | 100.0% | 84.4% | 100.0% | 0.918 | 99.2% | 1,755 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 85.9% | 91.1% | 68.2% | 87.0% | 0.766 | 89.2% | 1,016 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 92.7% | 98.4% | 73.4% | 99.5% | 0.852 | 92.5% | 2,193 |
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 99.0% | 100.0% | 76.6% | 99.0% | 0.864 | 98.3% | 753 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 99.0% | 99.5% | 80.2% | 98.4% | 0.884 | 98.3% | 721 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 81.8% | 98.4% | 41.7% | 91.1% | 0.606 | 80.0% | 1,226 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 77.1% | 99.5% | 63.0% | 99.0% | 0.760 | 82.5% | 1,781 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 80.7% | 99.0% | 42.2% | 91.1% | 0.608 | 79.2% | 1,219 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 77.1% | 99.0% | 63.5% | 97.9% | 0.763 | 82.5% | 1,773 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 83.3% | 97.9% | 41.1% | 92.2% | 0.605 | 81.7% | 1,224 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 77.1% | 99.0% | 63.0% | 97.4% | 0.759 | 82.5% | 1,774 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 71.9% | 86.5% | 35.9% | 73.4% | 0.508 | 70.8% | 1,040 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 62.0% | 89.1% | 46.4% | 93.8% | 0.642 | 60.0% | 2,269 |
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 96.9% | 100.0% | 46.9% | 91.7% | 0.646 | 95.0% | 716 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 95.8% | 98.4% | 40.1% | 91.7% | 0.601 | 94.2% | 523 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 74.5% | 94.3% | 34.4% | 83.3% | 0.530 | 71.7% | 1,214 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 64.1% | 92.2% | 50.0% | 90.6% | 0.649 | 67.5% | 1,783 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 75.0% | 93.8% | 34.9% | 84.4% | 0.534 | 72.5% | 1,210 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 66.1% | 92.2% | 50.5% | 91.7% | 0.659 | 69.2% | 1,771 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 75.0% | 94.8% | 31.8% | 84.4% | 0.515 | 72.5% | 1,210 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 66.1% | 90.1% | 52.1% | 88.5% | 0.665 | 69.2% | 1,777 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 66.7% | 82.8% | 34.9% | 71.4% | 0.487 | 61.7% | 1,044 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 48.4% | 87.0% | 38.0% | 95.3% | 0.572 | 49.2% | 2,211 |
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 97.4% | 100.0% | 40.1% | 91.1% | 0.599 | 95.8% | 693 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 96.4% | 100.0% | 34.4% | 78.1% | 0.537 | 94.2% | 529 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 70.3% | 90.1% | 30.2% | 80.2% | 0.492 | 63.3% | 1,204 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 60.9% | 82.8% | 45.8% | 81.8% | 0.606 | 63.3% | 1,695 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 65.6% | 89.1% | 28.1% | 78.1% | 0.474 | 58.3% | 1,193 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 61.5% | 83.9% | 45.8% | 82.8% | 0.605 | 62.5% | 1,684 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 66.7% | 88.0% | 28.1% | 78.6% | 0.477 | 59.2% | 1,197 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 62.5% | 84.9% | 42.7% | 81.8% | 0.590 | 63.3% | 1,693 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 66.1% | 82.8% | 29.2% | 68.8% | 0.441 | 60.8% | 1,021 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 54.2% | 83.9% | 37.0% | 91.7% | 0.571 | 50.8% | 2,214 |
| Pipeline | Chunks | R@1k tokens | R@2k tokens | Top-1 | Top-5 | MRR@10 | Table cells R@1k | Ctx tokens @5 |
|---|---|---|---|---|---|---|---|---|
| JDF · jdf chunk (section, 512 tok) | 192 | 96.9% | 100.0% | 63.0% | 94.8% | 0.758 | 95.0% | 728 |
| PDF → jdf convert → jdf chunk (section, 512 tok) | 216 | 99.5% | 99.5% | 54.2% | 94.8% | 0.694 | 99.2% | 552 |
| PDF · PyMuPDF get_text() · fixed 1000/200 | 153 | 78.6% | 97.4% | 31.8% | 90.1% | 0.528 | 81.7% | 1,220 |
| PDF · PyMuPDF get_text() · fixed 2000/200 | 90 | 56.8% | 94.3% | 46.9% | 92.7% | 0.630 | 69.2% | 1,793 |
| PDF · pdfplumber extract_text() · fixed 1000/200 | 150 | 80.2% | 97.4% | 32.3% | 90.6% | 0.534 | 83.3% | 1,215 |
| PDF · pdfplumber extract_text() · fixed 2000/200 | 90 | 57.8% | 95.3% | 47.4% | 93.2% | 0.634 | 70.0% | 1,778 |
| PDF · pypdf extract_text() · fixed 1000/200 | 151 | 77.6% | 98.4% | 30.2% | 91.7% | 0.519 | 78.3% | 1,215 |
| PDF · pypdf extract_text() · fixed 2000/200 | 90 | 57.3% | 94.8% | 46.9% | 91.1% | 0.631 | 69.2% | 1,790 |
| PDF · pdftotext -layout (poppler) · fixed 1000/200 | 268 | 74.0% | 87.5% | 34.4% | 76.6% | 0.504 | 73.3% | 1,037 |
| PDF · pdftotext -layout (poppler) · fixed 2000/200 | 96 | 44.8% | 89.6% | 34.4% | 95.3% | 0.547 | 48.3% | 2,296 |
R@1k tokens = the answer was inside the first 1,000 tokens of retrieved context (chunk-size neutral: a 2,000-character PDF chunk “hits” more often at top-1 simply because it is a quarter of the document, and then costs 3× the tokens). Hit = right document and the chunk contains both the answer and its row/subject key. 24 documents, 120 pages, 192 questions (120 table cells, 48 prose, 24 list). Embeddings run locally (nomic-embed-text via Ollama 0.32.6; BAAI/bge-small-en-v1.5 via sentence-transformers 6.0.1; sentence-transformers/all-MiniLM-L6-v2 via sentence-transformers 6.0.1; BAAI/bge-base-en-v1.5 via sentence-transformers 6.0.1). JDF chunk hashes verified; editing one paragraph re-embeds 1 of 192 chunks. Apple M5, 2026-09-13.
Same pipeline, only the input format differs: chunks → embeddings → vector store → top-5 context → LLM. Tokens counted from the chunks each pipeline produces, dollars from published prices; the benchmark never calls a paid API.
| Per 1,000 documents | JDF · jdf chunk (section) | PDF · pypdf extract_text() · fixed 1000/200 |
|---|---|---|
| Accuracy · answer in first 1,000 tokens (nomic-embed-text) | 99.0% | 83.3% |
| Accuracy · top-1 hit | 76.6% | 41.1% |
| Chunks | 8,000 | 6,291 |
| Embedding tokens, initial index | 1,305,820 | 1,390,646 |
| Embedding cost · OpenAI text-embedding-3-small | $0.0261 | $0.0278 |
| Local embedding time · bge-small-en-v1.5 (measured throughput) | 43.8 s | 32.4 s |
| Vector-store payload | 5.8 MB | 5.6 MB |
| Re-embed tokens when one paragraph changes in every document | 91,000 | 1,461,000 |
| Re-index cost · OpenAI text-embedding-3-small | $0.0018 | $0.0292 |
| LLM input tokens per 1,000,000 queries (top-5 context) | 753,364,583 | 1,223,625,000 |
| LLM input cost · Claude Sonnet 5 input | $1,506.73 | $2,447.25 |
1,000 files per format = the 24-document corpus cycled. Tokens are counted from the chunks each pipeline produces (ceil(chars/4), both sides); embedding time is measured throughput on Apple M5 (2026-09-13) applied to the totals. Prices from bench/prices.json (as of 2026-09-12); edit it for your provider — the benchmark never calls a paid API. The PDF column is the PDF pipeline that scored best in the accuracy run. Scaling: figures are measured at 1,000 documents and 1,000,000 queries; anything at other volumes (10,000 documents, 10M queries) is a linear estimate, not a measurement — index and re-index costs scale with documents, query cost scales with questions asked. Re-index: JDF re-embeds only chunks whose content hash changed (jdf embed --incremental); a PDF has no chunk identity, so an edit means re-chunking and re-embedding the whole document.
docs/bench.json, the same file this page uses. Source: demos/rag-benchmark.Synthetic corpus by design — a retrieval benchmark needs ground truth for every question, which public PDF corpora don't ship. The generator, every document, every question and every retrieved rank are in bench/. Run it on your own corpus and open an issue with the numbers.
If an AI-native document format becomes the default, documents themselves become a base layer for knowledge — not files you store, but interfaces software talks to.
A .jdf is a live data structure. Every paragraph, table, and figure has a stable identity that other systems can reference, query, and react to.
No more PDF parsers, no more layout heuristics, no more OCR fallbacks. The structure is in the file — every tool reads the same tree.
Retrieval skips the parse-and-reconstruct stage. An embedding pipeline ingests the document tree directly. The cited span is a real element, not a guess.
The way fax gave way to email and email to chat, the binary print-archive format gives way to the format every machine and model already speaks.
One schema, every consumer. Editors, viewers, summarisers, exporters, agents — all working off the same tree, no glue code per format.
Instead of “AI trying to understand humans’ files”, humans start writing in a format already built for AI reasoning.
The real endgame: documents stop being storage → they become interfaces.
Three unattended workflows the desktop app can't cover: legacy PDFs entering a pipeline, LLM-emitted JSON becoming a renderable document, and RAG ingestion — chunking and embedding, incrementally. All run in CI, all validate against the schema, all fail loudly on malformed input.
Same algorithm the desktop reader uses, packaged as a node CLI. Run it on a build server, in a Lambda, in a CI step — output is the structured tree your retriever wants instead of `pdfplumber` heuristics.
# 1360-page AWS API spec → 60k structured elements in 90s
$ jdf convert partner-api.pdf -o api.jdf --json
$ jdf validate api.jdf # schema gate — exit 1 fails CI
Models emit JSON. Hand them the JDF schema, drop the response into the CLI, and you get a validated `.jdf` you can render, ship, or audit. Three input shapes are accepted: a full document, a bare element array, or a partial.
# model output → validated, renderable document
$ openai-cli generate > report.json
$ jdf convert report.json -o report.jdf
# exit 1 if the model violated the schema
JDF already carries the heading hierarchy, so chunking reads structure instead of guessing. Deterministic chunks → stable hashes → --incremental re-embeds only what changed. Embed locally via Ollama (no data leaves the machine) or a remote API.
# section-aware chunks (tables kept as Header: value)
$ jdf chunk report.jdf # → report.chunks.jsonl
# local embeddings, skip unchanged chunks
$ jdf embed report.jdf --incremental
One step in your workflow rejects malformed JSON or PDFs with weird structure before they reach production:
- name: Convert & validate generated document
run: |
npx @uurtech/jdf-cli convert dist/output.json -o dist/output.jdf
npx @uurtech/jdf-cli validate dist/output.jdf
Embed a `.jdf` form on any web page, let users fill it, click Save, and they get the same `.jdf` back with their answers baked in. The downloaded file is still pure JSON — your backend, your RAG, your audit log all read the same tree.
One JSON document. type: "input" | "textarea" | "checkbox" | "select" | "signature". Each field has a stable name — the value lookup key. Same schema validates, same CLI imports, same TOC and search work.
One tag. save-button attribute drops a button in the corner. The user fills the form right in the browser; jdf.js mutates the in-memory document on every keystroke.
Browser downloads customer-form.jdf. Open it in any text editor — values are inline. Re-open it in jdf.js and the form starts pre-filled. Convert it to PDF with jdf export; pipe it to your RAG with jdf validate.
{
"$jdf": "1.0.0",
"meta": { "title": "Hello", "pageSize": "A4" },
"styles": {
"h1": { "fontSize": 22, "fontWeight": "bold" }
},
"pages": [{
"elements": [
{ "type": "text", "content": "Hello, JDF",
"heading": 1, "style": "h1",
"position": { "x": 0, "y": 5 },
"width": 166 },
{ "type": "list", "listType": "unordered",
"items": [
{ "content": "Just JSON" },
{ "content": "Diffable" },
{ "content": "Editable anywhere" }
],
"position": { "x": 0, "y": 25 }, "width": 166 }
]
}]
}
Double-click any paragraph, heading, list item, table cell, image — inline editor opens for that element only. Enter commits. Auto-saves to disk in 150ms.
Mouse over any element → floating bar appears with ↑ Move up · ↓ Move down · ⧉ Duplicate · × Delete. No right-click, no menu hunting.
Insert bar at the top of every page. Click to add: text, rich text, list, table, shape, image, collapsible section, auto-generated TOC.
Toggle View ↔ JSON. Two-way bound. Edit JSON directly with Cmd+S — visual render follows. Edit visually — JSON updates live.
Drag a PDF on the viewer. Every text run keeps its position, font, weight, color, opacity, link. Tables become real table elements — headers, rows, column widths, alignment — rebuilt from page geometry (120/120 tables, 99.3% of cells exact on the benchmark corpus). Embedded images extracted. Vector shapes preserved with fills and strokes. Looks identical to the original — and it's editable.
Round-trip back to .pdf. Honors page size, orientation, text colors, real TOC, embedded images. Renders text/list/table/collapsible/shape.
Open .md for a continuous-scroll, GitHub-style render with full GFM. Toggle to paged JDF view. Cmd+F highlights matches inline.
⌘Z · ⌘⇧Z — 100-step history covering every text edit, structural change, and JSON commit. ⌘N for a new window — compare two documents side-by-side.
Draft-07 schema. jdf validate file.jdf reports path-level errors with Ajv. CI on three OSes (macOS, Linux, Windows) on every PR.
git diffJSON.stringify(doc)grep, jq, ripgrepEvery JDF document has these top-level fields:
| Field | Required | What it does |
|---|---|---|
$jdf | yes | Format version (semver) |
meta | yes | title, author, page size, margins, language |
styles | no | Reusable named style definitions |
resources | no | Embedded fonts and base64 images |
header · footer | no | Repeating header/footer with template vars or full element trees |
pages | yes | Array of pages, each with its own elements |
Element types: text (with heading 1-6), richtext, image, table, list, shape, collapsible, toc.
Full schema: spec/jdf-schema.json ·
Working example: hello-world.jdf.
brew tap uurtech/jdf
brew install jdf
Upgrade later with brew upgrade --cask jdf. Homebrew strips the macOS quarantine attribute automatically — the app launches on first run.
No Homebrew? Download the .dmg from the latest release, drag JDF Reader.app into /Applications, then either right-click → Open once, or run:
xattr -cr "/Applications/JDF Reader.app"
open "/Applications/JDF Reader.app"
The build is unsigned — these commands clear the macOS quarantine flag added to downloads. The app itself is unchanged.
Grab a build from the latest release — .deb, .AppImage, .rpm, .msi, .exe all built by GitHub Actions on every tag.
git clone https://github.com/uurtech/jdf.git
cd jdf
pnpm install
pnpm tauri build
Requires Node 20+, pnpm 9+, Rust stable.