Xberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.
| Knob | What it controls | Values | Default |
|---|---|---|---|
--format |
How the CLI prints the result | text, json, toon |
text (extract), json (batch) |
--content-format |
How extracted content is rendered inside result |
plain, markdown, djot, html, json, doctags |
plain |
--token-reduction |
Strip whitespace / boilerplate for LLM contexts | off, light, moderate, aggressive, maximum |
off |
--format json returns an envelope wrapping the ExtractedDocument — the
document lives under .result for extract and under .results[] for
batch, with content, metadata, tables, and images as fields of
that nested document. --format text prints just content.
--content-format is what shows up inside that content field.
Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│ --format text --content-format markdown
├── Vector store / RAG indexer
│ --format json --content-format markdown
│ (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│ --format json --content-format plain
│ (cleanest text + structured metadata)
├── Human review / archival
│ --format text --content-format markdown
├── HTML re-rendering / web display
│ --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│ --format json --content-format djot
└── Token-budget-constrained pipeline
--format text --content-format plain
(drops markup; add --token-reduction moderate for further savings)
Feed a PDF directly into an LLM:
xberg extract paper.pdf --content-format markdown
Index a corpus into a RAG store with tables and headings preserved:
xberg batch docs/*.pdf --format json --content-format markdown \
| jq -c '.results[] | {content: .content, tables: .tables}'
Strip a file to bare text for a token-tight summarizer:
xberg extract long.pdf \
--content-format plain \
--token-reduction moderate
Pull metadata only, ignore content:
xberg extract file.pdf --format json | jq '.result.metadata'
markdown as the content format. It is the best
compromise across LLMs, RAG, and human review, and Xberg has the
most faithful renderer for it.plain only when downstream cannot tolerate any markup.djot only if you're already in a djot/pandoc pipeline.html only when re-rendering for the web.json for a heading-driven content tree, or doctags for Docling-compatible output.--token-reduction collapses whitespace, strips repeated headers/footers,
and trims boilerplate. It composes with any --content-format:
off (default), light, moderate, aggressive, maximum.Use moderate as a safe starting point for LLM context windows. maximum
is lossy — verify before relying on it.
See references/cli-reference.md for the full flag set and
references/configuration.md for the equivalent output_format and
token_reduction keys in xberg.toml.