Tabulus is a modular framework for digitizing scientific review tables into structured, citation-aware data. It starts from scientific PDFs, reconstructs reference-containing tables, extracts the paper bibliography, and resolves table citations through inspectable CLI stages and filesystem artifacts.
The rebuilt library currently runs through Stage 6 paper-level scholarly
reference resolution. Stage 7 resolved CSV export and a single end-to-end
tabulus run command are planned.
- PDF profiling and table detection with MinerU.
- Canonical table crops for reproducible adapter comparisons.
- Table reconstruction through a registry of OCR, table-structure, and document vision-language-model adapters.
- Deterministic reference-table classification and table-cell-to-bibliography matching.
- GROBID-backed bibliography extraction from the original PDF.
- Conservative Stage 6 scholarly reference resolution using deterministic evidence, Crossref/CORE metadata, and bounded LLM adjudication.
- Native table-reconstruction evaluation with Relative Mapping Similarity (RMS).
PDF
├─ Stage 1: profile PDF and export canonical table crops
│ └─ Stage 2: reconstruct tables
│ └─ Stage 3: classify reference-containing tables
└─ Stage 4: extract bibliography from the original PDF
Stage 3 selected tables + Stage 4 bibliography
└─ Stage 5: match table citations to bibliography positions
└─ Stage 6: resolve paper-level scholarly identities
└─ Stage 7: resolved CSV export (planned)
Stage 4 is a parallel PDF-level branch. Stage 6 resolves each linked bibliography entry once per paper, not once per table cell or reconstruction adapter.
git clone https://github.com/sciknoworg/tabulus.git
cd tabulus
python -m pip install -e ".[dev]"
tabulus --helpFor environment-specific setup, see the Windows/CPU, GPU server, and Python library installation guides.
PDF="/path/to/paper.pdf"
ARTIFACT_ROOT="/path/to/tabulus-artifacts/paper"
CROP_ROOT="/path/to/tabulus-output/table-crops/paper"
RECONSTRUCTION="$CROP_ROOT/reconstructions/tesseract-tatr"# Stage 1: PDF profiling and canonical crop export
tabulus profile --pdf "$PDF" --backend pipeline --method auto
# Stage 2: table reconstruction
tabulus reconstruct-tables \
--crops "$CROP_ROOT" \
--adapter tesseract-tatr \
--device cpu
# Stage 3: reference-table classification
tabulus classify-reference-tables \
--reconstruction "$RECONSTRUCTION"
# Stage 4: bibliography extraction
tabulus extract-bibliography \
--pdf "$PDF" \
--out "$ARTIFACT_ROOT" \
--grobid-url http://localhost:8070
# Stage 5: table-cell to bibliography-position matching
tabulus match-references \
--selected "$RECONSTRUCTION/selected_reference_tables.json" \
--bibliography "$ARTIFACT_ROOT/references/bibliography.json"
# Stage 6: paper-level scholarly reference resolution
tabulus resolve-references \
--bibliography "$ARTIFACT_ROOT/references/bibliography.json" \
--reference-matches "$RECONSTRUCTION/references/reference_matches.json" \
--out "$ARTIFACT_ROOT"Stage 6 requires Crossref, CORE, and OpenAI-compatible LLM configuration. The full option list and credential handling are documented in the Stage 6 tutorial.
| Stage | Main artifact |
|---|---|
| Stage 1 | tables_index.json and canonical crop images |
| Stage 2 | native/, parsed/, predictions/, batch_summary.json |
| Stage 3 | reference_table_classification.json, selected_reference_tables.json |
| Stage 4 | references/bibliography.json |
| Stage 5 | references/reference_matches.json |
| Stage 6 | references/reference_resolution.json |
| Stage 7 | resolved CSV export, planned |
See the Data Contracts for exact schemas and filesystem layout.
TabulusBench is the companion benchmark dataset used by the tutorials and evaluation documentation. It contains scientific PDFs, annotated reference-containing table crops, table gold CSVs, and bibliography gold for a subset of papers.
The documentation uses P4 as the canonical one-paper worked example. Benchmark
gold material is read-only; Tabulus runs should write their generated artifacts
outside the gold directories.
python -m pip install -e ".[dev]"
pytest
git diff --checkThe supported public library lives in src/tabulus. Retained legacy and
research directories are preserved for provenance, but they are not the current
public API unless the documentation says so explicitly.
Citation information will be added after the associated thesis or software release is published.
Tabulus is released under the MIT License.
