MinerU#
Official Resources#
Role In Tabulus#
MinerU is the external PDF profiling tool used by the current Tabulus Step 1
workflow. It is a document parsing toolkit for transforming PDFs and other
complex documents into structured Markdown and JSON. In Tabulus, MinerU
performs document/layout processing, table localization, and native table
extraction before Tabulus normalizes the canonical table-crop handoff used by
Step 2 reconstruction. The Tabulus CLI exposes only the MinerU options needed
for the current profiling contract, with CPU-compatible pipeline and
GPU-backed hybrid-engine backends.
For the guided profiling workflow and CLI examples, see Step 1: PDF Profiling.
Tabulus currently supports MinerU as the only PDF profiler:
tabulus profile --profiler mineru ...
MinerU performs document/layout processing, table localization, and native table extraction. Tabulus then:
locates MinerU
*_content_list.jsonfilesresolves MinerU table image paths
normalizes page numbers, bounding boxes, captions, footnotes, and provenance
retains MinerU
table_bodyas a native reconstruction candidateexports canonical table crops and
tables_index.json
The stable handoff for later Tabulus steps is:
tabulus-output/
table-crops/
<paper>/
tables_index.json
images/
Step 2 reconstruction adapters should consume this handoff rather than the complete MinerU-native directory.
MinerU Options Exposed By Tabulus#
tabulus profile exposes a small MinerU-specific surface:
Tabulus option |
MinerU meaning |
|---|---|
|
CPU-compatible MinerU backend. |
|
GPU-backed MinerU backend. |
|
Let MinerU choose text extraction or OCR handling. |
|
Ask MinerU to use native PDF text extraction. |
|
Ask MinerU to use OCR. |
|
Processing effort for |
If hybrid-engine is requested but GPU requirements are not satisfied,
Tabulus reports the reason and falls back to pipeline. Automatic output
paths use the resolved backend name.
Tabulus currently fixes these MinerU settings internally:
table=True
formula=False
image_analysis=False
They are not currently exposed as Tabulus CLI flags.
Native Output#
Tabulus owns only the profiling output root:
<PDF parent>/tabulus-output/mineru/<resolved-backend>/
MinerU owns the document/run hierarchy beneath that root:
tabulus-output/
mineru/
<resolved-backend>/
<paper>/
<MinerU-native run directory>/
images/
<paper>_content_list.json
<paper>_content_list_v2.json
<paper>_layout.pdf
<paper>_middle.json
<paper>_model.json
<paper>_origin.pdf
<paper>.md
images/MinerU-generated image assets, including table images referenced by structured output.
<paper>_content_list.jsonFlat structured content list currently used by the Tabulus table-discovery workflow.
<paper>_content_list_v2.jsonNewer structured representation produced by MinerU.
<paper>_layout.pdfLayout/debugging PDF useful for visually inspecting detected regions.
<paper>_middle.jsonDetailed intermediate parsing representation useful for debugging.
<paper>_model.jsonLower-level/model inference output primarily useful for debugging.
<paper>_origin.pdfMinerU’s copy of the original PDF.
<paper>.mdMinerU’s human-readable reconstructed Markdown representation.
For the detailed file contract, see MinerU Output Files.
Native Run Directory Discovery#
After successful MinerU execution, Tabulus discovers the actual
<MinerU-native run directory> from the generated *_content_list.json.
Tabulus must not construct, flatten, rename, or assume that native run
directory solely from --method.
Validated MinerU 3.4.5 examples:
pipeline/<paper>/auto/
hybrid-engine/<paper>/hybrid_auto/
These are observed MinerU-owned directory names from tested configurations, not universal Tabulus naming rules.
On success, Tabulus writes diagnostic files beside the discovered MinerU output:
mineru_stdout.log
mineru_stderr.log
tabulus_run.txt
If MinerU fails before a native run directory can be identified, diagnostics may be written at the document level instead.
Boundary#
MinerU is the profiling and canonical crop-generation step in the current workflow. It is separate from crop-consuming Step 2 reconstruction adapters, reference-table classification, bibliography extraction, reference matching, DOI resolution, and final resolved CSV export.