Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
-
Updated
Aug 31, 2026 - Python
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
A system for agentic LLM-powered data processing and ETL
Read and extract text and other content from PDFs in C# (port of PDFBox)
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
A curated list of resources for Document Understanding (DU) topic
Open-source platform for extracting structured data from documents using AI.
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.
Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser
This repository provides trainοΌtest code, dataset, det.&rec. annotation, evaluation script, annotation tool, and ranking.
[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.
Code for the paper "PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks" (ICPR 2020)
AssemblyLine 4: File triage and malware analysis
Official PyTorch implementation of LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding (ACL 2022)
Pandora is an analysis framework to discover if a file is suspicious and conveniently show the results
A package for parsing PDFs and analyzing their content using LLMs.
ππ Parse, extract, and analyze documents with ease ππ
Canary Detection
RObust document image BINarization
To associate your repository with the document-analysis topic, visit your repo's landing page and select "manage topics."