Production grade Document AI workflows run on Tensorlake
-
Updated
Jun 21, 2026 - Python
Production grade Document AI workflows run on Tensorlake
Agent skill for pixel-grounded chart data extraction
Python library for extracting content from PowerPoint files including embedded charts and SmartArt. Built for RAG and document processing pipelines.
CUDA-accelerated PDF -> Markdown/HTML converter using Docling + IBM Granite Vision chart extraction
A complete end-to-end pipeline for extracting structured data from chart and graph images
doc-textify: offline, CPU-only PDF/image to Markdown & LLM-ready text converter. OCR with CJK normalization, layout recovery, two-column reading order, table/chart/formula extraction, RAG-ready chunking — no vision LLMs, no GPU, no cloud.
Converters where figures survive. DOCX/XLSX to Markdown via native OOXML chart data: real numbers, OCR/VLM are only optional. CLI, MCP server, Docker image, and .mcpb bundle included.
Extract charts, figures, and tables from academic PDFs for AI agent analysis
Read the numbers out of graphics a page only shows as pictures
To associate your repository with the chart-extraction topic, visit your repo's landing page and select "manage topics."