The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
-
Updated
Sep 1, 2026 - Java
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
基于 Spring Boot 4.1、Java 25、Spring AI 2.0、React、PostgreSQL/pgvector、Redis 和 RustFS 构建的开源 AI 面试平台,支持简历智能分析、模拟面试、语音面试和知识库 RAG。
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
Elasticsearch File System Crawler (FS Crawler)
Spark-Crawler: Apache Nutch-like crawler that runs on Apache Spark.
A cross-platform command line tool for parallelised content extraction and analysis.
Use the Java Tika text extraction library on the .NET platform
Code for Machine Learning with TensorFlow: 2nd Edition Published by Manning Publications
Apache Tika bindings for PHP: extract text and metadata from documents, images and other formats
Tika-Similarity uses the Tika-Python package (Python port of Apache Tika) to compute file similarity based on Metadata features.
Interactive Image similarity and Visual Search and Retrieval application
RADiX overlay on Mnemosyne 1.11.0: Apache Solr 10, Apache Tika, HuggingFace TrOCR/Donut. Ingest in place, MIME/EXIF + OCR into Solr. Vue OPSUI at /opsui/. ImageSpace (search, CLIP similar, fg/bg, IQR) on :8090.
Extract and Visualize location from any file
R Interface to Apache Tika
Distributed, fault tolerant batch processing for Natural Language Applications and Search, using remote partitioning
Geographic Place, Date/time, and Pattern entity extraction toolkit along with text extraction from unstructured data and GIS outputters.
To associate your repository with the tika topic, visit your repo's landing page and select "manage topics."