Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
-
Updated
Aug 30, 2026 - Python
Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
A Python CLI to test, benchmark, and find the best RAG chunking strategy for your Markdown documents.
One library to split them all: Sentence, Code, Docs. Chunk smarter, not harder — built for LLMs, RAG pipelines, and beyond.
Steadlith reuses unchanged RAG chunks with content-defined identities, cache-aware planning, and transactional indexing.
Production-ready Snowflake RAG system with type-specific chunking
A practical guide to 6 document chunking strategies for RAG and LLM applications — Document, Fixed-Size, Recursive, Sentence, Semantic, and Agentic chunking with working code and plain-English explanations.
A lightweight Python library for metadata-rich document chunking in Retrieval-Augmented Generation (RAG) workflows. It leverages Azure AI Document Intelligence to enhance chunking by retaining hierarchical structure, page numbers, and bounding boxes for seamless integration with PDF viewers.
A Controlled Natural Language (CNL) for AI designed to "minify" language and make AI context denser.
"My complete LangChain learning journey — from basics to advanced RAG, LCEL, LangGraph, LangServe, LangSmith with hands-on code examples."
building a CPU-Only "PDF Q&A System" using hugging face, chromaDB vector search, and Python
Astra Vector DB on Python-paketti, joka tallentaa dokumentteja DataStax Astra DB -vektoritietokantaan ja suorittaa semanttista hakua.
FastAPI service for document chunking and sentence-transformer embeddings for RAG, semantic search, and vector database ingestion.
This repository provides a fully modular implementation of a Retrieval-Augmented Generation (RAG) pipeline tailored for Italian legal-domain documents.
Smart text chunking tool for RAG systems. Splits long texts into sentence-based chunks with ~10%-15% overlap for better context retention. Runs fully in-browser with a clean UI and copyable outputs.
KChunker is a lightweight, ultra-fast document parsing and chunking engine designed for RAG systems. It intelligently structures native/scanned PDFs, Excel files, Word documents, and email trails by preserving layout hierarchy, extracting tables, and generating dense vector embeddings for local search databases (ChromaDB and FAISS)
To associate your repository with the document-chunking topic, visit your repo's landing page and select "manage topics."