A robust Python tool to extract complex HTML tables and convert them into clean Markdown, JSON, or CSV formats.
Try it in your browser — an interactive Colab notebook, nothing to install.
Unlike simple formatters, this library features a Grid Logic Solver that correctly interprets rowspan and colspan attributes, normalizing complex HTML grids into perfectly aligned data structures.
There is an existing package called table2md on PyPI.
-
The existing package is a formatter. You feed it Python lists/dicts, and it draws a Markdown table.
-
This package is an extractor and parser. It takes raw HTML source code, uses
BeautifulSoupto parse the tags, mathematically resolves complex cell spans (rowspans/colspans), and builds an internal representation before exporting to Markdown.
Most parsers turn this <td rowspan="2"> into a misaligned mess:
| Header | Value |
|---|---|
| Spanned | Row 1 |
| Row 2 | | <-- Everything shifts!html-table-rescuer uses a grid solver to correctly normalize the matrix:
| Header | Value |
|---|---|
| Spanned | Row 1 |
| dito (Spanned) | Row 2 |Our pipeline ensures that complex HTML structures are safely converted without data loss or misalignment:
graph TD
n1["HTML Input"] --> n2["BeautifulSoup Parser"]
n2 --> n3["Grid Logic<br>(Rowspan/Colspan Solver)"]
n3 --> n4["ParsedTable Data Object"]
n4 --> n5["Markdown Export"]
n4 --> n6["JSON/CSV Export"]
n4 --> n7["LangChain/LlamaIndex/Haystack Wrappers"]
pip install html-table-rescuer
# With framework integrations
pip install "html-table-rescuer[langchain]"
pip install "html-table-rescuer[llamaindex]"
pip install "html-table-rescuer[haystack]"from html_table_rescuer import TableParser
html_content = """
<table border="1">
<tr>
<th colspan="2">Header</th>
</tr>
<tr>
<td>Data 1</td>
<td>Data 2</td>
</tr>
</table>
"""
# Initialize parser with your HTML
parser = TableParser(html_content)
# Parse all tables in the HTML
tables = parser.parse()
if tables:
table = tables[0]
# Export to Markdown (perfect for LLM context windows)
print(table.to_markdown())
# Export to JSON
# print(table.to_json())
# Export to CSV
# print(table.to_csv())The package installs an html-table-rescuer command that reads from a file, a URL, or stdin:
# From a file
html-table-rescuer page.html
# From a URL
html-table-rescuer https://example.com/page.html
# From stdin (pipe or '-')
curl -s https://example.com/page.html | html-table-rescuer
cat page.html | html-table-rescuer - --format jsonOptions:
| Option | Description |
|---|---|
--format, -f |
Output format: markdown (default), json, csv |
--strategy, -s |
Rowspan fill strategy: fill_dito (default), repeat, empty |
--dito-prefix |
Prefix used by the fill_dito strategy |
--table, -t |
Extract only the table with this index (0-based) |
--output, -o |
Write to a file instead of stdout; with CSV and multiple tables, writes name_1.csv, name_2.csv, … |
--no-links / --no-bold / --no-italic |
Strip the respective inline formatting |
--parser |
BeautifulSoup backend (lxml default, or html.parser) |
Tables with a multi-row header — a <th rowspan="2"> next to a grouped
<th colspan="3">, common in discographies and financial reports — get their
header levels merged, so the second header row never ends up as data:
table = TableParser(html).parse()[0]
table.headers # ['Title', 'Peak positions - US', 'Peak positions - AUS']
table.caption # 'Studio albums' (from <caption>, None if absent)The separator is configurable via ParseConfig(header_separator=" / "). The
caption is prepended to the Markdown as a bold line — turn it off with
table.to_markdown(include_caption=False).
Every table becomes its own document, so retrieval never splits a table in half.
from html_table_rescuer.integrations.langchain import HTMLTableRescuerLoader
docs = HTMLTableRescuerLoader("page.html").load()
print(docs[0].page_content) # Markdown table
print(docs[0].metadata) # {'source': 'page.html', 'table_index': 0, 'parser': 'html_table_rescuer'}from html_table_rescuer.integrations.llamaindex import HTMLTableRescuerReader
docs = HTMLTableRescuerReader().load_data("page.html")Works as a file_extractor in SimpleDirectoryReader, so entire folders are handled for you:
from llama_index.core import SimpleDirectoryReader
reader = SimpleDirectoryReader(
"./docs",
file_extractor={".html": HTMLTableRescuerReader()},
)
docs = reader.load_data()from html_table_rescuer.integrations.haystack import HTMLTableRescuerConverter
converter = HTMLTableRescuerConverter()
docs = converter.run(sources=["page.html"])["documents"]Drops straight into a pipeline — unreadable sources are skipped with a warning
rather than failing the run, and ParseConfig survives pipeline serialization:
from haystack import Pipeline
pipe = Pipeline()
pipe.add_component("converter", HTMLTableRescuerConverter())
result = pipe.run({"converter": {"sources": ["page.html"]}})All three accept a ParseConfig to control rowspan handling and inline formatting:
from html_table_rescuer import ParseConfig, RowspanStrategy
config = ParseConfig(rowspan_strategy=RowspanStrategy.REPEAT_VALUE)
docs = HTMLTableRescuerReader(config=config).load_data("page.html")- HTML parsing via
BeautifulSoup - Recursive inline-tag formatting (keeps links, bold, and italic tags alive even if nested in divs)
- Complex
rowspanandcolspangrid resolution (using flexible strategies like filling cells with "dito" to preserve context for LLMs) - Clean Markdown export
- Data Exports: JSON and CSV serialization from the
ParsedTableobject - CLI:
html-table-rescuercommand with file/URL/stdin input and Markdown/JSON/CSV output - Stacked headers: multi-row headers are merged into one (
Peak positions - US) instead of leaking into the data - Table captions:
<caption>is extracted intoParsedTable.caption, prepended to the Markdown, and passed to every integration's metadata - AI Integrations: Ready-to-use
LangChainDocument Loader,LlamaIndexReader (works withSimpleDirectoryReader), andHaystackConverter - Robust against broken real-world HTML: invalid
colspan/rowspanvalues, HTML comments, and oversized spans are handled gracefully
Measured on 30 real Wikipedia tables that use rowspan/colspan — full numbers,
methodology and reproduction steps in docs/BENCHMARKS.md.
| Result | |
|---|---|
dito fill overhead |
+1.5% tokens vs leaving cells empty (only 3% of cells are span continuations) |
| Markdown vs JSON output | JSON costs +78% tokens for the same tables |
Speed vs pandas.read_html |
pandas is 2.9x faster — use it if your HTML is clean |
| Malformed markup | survived 8/8 cases; pandas raises ValueError on 2 (colspan="abc", colspan="2.5") |
| Span resolution | identical grid to pandas on 27/30; the 3 remaining differences are not about spans |
Contributions are welcome! Please feel free to submit a Pull Request.
This project is licensed under the MIT License.
** Read the 1. blog:** How to Stop LLMs from Hallucinating on Complex HTML Tables ** Read the 2. blog:** How to Stop LLMs from Hallucinating on Complex HTML Tables