RegBot is an open-source tool for the Global Alliance for Genomics and Health Regulatory and Ethics Work Stream (REWS) and cross-border genomic data sharing. It complements the Alliance’s Regulatory & Ethics Toolkit by retrieving GA4GH and related policy provisions against researcher-supplied consent / data-use text and returning citation-grounded JSON for DPO, IRB, and DAC review—not compliance rulings or legal advice.
Release status: 0.1.0 release candidate. Implementation, corpus rebuild, the
41-query contributor-labelled benchmark, API/UI checks, and local Ollama validation are
complete. Independent mentor review of the gold set is intentionally still pending; the
scheduled benchmark remains informational until that review. Frontend and applicable Python
dependency audits are clean after the 2026-08-10 security update; the documented Chroma
server-only exception does not apply to RegBot's embedded client (see
docs/RELEASE_CHECKLIST.md).
docs/DESIGN.md— architecture, data model, evaluation plan (GSoC design doc)docs/eval_results.md— measured retrieval benchmark: metrics, tuning runs, threats to validitydocs/corpus_manifest.yaml— regulatory corpus inventory (85 documents;content_typemarks each asprimary,translation, orsummary)docs/CORPUS_SCOPE.md— inclusion criteria, regional coverage, exclusions, and the rule for reopening the corpusexamples/eval/gold_ga4gh.yaml— retrieval gold set (drafted; awaiting mentor review)examples/DEMO.md— local end-to-end demo
- Ingest policy PDFs or
.txtfiles into a local Chroma store plus a JSON manifest. Chunks carrysource,page,category,document_id,jurisdiction,framework,content_type(primarysource text, non-authoritative/referencetranslation, or contributorsummary— each badged in the UI), andsection(every heading the chunk spans) where the source is line-structured. - Hybrid retrieval: exact cosine embedding search + BM25, fused by reciprocal rank (best-channel
maxby default;REGBOT_FUSION=sumfor classic additive RRF).jurisdiction/framework/categoryfilters scope the candidate search itself, not just the output. Ranking is deterministic — identical inputs give identical results across runs. - Compliance pass: JSON-mode LLM via Ollama by default (e.g.
llama3, configurable withREGBOT_OLLAMA_MODEL). SetREGBOT_LLM_PROVIDER=openaiandOPENAI_API_KEYto use OpenAI instead. If no LLM is reachable, the keyword fallback keeps only recommendations with lexical support from a retrieved chunk and escalates unsupported rows for human review. - Web UI (recommended): FastAPI + Next.js in
frontend/— see Run the web UI below. - Streamlit UI (legacy): upload + paste flows (
src/streamlit_app.py). - CLI:
python -m src.main …(see below). - Citation grounding (programmatic): Each
recommendations[]item must be{ "text": "...", "evidence_chunk_ids": ["..."] }with ids taken only from retrieved chunks; optionalcitations[]must also respect the same allow-list. Failed LLM checks trigger automatic rewrite requests with the allow-list; both LLM and offline-fallback recommendations pass token-overlap filtering (REGBOT_MIN_TOKEN_OVERLAP). - Reviewable evidence: every recommendation carries
evidence[]with the source document, page, a verbatim quote from the cited chunk, the jurisdiction, and agovernance_hintnaming the body that normally reviews that scope (DPO / IRB / DAC). Quotes are copied, never generated. - Fail-safe escalation: when retrieval is thin, grounding fails, or the overlap filter drops any recommendation, the report sets
needs_human_reviewwith areview_reason(weak_retrieval/grounding_failed/low_overlap) instead of presenting incomplete output as an answer. - Retrieval benchmark:
benchmarksubcommand scores retrieval against a gold set (Recall@k / Precision@k / MRR) and can gate CI via--min-recall. Results:docs/eval_results.md. - PDF eval harness:
evalsubcommand ingests a real GA4GH PDF and prints retrieval hits for built-in or custom queries (manual inspection; usebenchmarkfor scored evaluation).
- Prerequisites: Python 3.10–3.13 (the range
pyproject.tomlaccepts; CI runs 3.11, which is the tested one). Python 3.14 is not supported yet for the full stack (native wheels for parts of the ML/Chroma toolchain often lag). Running the web UI also needs Node 20.9+ forfrontend/(CI uses Node 22). - Confirm that the interpreter used to create the environment is in the supported range,
then create the environment and install dependencies. The example names Python 3.11 to
match CI;
python3.10,python3.12, orpython3.13are also valid:
python3.11 --version
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
# Also install test/lint/type-check dependencies when developing or verifying a release.
python -m pip install -r requirements-dev.txt-
Configure environment variables:
- Export variables in your shell (recommended)
- If you use a local
.env, keep it private and do not commit it
-
Web login and roles (required for FastAPI / Next.js and Streamlit)
RegBot intentionally ships with no default password. The login page allows a public read-only user session by default; configure an administrator account for corpus management (account passwords require 12+ characters). A random in-memory session secret is generated automatically for local single-process use:
export REGBOT_ADMIN_USERNAME='admin'
export REGBOT_ADMIN_PASSWORD='replace-with-a-long-admin-password'Anyone can choose Continue as public user to browse, retrieve, run consent checks, and chat,
but public-user sessions receive HTTP 403 for write operations and custom-store access. admin
can upload/reset corpus content and select a custom store. Set
REGBOT_ALLOW_GUEST_VIEWER=0 to require named accounts for all access. Unauthenticated
API requests receive HTTP 401. Sessions use a signed HttpOnly, SameSite=Lax cookie and
expire after eight hours by default. Configure a stable 32+ character
REGBOT_SESSION_SECRET for multi-worker deployments or sessions that must survive a
server restart. For HTTPS deployment, also set REGBOT_COOKIE_SECURE=1.
-
LLM (default: local Ollama)
Install Ollama, runollama pull llama3(or another tag you set inREGBOT_OLLAMA_MODEL), and keep the daemon running (ollama serveorbrew services start ollamaon macOS). NoOPENAI_API_KEYis required for this path. -
Embeddings (first ingest)
The embedding model is downloaded from Hugging Face on first use. If downloads are slow or fail, try a longer timeout (HF_HUB_DOWNLOAD_TIMEOUT, seconds) or a mirror (REGBOT_HF_ENDPOINT=https://hf-mirror.com— setsHF_ENDPOINTfor the Hub client). -
Ingest a policy file into
./data/regbot_store(use--resetwhen reloading the same corpus):
python -m src.main ingest --path path/to/policy.pdf --reset- Batch ingest from the corpus inventory (source files live under
data/corpus/; seedocs/corpus_manifest.yaml):
python -m src.main ingest-manifest --dry-run # show documents absent from this local store
python -m src.main ingest-manifest # add documents absent from this local store
python -m src.main ingest-manifest --reset # clear and rebuild all 85 documents
# Build only P0 in a separate store; --store must precede the subcommand.
python -m src.main --store ./data/regbot_p0_store ingest-manifest --tier P0 --resetThe store's Chroma vectors are git-ignored and regenerated by this command; the text
manifest.json is tracked as the BM25 corpus and citation-audit record. The inventory's
ingested_at values are audit timestamps; incremental ingestion checks the active local
store instead of treating those timestamps as proof that vectors exist on this machine.
- Refresh the 83 reproducible source files from their publishers (GA4GH, EUR-Lex and regional legislative or government sites):
python tools/fetch_corpus.py --list # show targets
python tools/fetch_corpus.py --check # validate what is already on disk
python tools/fetch_corpus.py # refresh everythingEach fetch is validated before it is written — a target declares required phrases and a minimum length, so a publisher that serves a table of contents instead of the statute is rejected rather than ingested.
- Check a consent / data-use text file:
python -m src.main check --consent path/to/consent.txt
# Jurisdiction filters are repeatable; this searches the union of SG and GA4GH.
python -m src.main check --consent path/to/consent.txt \
--jurisdiction SG --jurisdiction GA4GH --top-k 8
# Framework filters are repeatable too and intersect with any jurisdiction scope.
python -m src.main check --consent path/to/consent.txt --framework GA4GH --top-k 8
python -m src.main status- Run the web UI (FastAPI + Next.js) from the repo root:
# Terminal 1 — API (repo root, venv active)
uvicorn src.api.app:app --reload --port 8000
# Terminal 2 — frontend
cd frontend && npm ci && npm run devOpen http://localhost:3000/login. The Next.js dev server
proxies /api/* and /health to the API on port 8000; set REGBOT_API_URL before
npm run dev if the API is not on http://127.0.0.1:8000.
A fresh clone ships manifest.json but not the Chroma vectors (git-ignored), so run
ingest-manifest --reset once before the UI can retrieve anything. The Corpus tab lists
the inventory and Browse can inspect portable manifest chunks, but Check and unscoped Chat
cannot retrieve evidence until this rebuild finishes.
The repository includes render.yaml for two Render web services:
regbot-apibuilds the Python environment, rebuilds the 85-document Chroma index, and starts FastAPI on0.0.0.0:$PORT.regbot-webbuilds and starts Next.js fromfrontend/, proxying/api/*and/healthto FastAPI.
Create or synchronize a Render Blueprint from the repository, then provide the generated
FastAPI public URL as the regbot-web service's REGBOT_API_URL. The API Blueprint
generates REGBOT_SESSION_SECRET, enables secure cookies and public read-only access, and
prompts for OPENAI_API_KEY; never commit these secrets.
The default deployment rebuilds the immutable baseline index during each build. Runtime
uploads are ephemeral. If administrator uploads must survive redeploys, attach a persistent
disk to the API service, set REGBOT_STORE to a directory on that disk, and initialize the
index from the API start command when the disk is empty. Render disks are not available to
build or pre-deploy commands.
After deploying, verify:
curl -fsS https://YOUR-API.onrender.com/health
# Expected: {"status":"ok"}Then enter through the Next.js public-user login, confirm that the sidebar says
Retrieval index ready, and run one consent check that returns at least one cited chunk.
- Run the legacy Streamlit UI:
python -m streamlit run src/streamlit_app.py- End-to-end sample (synthetic policy + consent under
examples/):
python examples/run_demo.pyEvaluate retrieval on a real GA4GH PDF (use --reset when reloading the same corpus):
python -m src.main eval --pdf path/to/ga4gh_policy.pdf --reset --top-k 8Use your own query list (one line per query):
python -m src.main eval --pdf path/to/ga4gh_policy.pdf --reset --queries-file examples/eval/queries_ga4gh.txtOptionally append a full compliance JSON report for a consent file:
python -m src.main eval --pdf path/to/ga4gh_policy.pdf --reset --consent path/to/consent.txtScore retrieval against the gold set (Recall@k / Precision@k / MRR):
python -m src.main benchmark --gold examples/eval/gold_ga4gh.yamlWrite a Markdown table, or optionally fail below a recall threshold (CI gate):
python -m src.main benchmark --markdown /tmp/regbot-benchmark.md
python -m src.main benchmark --min-recall 0.70The threshold applies to macro recall at the largest requested --ks value (default:
recall@8). Do not use it as a required release gate until the gold labels are independently
reviewed. Run python -m src.main --help or append --help to any subcommand for the full
CLI reference.
python -m pytest -qREGBOT_LLM_PROVIDER:ollama(default) — local LLM via Ollama’s OpenAI-compatible HTTP API (no OpenAI key). Set toopenaito use OpenAI’s hosted API instead.OPENAI_API_KEY: Required only whenREGBOT_LLM_PROVIDER=openai. Model:REGBOT_LLM_MODEL(defaultgpt-4o-mini).REGBOT_OLLAMA_MODEL: Tag known to Ollama (defaultllama3). Examples:llama3,mistral,mistral:latest.REGBOT_OLLAMA_BASE_URL: Ollama HTTP host only (defaulthttp://127.0.0.1:11434);/v1is appended automatically for the OpenAI-compatible routes.REGBOT_OLLAMA_API_KEY: Sent as the Bearer/API key to Ollama’s shim (defaultollama; ignored by Ollama).REGBOT_STORE: On-disk store directory (default./data/regbot_store).REGBOT_SESSION_SECRET: Optional random session-signing secret of at least 32 characters. When unset, RegBot creates a per-process secret, so sessions end on restart and cannot span multiple workers. Configure it for production or multi-worker deployments; rotate it to invalidate all sessions.REGBOT_ADMIN_USERNAME/REGBOT_ADMIN_PASSWORD: Administrator credentials; username defaults toadminwhen its password is set, and the password must contain at least 12 characters.REGBOT_VIEWER_USERNAME/REGBOT_VIEWER_PASSWORD: Optional named read/review credentials. The environment-variable and internal role names retainVIEWERfor backward compatibility; the UI calls this access level user. The username defaults toviewerwhen its password is set.REGBOT_ALLOW_GUEST_VIEWER: Allow the login page's account-free, read-only viewer entry (default1). Set to0to require a configured account for every user.REGBOT_SESSION_HOURS: Signed-session lifetime in hours (default8; clamped to1–168).REGBOT_COOKIE_SECURE: Set to1when serving over HTTPS so the login cookie is never sent over plaintext HTTP (default0for localhost development).REGBOT_API_URL: Read by the Next.js server (frontend/next.config.ts) to proxy/api/*and/health(defaulthttp://127.0.0.1:8000). Set it when the FastAPI process is on another host or port.REGBOT_EMBEDDING_MODEL: SentenceTransformers model id (defaultsentence-transformers/all-MiniLM-L6-v2).HF_HUB_DOWNLOAD_TIMEOUT: Hugging Face Hub download timeout in seconds (embedding model on first use). The app sets a higher default when unset; increase if you see read timeouts.REGBOT_HF_ENDPOINT: If set, copied toHF_ENDPOINT(e.g.https://hf-mirror.comwhere Hub mirrors are used).HF_HUB_OFFLINE: Set to1to skip Hub access and load the embedding model from the local cache.REGBOT_MIN_TOKEN_OVERLAP: Minimum token recall between each LLM or offline-fallback recommendation and its cited chunk texts (default0.06). Set to0to disable dropping low-overlap rows.REGBOT_SEMANTIC_CANDIDATES/REGBOT_BM25_CANDIDATES: Candidate pool sizes feeding reciprocal rank fusion (defaults12/48). Lexical weighting was measured to beat balanced pools — seedocs/eval_results.md§2.REGBOT_FUSION:max(default) takes each chunk's best channel;sumrestores classic additive RRF. Seedocs/eval_results.md§4c.REGBOT_MAX_CHUNKS_PER_PROVISION: How many chunks of one provision (document + section) may occupy the result list (default2;0disables). A reviewer wants distinct applicable rules, not repeated fragments of one — seedocs/eval_results.md§4g.REGBOT_CHROMA_ANONYMIZED_TELEMETRY: Set to1to enable Chroma client telemetry; default is off (0).REGBOT_OPENAI_MAX_RETRIES: Retries for the OpenAI Python client (used for both OpenAI API and Ollama’s compatible endpoint; default3).
- Core: Python 3, package under
src/regbot/— ingestion, hybrid retrieval, fusion, grounding, evidence, evaluation, jurisdiction and text utilities. - Embeddings:
sentence-transformers+ Hugging Face Hub (minimal file set; ONNX-heavy artifacts skipped where possible). - Vector store: Chroma persistent files under
REGBOT_STORE/chromaplusmanifest.json, which holds chunk text and metadata for BM25 and citation audit. - Retrieval: exact cosine ranking over the stored embeddings (loaded once from Chroma, ranked in process — the ANN index was approximate and made results irreproducible) +
rank-bm25, fused by reciprocal rank.jurisdiction/framework/categoryfilters restrict the candidate pool before the top-N cut. - LLM: Default: Ollama (
llama3orREGBOT_OLLAMA_MODEL) via OpenAI-compatible chat completions + JSON parsing. Optional:REGBOT_LLM_PROVIDER=openaiwithOPENAI_API_KEY. Fallback: keyword heuristic if OpenAI is selected without a key, or after LLM errors (e.g. Ollama not running). - API: FastAPI (
src/api/app.py) — signed-cookie authentication and role-protected corpus, chunk, ingest, check, and chat endpoints behind/api. - UI: Next.js in
frontend/(recommended); Streamlit (src/streamlit_app.py) retained as the legacy single-process option. - Access boundary: Next.js/FastAPI and the legacy Streamlit UI enforce the same roles. API sessions use signed cookies; Streamlit uses its server-managed session. This does not replace operating-system permissions: a user with shell and filesystem access can still run the local CLI directly.
- Optional / post-release: LangChain or LlamaIndex adapters on top of the same stores; larger independently labelled evaluation sets. A cross-encoder is not part of
0.1.0because the final benchmark does not show a provision-recall failure that justifies its model and latency cost.