Cross-provider AI code review for Claude Code — evidence-based confidence scoring with Codex, Gemini & Claude
-
Updated
Sep 6, 2026 - Shell
Cross-provider AI code review for Claude Code — evidence-based confidence scoring with Codex, Gemini & Claude
Uncertainty based selection of compatible inputs
Extract structured data from any document — PDF, DOCX, HTML, CSV, plain text — using LLMs with Pydantic schema validation, per-field confidence scores, and source grounding.
Open-source LLM evaluation engine with statistical confidence scoring
Runtime reliability intelligence designed specifically for OpenClaw frameworks and agents.
Zero-Noise utilities for safer product research and review signal analysis.
Multi-agent AI task delegation architecture for n8n: orchestrator routes natural-language commands to specialist agents with confidence scoring and human-in-the-loop gates.
RAG-based document intelligence system with semantic retrieval, grounded Q&A, citations, confidence scoring, and guardrails.
Research-grade Self-Correcting RAG agent built with LangGraph that retrieves knowledge, generates answers, evaluates grounding/relevance/completeness, and iteratively self-improves with confidence scoring and memory.
Deterministic structured extraction from noisy LLM/OCR output. Zero LLM round-trips, microsecond latency, confidence score on every result. msgspec · Pydantic · dataclasses.
System that aggregates outputs from multiple Large Language Models (GPT-4, Claude-3, custom models) to generate reliable, high-confidence results through consensus-based reasoning evaluation. Demonstrates sophisticated AI orchestration with 92.7% accuracy improvement over single-model.
AI-powered concierge that normalises guest messages from WhatsApp, Booking.com, Airbnb, Instagram and direct channels, drafts a reply with Claude, and routes responses through a deterministic confidence-scoring pipeline. Built with FastAPI + Claude Sonnet 4.
Verification system that catches coding agents falsely claiming task completion. Runs 4 parallel checks (file integrity, test quality, scope narrowing, optional LLM judge) over task+claim+diff and returns a weighted 0-100 confidence score with evidence.
Smart Document Conversion for the AI Era - CPU-only, fast, with confidence scoring. Converts PDF, DOCX, PPTX, HTML, EPUB to Markdown, JSON, HTML, Text.
Enterprise-grade Confidence-Driven State Reconciliation Platform leveraging Kafka Streams, Redis, PostgreSQL, and Spring Boot to preserve competing truths, compute evidence-backed consensus, deterministic replay, and operational analytics.
RAG service with a policy engine that decides whether to answer, ask a clarifying question, or refuse before generation ever runs. Lexical retrieval, SQLite-backed chunk storage, page-level citations, and a mocked generation layer shaped for a drop-in real LLM call.
MFGC confidence scoring and safety gates for AI agents. Zero dependencies.
Rigorous LLM research in two layers: a multi-agent dialectic synthesis prompt + an agent-loadable research skill. Hierarchy of evidence, verbatim citation gates, no fabrication.
7-axis weighted confidence function for AI output quality. Evidence, reasoning, calibration, source, domain, coherence, meta. (Companion repo — architecture and interface only. Source lives in the private engine.)
To associate your repository with the confidence-scoring topic, visit your repo's landing page and select "manage topics."