A Healthcare AI Benchmark Suite for Evidence Validation, Documentation Integrity, and Eligibility Reasoning
Version 1.0 | Final Portfolio Release | Meredith Reese, RN | July 2026
The Evidence Integrity & Eligibility Determination (EIED) benchmark suite evaluates whether AI systems can determine when available evidence actually supports a healthcare decision.
Rather than rewarding factual recall, polished language, or agreement with a dashboard label, EIED asks whether the evidence is sufficient, current, internally consistent, patient-specific, and authoritative enough to justify the conclusion.
Has the AI earned the right to reach this conclusion?
Claim → Evidence → Confidence → Decision
Every benchmark asks:
- What is the Claim?
- What is the Evidence?
- How confident should the model be?
- What Decision does the evidence justify?
Labels are claims—not evidence.
A record marked Complete, Signed, Reviewed, Attached, Scheduled, Verified, Green, or Closed may still be unsupported. The inverse matters too: a stale hold, preliminary note, or incomplete-looking record may be resolved by later authoritative evidence.
EIED therefore penalizes both:
- unsupported release, approval, inclusion, finalization, application, or closure; and
- unsupported hold, denial, exclusion, escalation, or deferral.
- 12 original healthcare AI benchmarks
- 360 fictional source documents
- 48 submission-ready batch archives
- 12 system/task prompts
- 12 strict grading rubrics
- 12 benchmark design notes
- Domains spanning documentation closure, prior authorization, research eligibility, population health, medical necessity, quality measurement, medication safety, case resolution, coding integrity, longitudinal reconciliation, and clinical guideline applicability
Each benchmark is distributed as a self-contained ZIP archive for direct download.
| # | Benchmark | Primary domain | Core decision |
|---|---|---|---|
| 001 | Discharge Documentation Closure Eligibility Audit | Clinical documentation integrity and care-transition closure | Has the record earned clean/complete closure status? |
| 002 | Prior Authorization Evidence Review | Payer criteria, imaging, specialty medication, and DME authorization | Has the request earned submission, hold, or escalation? |
| 003 | Clinical Trial Eligibility Validation | Oncology research screening and protocol eligibility | Has the candidate earned the right to proceed? |
| 004 | Registry Enrollment Audit | Population health registry and colorectal screening outreach | Has the patient earned inclusion, exclusion, suppression, or outreach? |
| 005 | Medical Necessity Documentation Review — Post-Acute Services | Home health, DME, wound therapy, oxygen, infusion, and outpatient therapy | Has the service earned a medically necessary determination? |
| 006 | Medical Necessity Documentation Review — Diagnostic and Procedural Utilization Management | Imaging, procedures, treatment history, and exception pathways | Has medical necessity been established for this service and patient? |
| 007 | Quality Measure Inclusion Validation | Hypertension quality-measure denominator, exclusion, and numerator review | Has the patient earned inclusion, exclusion, compliance, or hold status? |
| 008 | Medication Safety Verification | High-risk medication release, dispensing, and alert verification | Has the medication plan earned release or withholding? |
| 009 | Case Closure Readiness Audit | Patient safety, complaint, referral, privacy, communication, equipment, and follow-up cases | Has the case earned true resolution and closure? |
| 010 | Clinical Coding Evidence Validation | Diagnosis and principal-diagnosis coding integrity | Has the diagnosis or principal-diagnosis assignment earned final coding? |
| 011 | Longitudinal Evidence Reconciliation | Whole-record reconciliation across hospital, clinic, specialist, rehabilitation, home, pharmacy, and portal records | Does the longitudinal record support the conclusion? |
| 012 | Clinical Guideline Applicability Audit | Patient-specific guideline population, eligibility, exclusion, version, and safety-gate validation | Does the evidence actually justify applying this rule? |
- Portfolio overview — PDF
- Portfolio overview — editable DOCX
- One-page portfolio summary — PDF
- One-page portfolio summary — editable DOCX
Benchmark_001_...zipthroughBenchmark_012_...zip— twelve complete downloadable benchmark packagesUSING_THE_BENCHMARKS.md— quick guide for reviewers and evaluatorsBENCHMARK_INDEX.md— portfolio-facing summary of all twelve benchmarksCOVERAGE_MATRIX.md— cross-benchmark reasoning coverage and distinctive contribution mapEVALUATION_FRAMEWORK.md— EIED pillars and reasoning methodologyDESIGN_METHODOLOGY.md— construction principles and balanced failure designPACKAGE_MANIFEST.md— verified suite counts and public distribution structureRELEASE_NOTES.md— Version 1.0 release summaryRIGHTS_AND_USE.md— copyright and reuse termsDISCLAIMER.md— scope, fictional-data, and validation limitationsAUTHOR.md— author background and related benchmark suites
- HORB asks: Can the AI reason through realistic healthcare operations?
- EIED asks: Has the AI earned the right to reach this conclusion?
- UCM asks: Has the AI exceeded what it can justify?
HORB emphasizes operational reasoning. EIED emphasizes evidence-supported decisions. UCM emphasizes uncertainty boundaries and confidence calibration.
Meredith Reese, RN
Registered Nurse | Healthcare AI Evaluation & Benchmark Design
GitHub: https://github.com/reesemeres-jpg
All scenarios and records are fictional. EIED is a portfolio and evaluation-design project, not clinical guidance or a validated regulatory instrument. No protected health information is included.