Institutions increasingly allow AI systems to recommend, plan, communicate, or act across consequential workflows. Claims that these systems are trusted, trustworthy, or subject to human control often leave the evidentiary test unspecified.
This repository asks:
What evidence justifies reliance on an autonomous AI system, and what evidence shows that institutional authority can still detect, interrupt, correct, and repair its actions?
The project develops an evidence architecture for bounded reliance. It identifies the object of reliance, the action being permitted, the governing conditions, and the records an independent reviewer would need to inspect.
Version 0.17.0 adds a bounded bridge from the six-stage reconstruction method to selected public policy and standards sources. It shows which practical-control questions the EU AI Act, NIST AI RMF, ISO/IEC 42001 public overview, and selected United States rules can motivate, while preserving the distinction between an organizational requirement and proof that control worked in a particular event. The release proposes an EU AI Act Article 57 sandbox as a future test setting. It does not claim legal compliance, ISO conformity, policy effectiveness, or validated transfer.
Version 0.16.0 rebuilds the methods paper around the institutional problem and formalizes the case-level decision rule. It preserves the public v0.14.0 Zenodo preprint, the v0.15.0 venue package, and the released case states. The new result is derived from those states: Oko is unresolved, both Patriot cases fail, and no selected case passes the complete event-control rule.
Version 0.16.1 aligned the repository citation, formula register, figure metadata, audit protocol links, and Overleaf compile receipt with the v0.16.0 research package. Version 0.16.2 is a maintenance candidate that creates one current-paper entry point, adds an external-review guide, moves the retired v0.15.0 delivery files into a labeled archive, and restores three historical v0.14.0 delivery artifacts to their released hashes. It changes no manuscript claim, case state, case-level result, or figure interpretation.
- A conceptual model separating trust, trustworthiness, reliance, justified reliance, and calibration.
- A six-variable autonomy profile covering goal scope, action authority, temporal horizon, impact radius, oversight distance, and reversibility.
- A seven-level evidence ladder from assertion through longitudinal accountability.
- A documentary test for practical human control across information access, comprehension capacity, intervention authority, intervention feasibility, exercised judgment, execution propagation, correction, repair, and reform.
- A solo-validation suite containing 12 synthetic cases, 252 prespecified determinations, 12 mutation tests, three invariance tests, and sealed oracle artifacts.
- A frozen public-case selection protocol with preserved candidate inputs, search output, exclusions, and selection decisions.
- Three public evidence packets covering a successful pre-action intervention, formal authority without practical force, and an action sequence whose cause remains indeterminate.
- A frozen research agenda focused on practical authority, evidence sufficiency, and interacting control conditions.
- A publication figure set containing six main figures, four appendix figures, ten derived data tables, formal captions, reading guides, and artifact-integrity checks.
- A machine-readable map connecting 40 material claims to exact evidence locations, human support states, evidence-fitness judgments, dependencies, limitations, and reversal conditions.
- An executable integrity audit that applies five checks and detects 39 prespecified corruptions without changing the released case packets.
- A research-lineage record, activity log, audit report, and claim-evidence matrix that preserve authorship, AI assistance, open exceptions, and conclusion eligibility.
- A prereassessment Oko adjudication protocol, frozen evidence universe, six-stage reassessment, and machine-readable change ledger.
- A 60-source working literature matrix and sentence-level audit covering the registered literature propositions in the full review draft.
- A frozen formal search containing eight direct queries, fifteen citation seeds, 2,431 deduplicated records, and a closed 89-record author-decision gate.
- A full methods manuscript with results, discussion, institutional implications, ethics, limitations, and AI-assistance disclosure.
- A structured table package that preserves exact states and counts in Markdown and journal-ready
booktabsfragments. - A paper-readiness package that keeps independent assessment, inaccessible-record review, authenticated database coverage, and ethics guidance outside the supported claim set.
- A 27-source full-text ledger with 22 verified full-text records, three abstract-only records, two inaccessible records, and no open decisions.
- A frozen recovery and residual-risk protocol for 1,087 inaccessible records, plus accountable logs for five authenticated or disciplinary interfaces.
- Research-agenda discovery logs and a v0.11 SHA-256 manifest sealing 171 artifacts while preserving earlier audit and protocol checkpoints.
- A five-record direct-query retrieval tranche with route-level evidence, five bounded screening decisions, a zero-permission source-content boundary, and an executable ledger cross-check.
- A v0.11 human-review attestation and claim-control audit that support five bounded claims, publish four exceptions, and block one proposed transfer claim.
- A 102-record forward-citation retrieval tranche, route-level evidence file, 71-record author queue, deterministic builder, and claim gate that blocks all pending records from the manuscript.
- A frozen 71-record forward-citation screening protocol, complete decision ledger, author-accountability attestation, proposition-review boundary, and four added negative controls.
- A frozen 13-source proposition-review protocol with five bounded manuscript permissions, two background-only decisions, six quarantines, corrected source identities, and seven added negative controls.
- A professional single-column LaTeX package with navy-and-black journal styling, ten color figures, structured tables, a deterministic 12-member source archive, metadata, source lineage, and explicit compilation and author-review gates.
- A formal case-level rule, a deterministic result builder, a proposed timing margin, construct-derivation and institutional-interpretation tables, and an evidence-controlled manuscript rebuild.
- A five-claim, eight-source policy crosswalk with author attestation, nine detected mutations, a prospective regulatory-sandbox design, and an arXiv-ready v0.17.0 preprint package.
All 252 determinations and 12 original mutation tests pass under the committed contract. The v0.6.0 adjudication detects all six prespecified corruptions. The v0.16 integrity audit maps 40 material claims and detects all 39 prespecified claim-map corruptions. Five exceptions remain: no independent assessment, incomplete literature-search coverage, two direct-query source limits, untested contemporary transfer, and external venue status. The inaccessible-record recovery population contains 1,087 records; 107 outcomes are recorded and 980 remain open. Five forward-citation sources may support only their recorded propositions, two remain background-only, and six are quarantined. These results establish internal contract behavior and traceability for the included artifacts. They do not establish independent reliability, field validity, institutional effectiveness, source truth, universal originality, or improved outcomes.
The v0.10.0 release preserves every earlier release artifact and freezes the next evidence checkpoint before new results are known. Eight of the 27 retained-close sources have a recorded full-text review basis, leaving 19 open. The protocol controls those reviews, recovery of the 1,087 inaccessible records, a reproducible residual-risk sample, and five authenticated or disciplinary-interface searches. The earlier case packets, the v0.3.0 Oko assessment, and the released v0.9 claim audit remain unchanged. Independent assessment remains a separate validity question.
Post-release work has assigned a terminal state to all 27 retained-close sources: 22 verified full text, 3 abstract-only records, 2 inaccessible records, and no open decisions. This working result closes the first v0.10 evidence gate. It does not change the published v0.10.0 snapshot or resolve the 1,087-record recovery gate.
Version 0.11.0 freezes a 284-record residual-risk sample before retrieval outcomes are known. The sample contains 102 forward citations, 177 backward references, and 5 direct-query records selected by the declared SHA-256 ordering rule. The direct-query stratum has five retrieval outcomes, four screening decisions, and one open author review. The other 279 sampled records remain open. Its claim audit supports the bounded workflow count and four source descriptions. It keeps the proposed cross-domain mechanism outside the eligible conclusion set.
Version 0.12.0 records retrieval outcomes for all 102 forward-citation records. It recovered full text for 34 records and abstracts for 37, recorded 26 metadata-only outcomes, reconciled 3 duplicates, and left 2 unavailable. The release preserved all 71 recovered-content records outside manuscript claims pending screening.
The v0.13.0 candidate closes that frozen screening queue with 13 close records, 22 background records, 11 single-component exclusions, and 25 topic exclusions. All 71 decisions record a mechanism-specific rationale, source basis, locator, decision owner, date, assistance disclosure, and claim-permission state. Screening grants no proposition support. The 13 close sources now enter a separate locator-level review gate.
The v0.14.0 public preprint closes that locator-level gate. Five sources receive one bounded manuscript permission each, two remain background-only, and six remain quarantined. RS-DQ-004 is close for screening and has zero source-content permission. The version also adds the first repository-controlled source archive and a canonical 25-page Overleaf compilation with zero errors. It is archived on Zenodo under version DOI 10.5281/zenodo.21926005.
The v0.15.0 candidate adapts that same bounded paper for Preprints.org. The title page identifies the author as an independent researcher with Node & Norm, retains both authorized correspondence addresses, and records the Harvard University student relationship separately in the author note. The single-column presentation uses black body text, dark navy headings and rules, and light gray-blue table headers. All ten figures and seven tables appear before References. Preprints.org submission and screening remain external states and are not implied by the repository release.
The v0.16.0 working paper responds to the venue outcome by clarifying the research contribution without overstating novelty. It explains why each methodological control exists, adds the complete six-stage rule, derives the three case-level results from released data, and states what institutions may and may not infer from pass, fail, and unresolved outcomes. Its exact source compiles to a 30-page review PDF with zero errors, no overfull or underfull boxes, and no displays after References. Full author review remains open. The Preprints.org decline remains an external screening event and supplies no evidence that the internal controls failed.
Figure 2 compares whether assigned human authority became practical control in three historical cases.
The featured figure asks a simple question: Could the designated human actually change what the system did?
Each packet contains evidence that a human held a formal role. The assessed strength and practical consequences differ because authority is only one link in a longer chain:
- Did the person receive the relevant information?
- Could they understand it?
- Did they have authority to intervene?
- Was intervention realistically possible in the available time?
- Did they intervene?
- Did the intervention propagate into execution?
In the current v0.6.0 assessment, every Oko stage from access through execution propagation is partially supported. Retrospective participant accounts describe Stanislav Petrov receiving the warning, questioning it, reporting a false alarm, and affecting the decision path. No located contemporaneous command log or official incident record independently records those stages. The partial cells preserve both the account and that missing evidence. Under the v0.16.0 case-level rule, Oko is unresolved.
In the Patriot ZG710 case, a human authorized the engagement. The evidence indicates weak comprehension and no feasible or exercised challenge before launch. Execution propagation is also unsupported. The case fails the event-control rule.
In the F/A-18C case, the public record confirms human authority and some access to information. Missing records prevent conclusions about what operators saw, understood, or could have done in time. Open diamonds marked I mean “the packet cannot decide.” Gray crosses marked U record evidence that a condition failed. Unsupported execution propagation makes the case fail the event-control rule.
The later reforms shown in both Patriot cases indicate institutional learning. They could not repair the losses already caused.
The central lesson is that assigning a human role does not establish practical control. Institutions need separate evidence for timely information, comprehension, authority, opportunity, intervention, and execution propagation.
The current figure derives this pattern from the 27 plotted states. Oko records partial support across the six event-level stages. Both Patriot packets support formal authority, while the other practical conditions are unsupported or unresolved. The deterministic case result and figure methods preserve the formal derivation.
The paper-stage PR #11 pressure test identified the Oko protocol mismatch. The v0.6 adjudication resolves it through six reclassifications made under a protocol frozen before reassessment. The decision corrects the current assessment and does not add missing historical evidence.
The figure shows this pattern across three historical cases. It supplies no estimate of how often these failures occur and no prediction of performance in current AI systems.
The publication figure set contains six main figures, four appendix figures, derived data, formal captions, and plain-language reading guides. The structured tables preserve the exact states and counts behind the graphics.
Figure A3 asks a second question: Is a traceable claim fit to support a conclusion?
Every mapped claim passes traceability, which means its declared evidence locations resolve. Traceability is the first column. The later columns test separate questions: whether the artifact's integrity can be checked, whether a human reviewed support, whether the evidence fits the claim, and whether every dependency closes.
The Oko claim, PAPER-C04, passes because it reports partial support and preserves the missing contemporaneous-record limit. The dependent paper conclusion, PAPER-C09, also passes within the declared single-assessor procedure. TAE-C23 remains ineligible because no independent study has tested reliability or field validity. PAPER-C26 is now eligible within the 89-record queue because every author decision is recorded and Figure 5 resolves from the ledger. The central lesson is simple: claim eligibility depends on matching the conclusion to evidence that is fit for its exact scope.
The matrix uses categorical states and letter labels so color is not the only signal. It calculates no aggregate trust score. The derived data, figure specification, and v0.16 audit report preserve the exact path behind every cell.
| Path | Purpose |
|---|---|
README.md |
States the research question, current result, reading order, and validation commands. |
RESEARCH_STATUS.md |
Records the release state, completed artifacts, active work, and open empirical questions. |
CLAIMS.md |
Lists each proposition with its evidence, confidence, limits, and reversal conditions. |
LIMITATIONS.md |
Names validity threats and conclusions outside the current evidence. |
SOURCES.md |
Records the standards, papers, and public repositories used by the project. |
CITATION.cff |
Provides machine-readable authorship, release, license, and DOI metadata. |
CHANGELOG.md |
Tracks material changes to concepts, protocols, claims, and evidence requirements. |
release/v0.17.0-release-notes.md |
Explains the bounded policy crosswalk, prospective validation design, preprint package, integrity results, and claim limits. |
release/v0.17.0-manifest.json |
Seals the v0.17.0 manuscript, source package, compiled PDF, policy evidence, attestation, mutations, and audit artifacts. |
release/v0.16.1-release-notes.md |
Explains the maintenance alignment, preserved v0.16.0 findings, and version boundary. |
release/v0.16.1-manifest.json |
Seals the citation, formula, figure-metadata, audit-link, compile-receipt, and validation corrections with SHA-256 digests. |
release/v0.16.2-release-notes.md |
Explains the paper-workspace organization, archive boundary, and preserved v0.16.0 research result. |
release/v0.16.2-manifest.json |
Seals the navigation, archived v0.15.0 files, current paper paths, and updated validation controls. |
release/v0.14.0-release-notes.md |
Explains the v0.14.0 proposition review, preprint package, integrity controls, Zenodo DOI, and open external-validation limits. |
release/v0.15.0-release-notes.md |
Explains the v0.15.0 venue package, author metadata, version relationship, carried-forward evidence controls, and open submission gate. |
release/v0.16.0-release-notes.md |
Explains the manuscript rebuild, formal event-control rule, derived results, integrity controls, and open empirical gates. |
release/v0.16.0-manifest.json |
Seals the v0.16.0 manuscript, source package, compiled PDF, results, figures, claims, and audit artifacts with SHA-256 digests. |
CONTRIBUTING.md and GOVERNANCE.md |
Define contribution evidence, review rules, decision authority, and change records. |
research/ |
Contains the main conceptual paper on justified reliance, autonomy, and practical control. |
research/frozen-research-agenda.md |
Freezes three research topics and one project question for the next public-case cycle. |
research/agenda-discovery-log-v0.10.0.md |
Records findings that changed the work sequence while preserving the frozen topics and question. |
research/agenda-discovery-log-v0.11.0.md |
Records retrieval findings about inspectability, preserved provenance, parallel intervention paths, and the Patriot-adjacent close source. |
research/agenda-discovery-log-v0.12.0.md |
Records why retrieval, screening, access, duplicate handling, and independent assessment remain separate gates. |
paper/ |
Develops the methods paper, formal search, claim register, literature audit, structured tables, references, and publication decisions. |
paper/README.md |
Provides the single entry point to the latest paper, current source, GitHub edition, and earlier packages. |
paper/REVIEW.md |
Gives reviewers and prospective arXiv endorsers one current PDF, a category-fit summary, evidence paths, and focused review questions. |
paper/archive/ |
Indexes earlier paper packages and stores the retired v0.15.0 delivery files. |
paper/preprints/preprints-compiled-v0.16.0.pdf |
Provides the 30-page technical review PDF compiled from the exact v0.16.0 source archive. |
paper/preprints/compile-receipt-v0.16.0.json |
Records compiler identity, source and PDF hashes, page locations, visual inspection, and the remaining author-review gate. |
paper/preprints/overleaf-compile-receipt.json |
Records the 30-page v0.16.0 XeLaTeX compilation and full-page visual review in Overleaf. |
assessments/event-control-results-v0.16.0.json |
Stores the deterministic result: zero pass, two fail, and one unresolved case under the formal rule. |
paper/manuscript-reader.md |
Renders the manuscript's citation identifiers as clickable author-year citations with a reference list. |
paper/manuscript-pressure-test-v0.8.0.md |
Records citation, count, claim, reliability, ethics, and submission-gate findings. |
paper/review-record-v0.8.0.md |
Records author authorization, reviewed additions, support decisions, and publication limits. |
paper/review-record-v0.9.0.md |
Records the 89 author decisions, contribution decision, source boundary, and remaining search limits. |
paper/author-screening-completion-gate.md |
Records progress on the 89 author decisions and controls when final search-flow language becomes eligible. |
paper/next-evidence-gates-v0.10.0.md |
Reports the full-text, inaccessible-record, authenticated-interface, and independence gate states. |
paper/data/close-source-full-text-gate-v0.10.0.csv |
Records one full-text state for each of the 27 retained-close sources. |
paper/inaccessible-risk-sample-v0.11.0.md |
Explains the frozen 284-record sample, proportional allocation, reproduction rule, and claim boundary. |
paper/data/inaccessible-risk-sample-v0.11.0.csv |
Records each selected key, primary stratum, allocation, digest, rank, origin set, and source metadata. |
paper/data/inaccessible-risk-sample-v0.11.0.json |
Records the population hash, seed, allocation method, stratum counts, sample hash, and evidence boundary. |
paper/direct-query-retrieval-tranche-v0.11.0.md |
Explains the five direct-query recoveries, four screening decisions, source limits, paper effects, and next action. |
paper/data/direct-query-retrieval-evidence-v0.11.0.json |
Records the route, locator, review basis, source observations, decision, assistance, and limit for each direct-query record. |
paper/forward-citation-retrieval-tranche-v0.12.0.md |
Reports 102 retrieval outcomes, 71 open author decisions, access limits, duplicate handling, and the next gate. |
paper/data/forward-citation-retrieval-evidence-v0.12.0.json |
Records the route, outcome, source observation, review basis, assistance, and claim limit for every selected forward citation. |
paper/data/forward-citation-author-review-queue-v0.12.0.csv |
Holds 71 recovered-content records with blank author decisions and no claim permission. |
paper/data/author-screening-gate-v0.8.0.json |
Preserves the open author-gate checkpoint published in v0.8.0. |
paper/data/author-screening-gate-v0.9.0.json |
Stores the closed gate, final decision counts, and Figure 5 eligibility state. |
paper/tables.md |
Publishes compact exact-value tables with captions, notes, and interpretation boundaries. |
paper/tables/manuscript-tables.tex |
Provides journal-style booktabs fragments with three horizontal rules and no vertical rules. |
formulas/ |
Maps eight v0.16.0 formulas to their decision purpose, publication status, source locations, implementations, and limits. |
evidence/ |
Contains the trust evidence register, current and preserved claim maps, human-review attestation, research lineage, and AI-assisted activity log. |
evidence/claim-evidence-map.json |
Connects 40 material claims to exact locators, five fitness dimensions, dependencies, human review, limits, and reversal conditions. |
evidence/human-review-attestation-v0.11.0.json |
Records author review of five direct-query states and six added claims, with the limits of AI assistance. |
evidence/research-lineage.json |
Records people, software, research activities, artifacts, and relations using PROV-O-compatible concepts. |
protocols/ |
Defines solo validation, independent review, public-case reconstruction, practical control, and claim-evidence integrity procedures. |
protocols/coe-integrity-audit.md |
Defines the five claim gates, four adapted CoE checks, repository-specific closure check, negative controls, and conclusion rule. |
protocols/search-coverage-and-full-text-protocol-v0.10.0.md |
Freezes full-text verification, inaccessible-record recovery, residual-risk sampling, and authenticated-interface completion rules. |
protocols/public-case-reconstruction-protocol.md |
Freezes the source cutoff, candidate pools, eligibility rules, screening order, and reconstruction procedure before case selection. |
cases/ |
Publishes three case packets, their provenance manifests, assessments, hashes, and admissibility requirements. |
cases/public-case-selection-register.md |
Preserves the frozen collection hashes and every inclusion or exclusion in screening order. |
cases/data/candidate-search-output.json |
Preserves the deterministic search result from the two candidate collections without redistributing article text. |
schemas/ |
Defines machine-readable contracts for synthetic cases, public cases, adjudication, literature support, claim maps, lineage, mutations, and audit results. |
fixtures/ |
Contains 12 synthetic cases, the original mutation suite, six v0.6 adjudication controls, and 39 current claim-integrity controls. |
oracles/ |
Stores prespecified expected decisions and the SHA-256 manifest that seals them. |
analysis/ |
Implements deterministic assessment logic and builders for the publication figures and claim-evidence matrix. |
assessments/ |
Stores generated results plus the current v0.6 Oko assessment and change ledger. |
reports/ |
Publishes the solo-validation and three-case reconstruction results with explicit claim boundaries. |
figures/ |
Publishes six main figures, four appendix figures, ten derived CSV files, plotting specifications, and plain-language reading guides. |
reports/figure-methods.md |
Records formal captions, transformations, missingness treatment, and prohibited interpretations for the figure set. |
audits/v0.16.0/ |
Publishes the current audit plan, machine-readable result, plain-language report, and five open exceptions. |
audits/v0.9.0/ |
Preserves the prior 20-claim audit as version history. |
audits/v0.8.0/ |
Preserves the open author-screening checkpoint as version history. |
audits/v0.6.0/ |
Preserves the earlier 15-claim audit as version history. |
scripts/ |
Contains candidate-search, packet-sealing, release-manifest, repository-validation, paper-validation, and integrity-audit utilities. |
scripts/build_forward_citation_tranche_v0_12_0.py |
Rebuilds the 102-record evidence file, population-ledger rows, and 71-record author queue from the frozen sample. |
release/ |
Seals each versioned research package with SHA-256 digests while preserving earlier releases. |
release/v0.10.0-release-notes.md |
Explains why the protocol checkpoint is released before the new evidence gates close. |
release/v0.11.0-release-notes.md |
Explains the direct-query evidence, claim-control result, exceptions, and next gate. |
release/v0.12.0-release-notes.md |
Explains the forward-citation retrieval state, pending author gate, controls, exceptions, and next work. |
mappings/ |
Relates this work to GDI, HIT, CDFI, and CDCF governance artifacts. |
.github/ |
Defines automated validation, the pull-request checklist, and structured issue forms. |
requirements-dev.txt and LICENSE |
Pin the validation dependency and state the Apache-2.0 license. |
Start with Trust, Autonomy, and Evidence. Use the Trust Evidence Register to translate a reliance claim into inspectable evidence. Read the current public-case report, CLAIMS.md, and LIMITATIONS.md before applying the model.
The five protocols define solo validation, public-case reconstruction, practical-control assessment, claim-evidence integrity, and future independent evaluation. The public-case selection register must record its freeze commit before candidate screening begins. The mapping files show how this project relates to existing public artifacts without transferring claims among them.
Install the pinned development dependency and run the validator with Python 3.10 or later:
python -m pip install -r requirements-dev.txt
python scripts/validate_repository.pyThe repository validator checks required release files, internal links, version alignment, schemas, source references, sealed packet and release hashes, selection invariants, current case interactions, 252 oracle comparisons, 12 original mutation tests, six adjudication controls, 22 claim-integrity controls, formal-search consistency, literature support, and figure integrity. The paper validator checks author identity, question alignment, at least 45 bibliography entries, the archived v0.6 DOI, current claim eligibility, and the originality-language boundary. Successful runs end with repository validation: PASS, chain-of-evidence audit: PASS_WITH_EXCEPTIONS, and paper validation: PASS.
This repository does not establish that:
- an AI system is generally safe or trustworthy;
- a complete decision record is truthful;
- a reviewer understood the evidence;
- formal human authority had practical force;
- a governance artifact satisfies a legal or normative requirement;
- the proposed evidence architecture improves outcomes.
Each proposition requires evidence from the deployment, institution, decision, and review context in which the claim is made.
Contributions should identify the proposition being changed, the evidence supporting the change, the limits of that evidence, and the conditions that would reverse the conclusion. See CONTRIBUTING.md and GOVERNANCE.md. Case material must exclude personal, confidential, and institutionally restricted information unless the contributor has documented authority to publish it.
The current working paper has no v0.16.0 DOI. Cite the versioned GitHub release so the cited manuscript and research package remain identifiable:
Banasihan, M. J. (2026). From Formal Authority to Practical Human Control: A traceable method for reconstructing human control in automated decisions (Version v0.16.0) [Working paper]. GitHub. https://github.com/mj3b/trust-autonomy-evidence/releases/tag/v0.16.0
The DOI 10.5281/zenodo.21926005 identifies the earlier v0.14.0 preprint. The concept DOI 10.5281/zenodo.21841127 identifies archived repository versions. Neither DOI currently identifies the v0.16.0 manuscript. Machine-readable metadata and the preferred paper citation are in CITATION.cff.
Mark Julius Banasihan is an independent applied researcher studying decision authority, human influence, evaluation, and assurance in AI-mediated institutional systems.

