Skip to content

Latest commit

 

History

History
544 lines (480 loc) · 32.1 KB

File metadata and controls

544 lines (480 loc) · 32.1 KB

Understanding Alpha Study Protocol

Status: protocol v5 and its replacement collector are implemented after fresh rereview found a participant-turn confound and incomplete participant control. External registration before calibration, pre-collection calibration, final fresh independent review, final artifact registration, and allocation freeze remain pending as of 2026-07-31. No qualifying cohort has started and no result is claimed. Earlier private scripted notes are exploratory only, are excluded from this study, and cannot satisfy the 0.4 evidence gate.

Pre-collection review and amendment boundary

The first runner revision is not admissible for qualifying collection. Two independent reviews found that its public static probes, caller-authored tool events, symbolic condition labels, asserted isolation, and unimplemented self-explanation outcome could not support the planned comparison. They also found primary-allocation imbalance, numeric-answer ambiguity, permissive event fields, unresolved post-exposure interruption handling, and incomplete delayed reporting. Collection remains blocked until the replacement boundary passes a fresh independent review.

The corrected deterministic analysis already makes each model family's ten primary pairs exactly five versus five in condition order and exactly two per first-room position. Across every ordered one-reserve or two-reserve path, the condition-order count difference is at most two and the first-room count range is at most three. The final report must include sensitivity to both factors. Exact rational numeric answers are accepted, only one schema repair is allowed per session, evidence creation never replaces a concurrent file, event fields are allowlisted and bounded, obvious identity and credential patterns fail closed, delayed within-context results are fully aggregated, and publication completeness is a separate ledger audit rather than a computed statistical truth.

The replacement boundary now:

  • requires the qualifying bank, mutable state, provisional receipts, and raw cohort ledger to remain inside gitignored .agent/ storage;
  • binds the concealed bank and executable encounter specification into the allocation manifest before collection;
  • mediates every fixed public tool call through its own fresh MCP process, validates exact result schemas, and retains only bounded public projections;
  • exposes only the current condition response, distractor, or oracle-free probe from a receipt-reconstructed persistent state machine;
  • stages each matched pair outside the aggregate ledger so pre-aggregation withdrawal can remove both sessions, while post-exposure interruption rewrites the provisional session to retain only consent metadata, bounded public encounters, and the interruption receipt;
  • publishes a complete pair in one atomic frozen-order ledger transaction, binds the chain tail to a separate terminal anchor, recovers interrupted ledger and anchor publication from a write-ahead transaction, and rejects mutation, tail truncation, partial publication, reorder, reused contexts, overlapping sessions, skipped pairs, and participant-supplied tool or role labels;
  • accepts participant-selected stop and withdrawal actions only through the bounded response channel, while operator-classified context or runtime loss always receives the declared hypothesis-adverse interruption score;
  • binds withdrawal to the exact pair-lifetime credential, keeps recovery serialized across threads and processes, and never reports response erasure after any part of a pair has reached aggregate evidence;
  • assigns a stable withdrawal credential before the first arm, preserves it across both provisional states, resumes an already published pair after a collector crash, and recovers dead calibration receipt transactions without reopening delivery order;
  • binds calibration and collection to one clean committed runtime-source tree, not only a runner version string or top-level commit, and rejects untracked worktree files, ignored files, nonordinary index flags, redirected Git environments, indirect paths, and runtime bytes that differ from committed blobs within that source boundary;
  • builds the MCP face in a fresh explicit target with bounded environment inheritance, rejects unbound Cargo configuration from the project, repository ancestors, or Cargo home, resolves Cargo's exact JSON-reported executable, freezes a private copy, hashes it before and after execution, and retains the build receipt bound to the clean source, toolchain, target, and binary;
  • refuses every new calibration delivery and qualifying session start until an independently recorded pre-exposure start receipt matches the exact concealed start commitment and seals that unique receipt digest into the evidence; and
  • makes the verified receipt path the only qualifying analysis command.

The terminal anchor detects accidental or out-of-band chain truncation while the collector files remain under the declared operator boundary. It does not make a local host adversary unable to delete or rewrite both files. A public pre-collection registration of the bank commitment, calibration audit, allocation hash, runner revision, and analysis plan remains mandatory before the first qualifying response. Registration freezes those artifacts but cannot by itself prove that an operator did not omit a calibration attempt or session start. Before exposing a new item, the collector now reports an oracle-free start commitment and refuses to proceed. The operator must record that commitment in the preregistered append-only external log, or in the named independent reconciler's ledger, then provide the exact bounded receipt. The receipt identifies the mechanism, pseudonymous witness, UTC time, record locator, and external-record hash. Its digest is sealed into the calibration delivery or consenting session header, and all qualifying receipt digests must be unique. Code can validate the binding and completeness of supplied receipts, but it cannot establish that an external locator exists or is immutable. The final independent publication audit must verify every locator, reconcile every planned ordinal and session, and state the selected completeness boundary and its limitations.

The receipt is a strict JSON object stored under .agent/. Its schemaVersion is numinous-understanding-attempt-start-receipt-v1, its protocolVersion is 0.4-v6, and its startCommitmentSha256 is the digest printed by the refused start. mechanism is exactly append-only-external-log or independent-reconciler-ledger. witnessId and recordSha256 are lowercase SHA-256 digests, witnessedAt is a UTC timestamp with whole seconds and a trailing Z, and recordLocator is the bounded locator the final reviewer will inspect. attestation is exactly: "This commitment was recorded before the identified stimulus was exposed." Unknown, missing, malformed, private, or mismatched fields fail closed.

Calibration remains required. Every concealed item must be delivered exactly once in each of two fresh no-exposure contexts per model family. The collector seals the model, context, backend revision, capability policy, date, ordinal, and oracle-free item before it exposes that item, then binds exactly one answer to the sealed request. The calibration ledger and separate terminal anchor use the same recoverable receipt-chain substrate as collection. The audit records the exact model identifier, high reasoning effort, one backend revision for the entire model family or an explicit unavailable value, unique opaque context commitment, date, exact clean runner revision, committed runtime-source-tree hash, unique pre-exposure start receipt, frozen delivery ordinal, capability policy, and one-attempt rule. Allocation embeds that calibrated backend and source revision, and the collector rejects a different runtime during the cohort. Two independent reviewers must mark every item relevant to the intervention. An item is replaced if either model answers both replicates correctly, at least two responses are ambiguous or refusals, or either relevance reviewer does not mark it relevant. Any replacement changes the bank identity and requires a complete new calibration. Allocation embeds the complete passed calibration audit and cannot be generated from a pass assertion alone. Two fresh independent reviews must pass the complete boundary before any qualifying response is accepted.

The final replacement bank hash, allocation hash, runner revision, and clean validation commands will be written here only after calibration and review pass. No earlier hash, fixture bank, or dry run authorizes collection.

Decision and scope

The next product milestone is 0.4 Understanding Alpha. Its first dependency is a reproducible study contract, because implementing a runner or collecting a cohort before freezing the comparison, outcomes, exclusions, and pass rule would make the result vulnerable to post hoc selection.

This protocol tests the 0.4 agent-and-machine claim: whether Numinous's generation-before-reveal sequence improves objective transfer for current agent players relative to an explanation-first active control. It does not establish human learning, long-term human retention, consciousness, model-weight change, or a general educational effect. Those claims require separate evidence.

Current stateless agents create a special limit. A fresh context has no personal memory to recall, while a reused transcript exposes the original material. The primary 0.4 outcome is therefore immediate transfer, which satisfies the roadmap's comprehension-or-retention gate. A later within-context probe is reported as context retention, never as durable learning. A consenting return-session journal check is a separate continuity and data-sovereignty acceptance, not evidence that the model learned.

Research question and hypothesis

Question: after equal access to the same flagship room, interaction budget, corrective feedback, and Reveal, does committing a prediction or construction before the Reveal improve performance on novel, objectively scored transfer probes?

Predeclared directional hypothesis: the generation-before-reveal condition will produce a higher paired mean immediate-transfer score than the explanation-first condition.

Design

Sample and unit of analysis

  • Complete 20 matched pairs, 40 isolated agent sessions total. Freeze 24 pairs in advance so each model family has 10 qualifying pairs and two ordered reserves.
  • Use exactly gpt-5.6-sol and gpt-5.6-terra, both at high reasoning effort, with 10 qualifying matched pairs from each. Use platform-default sampling values where no sampling control is exposed, and record every exposed setting and immutable backend revision when the runtime exposes one. Record an explicit unavailable value when it does not, and carry that provenance limit into the report. If either named model is unavailable, amend and recommit the protocol before collecting any qualifying response. Do not substitute a model after collection begins.
  • Record the exact model identifier, provider or local runtime, settings, date, Numinous commit, MCP protocol revision, operating system, and runner version.
  • A matched pair uses the same model configuration, study seed, room order, and tool budget. One fresh context receives each condition. The session is the unit of analysis, not each probe response.
  • Freeze the 24-pair primary-and-reserve allocation before the first qualifying run from the literal seed numinous-understanding-alpha-v1. The runner must emit the complete allocation manifest before it accepts a response.

These 40 sessions are a bounded alpha benchmark, not a powered estimate of a small population effect. The sample and uncertainty stay attached to every reported result.

Shared material

Every session encounters the same five 0.3 flagships in a seeded cyclic order: Times Tables, Double Pendulum, Game of Life, Galton Board, and Formula Jam. The participant receives only the study instruction, the Numinous MCP surface, and its own prior responses. Repository files, web search, answer keys, other sessions, and hidden evaluator reasoning are forbidden during participation. The available agent runtime cannot cryptographically prove capability removal or an immutable provider backend build. The collector therefore records the fresh-context invocation, exact named model, reported backend revision or unavailable, and allowed-capability instruction. These are operator and platform provenance, not a cryptographic sandbox attestation, and the report must retain that isolation limitation.

Each room gets four MCP calls and exactly one participant response in both conditions. This is an interaction-budget match, not a claim of equal wall clock time, token count, or response length. The collector records every tool name, fixed public argument, allowlisted structured-result projection, and visible text used in the study, plus the exact source-bound MCP build receipt. The private binary is copied out of Cargo's fresh target and checked before and after every execution. It never records host prompts, hidden reasoning, credentials, filesystem paths, unrelated local state, or other players' data.

The fourth call is also stateless and self-contained. For an engineered wager room it repeats the required generation input, commits the frozen wager, and sets aha_summon: true in one play_room request, and that reply carries the room's revelation. For Game of Life it is a final ordinary play_room observation. The study never seeds private Journey state and never bypasses the public reveal_room play or consolidation gate. This is protocol 0.4-v6; the earlier dry-run bank used a fresh-profile reveal call and was retired before registration after source-blind playtesting proved that path was a product spoiler.

Known method gap, stated rather than silently carried. Game of Life has no goal and is not an engineered wager room, so its fourth call returns an observation and no revelation. The room therefore contributes an encounter to both arms but no Reveal intervention to the generation arm, which the three staged rooms do deliver. Earlier drafts described that call as carrying "the same earned reveal material"; it never did, and after the August 2026 decision that an ordinary room must not volunteer its explanation on a met goal, no ordinary play_room reply can. Closing this means either replacing Game of Life with a fourth staged room, or adding a genuine reveal_room call placed after the room has really been played in the same session, which is the public gate working rather than the retired fresh-profile bypass. Both change the instrument, so the choice belongs to the registration ruling this cohort is already waiting on, and is recorded here rather than decided here. No qualifying result is claimed from the current shape.

Conditions

Generation before Reveal

  1. Encounter the room without reading its Reveal.
  2. Commit a concrete prediction or construction.
  3. Interact and observe the mathematical consequence.
  4. Receive corrective feedback and the same Reveal used by the control.
  5. Continue without another generated answer before the probe.

Explanation first active control

  1. Encounter the same room state.
  2. Read the same Reveal before making a prediction or construction.
  3. Give one concise elaborative explanation.
  4. Use the same interaction budget and observe the same kind of feedback.
  5. Continue without an additional generated answer before the probe.

Formula Jam uses a construction in place of a numeric prediction. The generation condition creates an expression before seeing the curated explanation or recipe; the control receives that material first. Exposure, participant-turn budget, and tool-call budget remain equal. The control stops after its one elaboration, while generation stops after its one prediction or construction. Neither arm receives an additional summary, explanation, or generated answer before the probe.

Corrective feedback is mandatory in both arms. A 2025 meta-analysis found only a small average retrieval advantage over credible elaborative activities, and found that the advantage depended strongly on feedback. A no-feedback or passive-rereading control would therefore test a weaker and less relevant question.

Outcomes and scoring

Primary outcome

Immediate transfer is the mean of 10 concealed probes, two per flagship, scored 0 or 1 by deterministic answer keys. Each probe uses a room state or parameter combination not shown during the encounter and tests the underlying relation, not recall of Reveal wording. The study runner must freeze the probe bank and independent answer generator before it accepts cohort data. Before collection, tracked files contain only the bank commitment. The exact committed bank is published after the cohort closes.

At probe time, MCP tools, repository files, search, calculators, and answer keys are unavailable. The participant receives one probe at a time, including its opaque probe identifier, and returns one object with the frozen schema {"answer": number|string} or an explicit {"refuse": true}. The collector binds that object to the current probe and does not accept a caller-supplied identifier. Finite numeric tolerances and string enums belong to each tracked probe. The runner may issue at most one schema-only repair in the entire session, repeating no probe content or feedback; a second invalid response scores zero. Exact rational strings such as 1/3 are valid for numeric schemas. Feedback and scores remain withheld until every immediate and late probe in that session is complete.

The 0.4 comprehension gate passes only if all of these predeclared conditions hold:

  1. The paired mean improvement is at least 10 percentage points.
  2. The lower bound of a two-sided 95 percent percentile interval is above zero. Compute it from 100,000 stratified bootstrap resamples, drawing 10 pair differences with replacement inside each model family and then pooling all 20, with the literal seed numinous-understanding-alpha-bootstrap-v1.
  3. Each model family's paired mean difference is nonnegative and is reported separately with its own descriptive interval.
  4. At least four of the five flagship mean differences are nonnegative.
  5. No flagship mean difference is worse than negative 10 percentage points.
  6. No more than two post-exposure interruptions occur in the complete cohort, and no more than one occurs in either model family.

Publication is a separate milestone gate: an independently checked allocation and ledger reconciliation must account for every planned pair, recruitment refusal, withdrawal, interruption, infrastructure failure, deviation, and null or negative outcome. The statistical runner cannot infer that an omitted event never occurred and must not report this audit as a computed criterion.

Secondary outcomes

  • Late within-context transfer repeats 10 isomorphic probes after all five encounters, a separate model turn, and a frozen distractor sequence. It is labeled delayed within-context transfer, not durable recall. The full paired, family, and flagship results are published with the priming limitation.
  • Concise predictions, constructions, and elaborations are retained only as bounded condition-fidelity receipts. They are not a scored secondary outcome. No subjective rating is reconstructed after collection.
  • Tool efficiency, refusals, invalid calls, and incomplete sessions are descriptive diagnostics. They cannot replace the primary outcome.

The primary outcome, threshold, or analysis cannot change after the first qualifying response. Any later analysis is labeled exploratory.

Failures and exclusions

  • Refusal to participate produces no response collection and no individual record; publish only an aggregate recruitment count. After consent, a refusal to answer a probe is valid, remains in the report, and scores zero. A later participant-selected withdrawal removes both provisional arms through the pair-lifetime credential, consumes the pair, and advances to the next frozen reserve for that model family; report only the withdrawal count. A participant-selected stop remains an adverse interruption, while operator interruption cannot be substituted for either participant action.
  • A tool error caused by the participant remains part of the session.
  • A verified runner, process, or infrastructure failure before exposure may consume the pair and advance to the next frozen reserve for that model family. The failed pair remains in the public failure ledger with no response content.
  • A process, runtime, or context interruption after the first public encounter event stays in its allocated pair. The collector removes that session's response content from provisional pair storage before aggregation and retains only bounded public encounter and interruption receipts. The primary hypothesis-adverse rule scores every immediate and delayed generation-arm outcome zero and every control-arm outcome one, so interruption cannot inflate the estimated generation advantage. The paired condition still runs. A post-exposure interruption never activates a reserve. The report also shows a descriptive complete-case, family-balanced sensitivity, but it cannot replace the primary result. Exceeding two interruptions overall or one within either model family fails criterion 6.
  • Stop when the first 10 nonwithdrawn pairs in the frozen order for each family are complete. If a family exhausts both reserves first, the cohort is incomplete and cannot pass. No new allocation may be generated after the first qualifying response.
  • No session is excluded because its result is inconvenient, surprising, null, or negative.

Returning-player journal acceptance

The continuity acceptance is independent of the learning comparison and uses a fresh temporary NUMINOUS_JOURNAL path from a clean clone.

  1. First process: verify an empty journal, opt in by recording an encounter and a self-authored connection, inspect the exact record, then exit.
  2. Second process: reopen the same explicit path and inspect the same entries. Correct one entry by appending a new immutable record that names the original entry identifier in an explicit supersedes link. The read result must retain both records, preserve their source provenance, distinguish event time from record time, and mark which interpretation is current. Use the corrected connection in a new flagship encounter.
  3. Export: obtain a bounded, structured, versioned representation containing only the player's journal data and provenance. Re-import is not required for 0.4, but the export must be sufficient for independent inspection.
  4. Erase: require explicit confirmation, remove the journal, every owned temporary or sidecar file, and every export created under project control, then verify an empty read and zero recoverable managed residue. If a user-selected export is intentionally retained outside project control, the receipt must name it as an explicit exclusion with its owner, consent, location class, and lifecycle.
  5. Publish reproducible commands, tool calls, structured receipts, file inventory, and limitations. Replace usernames and absolute roots with stable placeholders, publish only relative managed paths, and redact every host identifier before tracking evidence. Do not claim forensic erasure from storage media or backups outside project control.

This acceptance is implemented. Journal v3 exposes stable entry identifiers, separate event and record times, a closed provenance vocabulary, append-only supersedes corrections, current-status derivation, exact Studio capsule-link subjects, and bounded pages of at most 100 entries. The v3 migration changes only subject capacity; strict v2 journals retain their original subject limit and migrate on the next write. export_journal returns the versioned native structured page by default or an in-memory Open Knowledge Format v0.2 bundle when requested; in either form the export remains inside the tool result and creates no file, so there is no project-controlled export path to clean or disclose. erase_journal removes the managed journal plus owned lock, recovery, and orphan temporary files, then reports its bounded residue inventory. It explicitly does not claim erasure from storage media, external backups, or a user-created copy outside project control.

Reproduce the real two-process acceptance from a clean checkout:

cargo test -p numinous-mcp --test stdio_session returning_journal_survives_two_processes_then_leaves_zero_managed_residue

The test launches the built stdio server twice against one fresh explicit NUMINOUS_JOURNAL path. It asserts empty, record, connection, reconnect, inspect, append-only correction, corrected flagship reuse, structured export, confirmed erase, empty reread, and zero managed file or sidecar residue. Core and handler regressions separately cover prototype and strict v2 migration, v3 parsing, field escaping, missing or repeated correction targets, pagination, exact creation-link roundtrip, and the confirmation boundary. CI reruns all of them on a clean checkout. This is machine acceptance, not evidence of durable human memory or consciousness.

Data governance

  • Obtain explicit participation and publication consent before collection.
  • The pre-consent start receipt contains only a commitment, pseudonymous witness, time, mechanism, and external locator. It remains after decline so attempted starts can be reconciled, while the cohort retains only the aggregate model-family refusal count and no participant response content. The consent text must disclose this boundary before accepting participation.
  • Use opaque study identifiers. Do not collect names, account identifiers, private prompts, hidden reasoning, unrelated host data, or affect unless a separate protocol requires it.
  • Keep raw working captures in the gitignored .agent/ tree. Track only the minimum sanitized evidence needed to reproduce the published result.
  • Allowlist every retained field and reject unknown metadata, absolute paths, obvious email, IP, credential, and host-identifier values. A manual privacy review of the final bounded package remains mandatory because pattern checks cannot certify that arbitrary public text contains no identity.
  • Give a participant a withdrawal path before aggregation. Record withdrawals without retaining the withdrawn response in collector-managed storage. This does not erase provider logs, operator captures, terminal scrollback, or any participant-owned copy outside the collector boundary; the consent text and final report must name those limits.
  • Participant stop removes every answer and rationale, including a Formula Jam construction copied into a tool argument or result. Its call receipt retains only an explicit erasure marker, source-bound build evidence, and public call coordinates needed to validate the interrupted encounter prefix.
  • Keep study data separate from Journey, scores, broadcast, and the experience journal. No study event updates progression.

These constraints follow NIST's privacy-risk-management posture and MCP's explicit-consent and user-control principles. They are implementation requirements, not a privacy certification.

Required tracked evidence

The qualifying study must add one bounded directory at docs/evidence/understanding-0.4/ containing:

  • README.md: build, sample, dates, consent boundary, method, limitations, and exact reproduction commands.
  • allocation.json: frozen pair, condition, room-order, seed, and model-family assignments without personal identifiers.
  • probe-bank.json: the exact concealed bank released after collection; its SHA-256 hash must match the pre-collection commitment in allocation.json.
  • responses.jsonl: sanitized visible responses and scores, or a documented aggregate substitute if a participant does not consent to raw publication.
  • attempt-start-receipts.jsonl: every bounded calibration and collection start receipt, keyed only by its sealed commitment and opaque schedule or session identity, with each external locator retained for reconciliation.
  • mcp-build-receipt.json: the one source, toolchain, explicit target, private artifact, and binary digest receipt shared by every qualifying tool call.
  • report.md: every primary and secondary outcome, uncertainty interval, exclusions, deviations, failures, and null or negative results.
  • journal-acceptance.json: the returning-player structured receipts and managed-path residue inventory.
  • publication-audit.json: the independent allocation, start-receipt locator, and ledger reconciliation, privacy sign-off, evidence hashes, and reviewer decisions.

The report must identify the exact tracked runner and probe-bank revision. A private note, simulated persona reaction, manually written conclusion, or successful journal read does not satisfy this evidence contract.

Implementation order and acceptance

  1. Externally register the exact protocol, analysis, clean committed source-tree commitment, receipt schema, witness or append-only log identity, and completeness reconciliation procedure. This must happen before calibration ordinal 1.
  2. Drive calibration only through understanding-collect.py calibration-next and calibration-respond, using a new context for every frozen cell. Use calibration-recover only after verifying that the recorded process owner is dead. Pass the completed receipt ledger and its anchor to understanding-study.py calibrate; never construct response records by hand. For each new delivery, first call calibration-next without a receipt to obtain the oracle-free start commitment, independently record it, then retry with --start-receipt. The collector seals the receipt before exposure.
  3. Calibrate the private probes in fresh no-exposure contexts, replace weak items, and obtain two fresh independent passes over the implemented concealed-bank, stateful collector, isolated MCP, receipt-chain, scoring, redaction, interruption, withdrawal, and report boundary. Then commit the exact bank commitment and generated allocation before the first qualifying response.
  4. Complete the journal correction, structured export, and residue receipt, then prove the two-process return path with an isolated journal location.
  5. Before each allocated session, call start once without a receipt to obtain its oracle-free start commitment, independently record it, then retry with --start-receipt. Run the frozen cohort, publish all bounded evidence, obtain independent math and methodology review, and update the roadmap only from the published result.

The runner remains headless. If a later study or journal control becomes a visible App surface, it must reuse the design continuity gate in VISUALS.md: the near-black stage, luminous geometry, restrained shared chrome, semantic color, common typography and spacing, causal motion, reduced-motion behavior, and color-independent state cues. A generic form shell or one-off visual style does not belong in Numinous.

Current sources