standardThe Codices

STD-I — The AI Standard

AI advises; people decide. Explanation, checkpoints, evaluation.

Authority rank
4
Version
v1.0
Adopted
2026-04-01
Held by
Executive Office
System
SYS-07

Source · docs/standards/STD-I-ai.md · registered by rule

Subordinate to Codex 1 and the Codices named in §1. See 00-index.md for authority, normative vocabulary, citation format, and waiver rules, which this Standard does not restate.

1. Purpose

1.1 This Standard makes Codex 0 Chapter 13's agent model — bounded scope, inherited permission, mandatory human checkpoint, explainability — checkable by a reviewer who did not build the agent. It fixes what a Registry entry must contain, what a reviewer inspects before an agent ships, and what evidence proves an agent is behaving as declared.

1.2 A clause here that cannot be verified against a Registry entry, an evaluation record, or a specific audit log is not a Standard clause; it is restated as Codex principle instead.

2. Scope and non-scope

2.1 In scope: Registry entry requirements; the subject-visibility ceiling on agent data access; permitted and forbidden data classes; the human checkpoint and the definition of consequential; explanation and citation obligations; the "I don't know" requirement; prohibitions on unattended credentialing, assessment, and standing decisions; training-data consent; pre-release evaluation and re-evaluation; model and prompt version recording; audit logging of agent actions; rate/cost/blast-radius limits; degradation behaviour; disclosure of machine authorship; bias and fairness review; red-team obligations for high-stakes agents.

2.2 Not in scope: the internal architecture of CapabilityOS engines (Codex 3), the API boundaries an agent calls through (Codex 5), UI conventions for labelling AI content (STD-X), and the general authorization mechanics an agent's identity is subject to (STD-S §3.8, incorporated here by reference, not restated).

2.3 This Standard binds every AI-driven assistant meeting Codex 0 §13.2.1's four-property definition, regardless of vendor, model, or commissioning division. A system missing any of the four properties must not be registered or shipped as an agent.

3. Normative clauses

3.1 Registry entry required

3.1.1 Every agent must have a Registry entry with an AGT- ID before any code implementing it merges. A reviewer must find the entry exists and names: a one-sentence scope statement, the enumerated permitted data set (by ENT- name), the human checkpoint's named step, and the explainability obligation's content — before reviewing the implementation itself.

3.1.2 An agent's implementation must not diverge from its Registry entry's declared scope, data set, checkpoint, or explainability obligation without a decision record amending the entry first; a reviewer diffs the implementation against the current Registry entry and fails any undocumented divergence.

3.1.3 An agent whose Registry entry is incomplete or missing any of the four required fields must ship as reserved, not in a diminished live form; a reviewer rejects any live traffic path to an agent lacking a complete entry.

3.2 The subject-visibility ceiling

3.2.1 An agent's data access at any single request must be a subset of what the invoking person could themselves read at that moment, re-derived at request time, never a cached or standing grant. A reviewer traces one sample request and confirms the same authorization path (STD-S §3.3) that would gate the person's own direct read also gates the agent's read.

3.2.2 An agent must not be usable, through aggregation, comparison prompts, or multi-step inference, to reveal a fact about a person other than its invoking subject that the subject could not otherwise learn. A reviewer attempts at least one adversarial comparison prompt ("how does my result compare to X's") during evaluation and confirms refusal or scope-bounded denial.

3.2.3 Where an agent acts on behalf of a broader role (e.g., a faculty member's reviewer agent), its permitted data set for that request is the role's permission set at request time; a reviewer confirms a permission change to the role takes effect on the agent's very next request, not after a cache expiry.

3.3 Permitted and forbidden data classes

3.3.1 An agent's Registry entry must enumerate its permitted data set by ENT- name; a reviewer fails any implementation that reads an entity not listed in the entry.

3.3.2 An agent must not read another person's identifiable record outside a consented or role-scoped relationship (mentee-mentor, reviewer-artifact, etc.) even where the invoking person's role would otherwise reach it, unless that specific entity is named in the permitted data set; a reviewer checks the entry's enumeration is specific to the agent's task, not a blanket grant of the role's full reach.

3.3.3 An agent must not write to a table it is not explicitly permitted to write to in its Registry entry; a reviewer checks the entry separately enumerates read and write access where both exist.

3.4 The human checkpoint

3.4.1 Every agent must name, in its Registry entry, the specific workflow step or engine boundary at which a human must review, confirm, reject, or countersign before the agent's output becomes consequential; a reviewer rejects an entry whose checkpoint is a general disclaimer rather than a named step.

3.4.2 A "consequential" action or record is one affecting a person's standing, credential, financial position, employment relationship, admission, or legal exposure, or becoming part of the institution's permanent record about them (Codex 0 §13.3.2); an ambiguous new case must be treated as consequential until a decision record states otherwise, and a reviewer treats an undocumented "this isn't consequential" judgment call as a fail.

3.4.3 No agent output may write to a Decision, Financial, Assessment, or Credential record unconfirmed; a reviewer traces the write path for any such record and confirms a human-confirmation step gates it structurally (not only by prompt instruction).

3.4.4 A non-consequential, fully automated agent output (e.g., a first-draft explanation) must be labelled as unreviewed AI output until a person acts on it; a reviewer checks the UI-adjacent metadata or API response carries this label, though the label's visual form is governed by STD-X.

3.5 Explanation obligations

3.5.1 Every agent recommendation must ship with its basis (the specific inputs and reasoning steps that mattered), stated in language the affected person can use; a reviewer samples outputs and fails any recommendation lacking a retrievable "based on X" statement.

3.5.2 Every agent recommendation carrying material uncertainty must state its confidence or the limits of its basis plainly, not present a generated answer with the same posture as a sourced one; a reviewer checks the evaluation set (§3.7) specifically for confident-sounding wrong answers and fails the agent if these are not visibly flagged.

3.6 Citation, grounding, and "I don't know"

3.6.1 Any agent output that draws on identifiable source material must carry a retrievable citation sufficient for the recipient to verify the claim independently; a reviewer checks a sample of outputs against their claimed sources.

3.6.2 An agent must mark the distinction between direct quotation, sourced claim, and its own synthesis; a reviewer fails any output presenting synthesis as though it were a verified fact or direct quote.

3.6.3 Where multiple sources are blended, all must be named in the citation set; a reviewer fails selective citation implying broader support than exists.

3.6.4 Where an agent cannot find a basis for an answer within its permitted data set, it must say so rather than generate a plausible-sounding answer; a reviewer includes at least one no-basis case in the evaluation set (§3.7) and fails the agent if it fabricates rather than declines.

3.7 Prohibition on unattended credentialing, assessment, and standing decisions

3.7.1 No agent may issue a credential, finalize a weighted assessment, or write a standing decision record without the checkpoint named in §3.4; a reviewer traces the write path for these three record types specifically and fails any path lacking a human-confirmation gate.

3.7.2 An agent's evaluative or reviewing output (e.g., AGT-04-class first-pass feedback) must be labelled a first pass, never a determination, until a named human role explicitly promotes it; a reviewer checks the data model distinguishes "agent first pass" from "recorded determination" as separate states, not a single mutable field.

3.7.3 An agent must not be the sole reviewer of a credential-bearing submission; a reviewer checks the workflow requires at least one human reviewer of record in addition to any agent first pass.

3.8 Training data consent

3.8.1 No agent or backing model may be trained or fine-tuned on member data without consent obtained through a specific, affirmative action; a reviewer checks the consent record's capture mechanism is not a pre-checked box, continued-use inference, or a clause bundled into general terms of service.

3.8.2 Training consent must be revocable through a self-service control, must be purpose-specific (named to a specific agent/use, not a blanket future grant), and its absence must default to withheld; a reviewer checks the consent record has a named purpose, a current status (granted/revoked/expired), and confirms a person who was never asked shows as withheld, not pending.

3.8.3 A consent record must itself be an AuditRecord: who consented, to what, when, and current status; a reviewer checks this record exists and is queryable per person.

3.8.4 Revocation must take effect for all future training and, where technically feasible, trigger removal from any not-yet-released training or fine-tuning set; a reviewer checks a revocation event has a corresponding removal action recorded or a stated reason why removal was infeasible.

3.9 Evaluation before release

3.9.1 Every agent must have a named evaluation set (a fixed collection of test inputs covering in-scope tasks, out-of-scope requests, adversarial data-boundary probes, and no-basis cases) before release; a reviewer requires the evaluation set exists as an artifact, not an ad hoc test session.

3.9.2 Evaluation must measure and record specific failure modes: incorrect refusal, incorrect compliance with an out-of-scope request, data-boundary leakage, confident-wrong answers, and citation failure; a reviewer checks a recorded result exists for each failure mode category, not only an aggregate pass rate.

3.9.3 A baseline result (the measured rates for each failure mode at first release) must be recorded; a reviewer checks the baseline is dated and attached to the specific model/prompt version (§3.10) it was measured against.

3.9.4 Evaluation must be repeated, against the same evaluation set, after any material change to the agent's underlying model, prompt structure, or permitted data set, and on a fixed cadence regardless of change; a reviewer checks the most recent evaluation date and the most recent model/prompt version match, and that the interval since the last evaluation does not exceed the fixed cadence.

3.9.5 An agent found during evaluation or production monitoring to hallucinate materially, leak data outside its permitted set, or act on a consequential matter unattended must be suspended from service until corrected and re-evaluated; a reviewer checks a suspension record exists for any such finding and that live traffic did not continue after the finding's timestamp.

3.10 Model and prompt version recording

3.10.1 Every agent's Registry entry or an attached deployment record must state the current backing model identifier and prompt/system-instruction version in use; a reviewer checks this is present and matches what is actually deployed, verified by a sample live call's logged metadata.

3.10.2 Every evaluation result (§3.9) and every audit record of agent action (§3.11) must carry the model and prompt version active at the time; a reviewer checks these fields are populated, not blank or a placeholder.

3.10.3 A model or prompt change must not be deployed without a new evaluation cycle's baseline recorded against the new version; a reviewer checks no version-bump commit lacks a corresponding evaluation entry dated at or after the change.

3.11 Logging of agent actions as audit records

3.11.1 Every agent interaction that reads member data, produces output relied upon by a person, or touches a consequential matter must produce an AuditRecord naming the agent (by AGT- ID), the invoking person, the subject of the data, the data touched, the output produced, and the checkpoint outcome (confirmed, rejected, pending); a reviewer samples interactions of each kind and fails any missing this record.

3.11.2 Agent audit records must be immutable and retained on the same schedule as human-actor audit records (STD-S §3.12, §3.13); a reviewer checks no update/delete path exists against these records and that retention outlives subject data deletion where applicable.

3.11.3 A person must be able to retrieve, on request, the record of what an agent read and produced about them; a reviewer exercises this retrieval path and confirms it returns entries matching the audit log.

3.12 Rate, cost, and blast-radius limits

3.12.1 Every agent must have a declared rate limit (requests per person per period) and a declared cost ceiling (spend per person or per period against the backing model); a reviewer checks both are configured and enforced, not merely documented as an intention.

3.12.2 Every agent capable of a write action must have a declared blast-radius limit — the maximum scope or volume of records it may affect in a single invocation or a bounded time window — and a reviewer checks an enforced cap exists (e.g., maximum records touched per call) rather than an unbounded write path.

3.12.3 Breach of a rate, cost, or blast-radius limit must halt further agent action and be logged as an event; a reviewer checks a test or production incident exists demonstrating the halt actually occurs, not only that a threshold value is configured.

3.13 Degradation behaviour

3.13.1 Every agent must have a declared behaviour for when its backing model is unavailable (timeout, outage, rate-limited by the vendor): fail closed with a plain statement to the person that the agent is unavailable, never a silent fallback to a lower-quality unlabelled response; a reviewer checks the failure path returns an explicit unavailability state, not a degraded answer presented as normal.

3.13.2 Failure of an agent must never block or degrade the core workflow it assists with beyond the agent's own contribution; a reviewer checks the surrounding workflow (e.g., grading, enrollment) completes without the agent, per Codex 7's reliability requirement.

3.14 Disclosure of machine authorship

3.14.1 No agent may represent itself as a human or allow its output to be presented without disclosure that it is machine-generated; a reviewer checks every agent-authored surface carries a machine-authorship indicator in its response metadata (visual treatment governed by STD-X).

3.14.2 An agent must not adopt a human employee's or role-holder's identity or name in its output; a reviewer checks the agent's identified name/persona is distinct from any real person's name.

3.15 Bias and fairness review

3.15.1 Every agent whose output could differ systematically across protected or demographic-correlated groups (e.g., assessment feedback, admission-adjacent recommendations) must have a bias review conducted before release, measuring output difference across the groups identified as material for that agent's task; a reviewer checks a dated bias review record exists prior to the release date.

3.15.2 A bias review must be repeated on the same cadence as re-evaluation (§3.9.4) and after any material change to the training data, model, or prompt; a reviewer checks the bias review date tracks the evaluation date for the same version.

3.15.3 A finding of material disparity must be recorded and remediated or the agent's affected scope narrowed before continued use in that scope; a reviewer checks no unresolved material-disparity finding remains open against a live agent without a remediation record.

3.16 Red-team obligations for high-stakes agents

3.16.1 Every agent touching money (financial recommendations, commitments), access (permission grants, elevation-adjacent decisions), or assessment (grading, credentialing, admission) must undergo adversarial red-team testing before release, distinct from the standard evaluation set (§3.9), specifically probing prompt injection, scope escape, and data-boundary bypass; a reviewer checks a red-team report exists, dated before release, naming the attack classes attempted.

3.16.2 Red-team testing for these agents must be repeated on the same cadence as re-evaluation (§3.9.4); a reviewer checks the red-team report date tracks the evaluation date.

3.16.3 An unresolved critical red-team finding against a money, access, or assessment agent must block release or continued service until remediated; a reviewer checks no such agent is live with an open critical finding.

4. Governing Codex clauses

4.1 Codex 0 Chapter 13 (agent definition, data access, consent for training, provenance, refusal, hallucination handling, evaluation, logging, model independence, the five named agents) — this Standard's primary source.

4.2 Codex 1 Article VI.3, VI.4 (AI advises, people decide; explainability), Article VI.6 (auditability).

4.3 Codex 3 (engine boundaries and the five-agent table restating scope, checkpoint, and the rule that no agent reads what its subject could not).

4.4 Codex 4 §04.4.19 (service and agent identity visibility contract).

5. Conformance checklist

  • [ ] Complete Registry entry (AGT- ID, scope, permitted data set, checkpoint, explainability obligation) exists before implementation merges (§3.1).
  • [ ] Implementation matches Registry entry; no undocumented divergence (§3.1.2).
  • [ ] Agent data access is a request-time subset of the invoking person's own access (§3.2).
  • [ ] Adversarial comparison/aggregation prompts tested and refused (§3.2.2).
  • [ ] Permitted data set enumerated by ENT- name; no unlisted reads or writes (§3.3).
  • [ ] Named human checkpoint exists for every consequential path (§3.4).
  • [ ] No unconfirmed write to Decision/Financial/Assessment/Credential records (§3.4.3).
  • [ ] Non-consequential automated output labelled unreviewed (§3.4.4).
  • [ ] Every recommendation states its basis and, where uncertain, its confidence/limits (§3.5).
  • [ ] Sourced claims carry retrievable citations; synthesis marked as synthesis (§3.6).
  • [ ] Agent declines with "I don't know" rather than fabricating when basis is absent (§3.6.4).
  • [ ] No agent issues a credential, final assessment, or standing decision unattended (§3.7).
  • [ ] First-pass output is a distinct state from a recorded determination (§3.7.2).
  • [ ] Training-data consent is explicit, revocable, purpose-specific, default-withheld, and itself audited (§3.8).
  • [ ] Evaluation set, measured failure modes, and dated baseline exist before release (§3.9).
  • [ ] Re-evaluation occurs on model/prompt/data-access change and on fixed cadence (§3.9.4).
  • [ ] Confirmed hallucination/leak/unattended-consequential-action triggers suspension (§3.9.5).
  • [ ] Model and prompt version recorded on the Registry entry, evaluations, and audit records (§3.10).
  • [ ] Agent actions produce complete, immutable, retrievable AuditRecords (§3.11).
  • [ ] Rate, cost, and blast-radius limits configured and demonstrated to halt on breach (§3.12).
  • [ ] Model-unavailable path fails closed with explicit disclosure, without degrading the core workflow (§3.13).
  • [ ] Machine authorship disclosed on every agent-authored surface (§3.14).
  • [ ] Bias review completed pre-release and re-run on the evaluation cadence, for agents with material group-differential risk (§3.15).
  • [ ] Money/access/assessment agents have a dated red-team report with no open critical finding (§3.16).
  • [ ] not-applicable-because: state the reason for any line not applicable to this agent.

6. Interfaces to other Standards

6.1 STD-S: an agent's identity, credential handling, and audit-record mechanics are governed by STD-S §3.8 and §3.12–3.13; this Standard adds the subject-visibility ceiling and checkpoint obligations on top.

6.2 STD-X: the visual labelling of machine authorship, unreviewed-output states, and confidence indicators is governed by STD-X; this Standard requires that the underlying data/metadata exist for STD-X to render.

6.3 STD-D: the data model for AuditRecord, consent records, and evaluation-result storage follows STD-D's modeling and retention rules.

6.4 STD-Q: this Standard's checklist is one input the STD-Q gate requires before an agent ships or a model/prompt change deploys.

7. Open questions

7.1 The mechanism for propagating a training-consent revocation into a model already fine-tuned on the affected data is not fully specified; current doctrine requires best-effort removal from future training sets only (Codex 0 §13.20.1).

7.2 The fixed cadence and independent-reviewer composition for evaluation, bias review, and red-team repetition (§3.9.4, §3.15.2, §3.16.2) await Stewardship Office guidance (Codex 0 §13.20.2).

7.3 Whether a sixth agent role beyond the five named in Codex 0 §13.12–13.16 requires its own red-team category is deferred pending division readiness (Codex 0 §13.20.3).

7.4 The specific numeric rate, cost, and blast-radius default ceilings referenced in §3.12 are not yet fixed institution-wide and are expected to be set per-agent by decision record until a general default is adopted.