Closeout Specification, Draft 0.1
Status: Draft. An evaluator can implement this edition. A later edition that changes behavior will use a new specVersion.
1. Purpose
Closeout is an open specification for declaring what must be verified before an agent's work can be accepted, and recording the evidence used to make that decision.
A skill explains how to perform work. A closeout policy makes a requirement mandatory at a gate. Installing a skill does not add a requirement. An agent that never loads a skill does not remove a requirement.
The interoperability rule is:
Given the same effective policy, candidate identity, and evidence, compliant evaluators reach the same acceptance decision.
Two reviewers do not have to produce the same findings. Evaluators that receive the same findings have to agree on whether those findings satisfy the policy.
2. Files
A repository keeps policy in version control and keeps results out of it.
| Path | Role |
|---|---|
AGENTS.md |
Project instructions. Point at the policy. Do not copy the policy into prose. |
.agents/skills/ |
Procedures for doing or reviewing work. Agent Skills packaging: SKILL.md, plus optional references/, scripts/, and assets/. |
.agents/closeout.yaml |
Public entry point. The acceptance requirements. |
.agents/closeout/*.yaml |
Policy documents imported by the entry point. |
.closeout/ |
Evidence and decisions written by a runner. Not policy. |
Paths inside a policy are relative to the repository root. A path is one or more segments separated by /. A segment is not empty, ., or ... A path does not start with / or ~, and it does not contain a backslash. The resolved path stays inside the repository root.
~/.agents is not a policy source. Personal files there do not add, remove, or replace a repository requirement.
The entry is .agents/closeout.yaml.
When that file does not exist, the repository has no requirements. validate succeeds. A decision for any gate is accepted, with message no closeout policy.
When a file exists and loading or validation fails, acceptance is blocked. The failure is not treated as zero requirements.
3. Policy document
A public policy file is YAML 1.2. Duplicate keys are an error. The JSON value after parsing matches schema/policy.schema.json.
specVersion: "0.1"
description: Requires Bun and a typecheck script.
imports:
- path: .agents/closeout/bun-quality.yaml
as: quality
items:
- id: adversarial-review
kind: review
gate: beforePR
skill: .agents/skills/adversarial-review/SKILL.md
independence:
differentSession: true
differentModel: true
failOn: P1
specVersion is the string 0.1. Unknown fields are errors. description is optional prose for humans, including prerequisites of a shared collection. It does not change evaluation except through the file hash in the policy identity.
imports and items default to empty arrays. Every item that resolves is required at its gate. This edition has no optional item.
An imported file uses the same schema. This edition defines one gate, beforePR.
Item ids and import names match ^[a-z][a-z0-9]*(-[a-z0-9]+)*$.
4. Imports
An import adds that file's requirements under a namespace. There is no override and no last-file-wins rule.
Resolution is depth-first. For an import { path, as } inside a document whose namespace prefix is P:
- The child prefix is
aswhenPis empty, otherwiseP/as. - The child's imports are resolved first.
- Each of the child's items is then added with id
childPrefix/item.id.
Items of the entry file keep their own ids. Import path values are repository-root-relative, including imports nested inside another file.
These are errors:
- An import path that breaks section 2.
- A missing import file.
- An import cycle. The cycle is the chain of policy files.
- Two imports in the same file with the same
as. - Two resolved items with the same qualified id.
This edition imports only files inside the same repository. A shared collection is vendored into the repository and imported by path.
5. Requirements
5.1 command
id: typecheck
kind: command
gate: beforePR
exec: ["bun", "run", "typecheck"]
timeoutSeconds: 120
exec is an argv array. The runner does not invoke a shell. timeoutSeconds is required, from 1 to 86400.
The runner executes the argv in a new detached git worktree checked out at the candidate commit, then deletes that worktree. The command's exit, timeout, and whether it left the worktree dirty or moved HEAD, are the evidence. A statement in a chat transcript is not command evidence.
A shared collection names its prerequisites in description. A command that assumes a package manager does not become valid in a repository that does not provide it.
5.2 review
id: adversarial-review
kind: review
gate: beforePR
skill: .agents/skills/adversarial-review/SKILL.md
independence:
differentSession: true
differentModel: true
failOn: P1
skill is a repository-root-relative path ending in /SKILL.md. The file has to exist. The skill directory is part of the policy identity, including references/ and any other regular files under that directory. Symlinks inside the skill directory are an error.
failOn is P0, P1, P2, or P3. Severity order from most severe to least is P0, P1, P2, P3. A finding fails the requirement when its severity is as severe as failOn or more severe. failOn: P1 fails on P0 and P1.
independence.differentSession and independence.differentModel are required booleans. true means the reviewer's value and the candidate's value are both non-empty and different. An empty or unknown value does not satisfy true.
The candidate session, model, and provider are inputs supplied by the orchestrator. The reviewer session, model, and provider are fields on the review record. The YAML flags are requirements. The runner records the identities it compared. It does not infer them from a model name it cannot see.
A review record carries findings. It does not carry an outcome. The evaluator computes the outcome from the findings, failOn, and independence.
Each finding has:
| Field | Meaning |
|---|---|
severity |
P0, P1, P2, or P3. |
location |
Where the reviewer looked. |
explanation |
What is wrong, or what the reviewer checked. |
evidence |
The observation that supports the finding. |
6. Effective policy identity
The identity is sha256: plus 64 lowercase hex characters. It is the SHA-256 of the UTF-8 bytes of a canonical JSON document:
{
"specVersion": "0.1",
"files": [{ "path": "repo-relative path", "sha256": "64 hex characters of the file bytes" }],
"items": []
}
files is sorted by path. It contains each policy file that was read, and every regular file under each referenced skill directory, once each. items stays in execution order.
Canonical JSON for this edition:
- Objects have keys sorted by UTF-16 code unit, ascending.
- Arrays keep their order.
- Strings use JSON string encoding.
- The only numbers are safe integers, written in base 10 without an exponent.
- There is no insignificant whitespace.
A resolved command item is { id, kind, gate, exec, timeoutSeconds }. A resolved review item is { id, kind, gate, skill, independence, failOn }. independence contains differentSession and differentModel.
Changing a policy file, an imported file, or a file in a referenced skill directory changes the digest. Evidence bound to the previous digest does not satisfy the new one.
7. Evidence
A record matches a requirement only when itemId, base, head, and policyDigest are equal to the evaluation inputs. base and head are lowercase hex git commit ids, 40 or 64 characters, resolved by the runner before execution. attempt is a positive integer. Two records with the same identity and the same attempt are an error for that item.
The evaluator uses the matching record with the greatest attempt.
Command records are accepted only from a producer the evaluator has been configured to trust. The reference runner trusts producer name closeout-reference at its own version, and only when exec.argv equals the policy argv. A content hash binds the record to its bytes. It does not prove who wrote the file.
Review records that match the evidence schema are intake. The evaluator recomputes pass or fail. Any schema-invalid file in the evidence directory blocks the decision.
The reference runner stores records under .closeout/evidence/ and command logs under .closeout/logs/. Those directories are outside the committed policy.
8. Evaluation
For each resolved item at the requested gate, in order:
| Condition | Item state | Decision class |
|---|---|---|
Kind is not command or review |
unsupported |
blocked |
| No records for the item | missing |
blocked |
| Records exist, none match this base, head, and digest | stale |
blocked |
| Duplicate attempt among matching records | invalid |
blocked |
| Command producer is not trusted, or argv differs | untrusted or invalid |
blocked |
Command timed out, could not start, exited non-zero, moved HEAD, or left the worktree dirty |
failed |
rejected |
| Command exited 0 in a clean worktree still at the candidate | passed |
satisfied |
A finding meets failOn |
failed |
rejected |
| Independence is not satisfied | independence |
blocked |
Review findings do not meet failOn, and independence holds |
passed |
satisfied |
The decision is blocked when any item is in the blocked class, or when policy loading failed, or when the evidence directory is unreadable. Otherwise the decision is rejected when any item failed. Otherwise the decision is accepted.
A candidate commit that changed cannot reuse evidence recorded for the previous commit. That evidence is stale.
An unsupported required kind blocks acceptance. The runner does not skip it.
No reviewer identity that satisfies independence blocks acceptance. The runner does not substitute the candidate's own session or model.
9. Result and exit codes
The decision document matches schema/decision.schema.json. The reference runner prints it with --json. Without --json, it prints the decision, the commits, the policy digest, and one line per item.
| Exit | Meaning |
|---|---|
| 0 | accepted. Also validate when the policy is valid or absent. |
| 1 | rejected. A supported requirement failed. |
| 2 | The invocation is invalid. |
| 3 | blocked. Policy invalid, configuration conflict, unsupported kind, missing or stale evidence, independence unmet, or the commits do not resolve. |
Reference commands:
closeout validate
closeout run \
--gate beforePR \
--base <base-commit> \
--head <candidate-commit> \
--json
closeout decision \
--gate beforePR \
--base <base-commit> \
--head <candidate-commit> \
--candidate-session <id> \
--candidate-model <id> \
--json
closeout evidence add \
--gate beforePR \
--item <qualified-id> \
--base <base-commit> \
--head <candidate-commit> \
--session <reviewer-session> \
--model <reviewer-model> \
--findings findings.json
findings.json is an array of finding objects. evidence add writes the next attempt. It does not decide acceptance.
--candidate-session, --candidate-model, and --candidate-provider supply the candidate identity used for independence. Omitting them leaves the value empty, which does not satisfy a true independence flag.
A hook may consult closeout decision and hand the result back to an agent. A hook that allows the session to end does not write an accepted decision and does not grant permission to open or merge a pull request. The orchestrator checks the current decision at its own boundary.
Human review of a pull request is a separate decision from this document.
10. Conformance
conformance/<name>/repo/ is a sample repository. conformance/<name>/expect.json is the outcome a compliant evaluator reproduces.
check is validate, decision, or digest. A decision fixture with git: true is committed by the harness, and that commit is both base and head. Evidence templates may use $base, $head, $digest, and $version for those commits, the effective policy digest, and the reference runner version. A digest fixture names touch. The digest taken before and after appending a newline to that file must differ.
The cases cover malformed policy, duplicate ids, import cycles, import conflicts, path escape, stale evidence, a missing reviewer, independence, findings that meet failOn, and a referenced criteria file changing the digest.
The reference runner in this repository is the Rust binary closeout. cargo test runs these fixtures against it.