Verdicts
A verdict is a judgement about a run, an attempt or an eval case run, against a named standard. superwitness stores verdicts in its own Postgres and never changes one once it is written. Concepts introduces them; this page is how to record them.
Record a verdict
Section titled “Record a verdict”Post one JSON object to POST /v1/verdicts, or call the record_verdict MCP tool with the same
fields:
curl -X POST -H 'Authorization: Bearer <token>' -H 'Content-Type: application/json' \ https://superwitness.example.internal/v1/verdicts -d '{ "idempotency_key": "review-brd_01-run_01-1", "subject_kind": "run", "subject_ref": "superpipeline:brd_01/run_01", "standard": "stage:draft", "value": {"decision": "approved"}, "comment": "Reads well." }'| Field | Required | Meaning |
|---|---|---|
idempotency_key |
yes | a key you choose, 1 to 200 bytes; see Retries |
kind |
no | grader, review, eval or calibration; left out, review for a human judge and grader for anyone else |
subject_kind |
yes | run, attempt or eval_case_run |
subject_ref |
yes | superpipeline:<board>/<run>, attempt_<id>, or case:<id>@sha256:<64 hex> to match subject_kind |
standard |
yes | what the verdict judges against; see Standards |
value |
yes | the judgement; see Values |
comment |
no | free text, at most 10000 characters |
evidence_refs |
no | up to 100 pointers to evidence; see Evidence references |
supersedes |
no | the id of your earlier verdict that this one corrects |
The HTTP body also accepts judge, which must equal the caller’s own principal id if it is
present, and refuses judge_kind outright (400 judge_kind_not_accepted): the judge is always
the caller, and its kind comes from the caller’s principal record. Any other field is refused as
400 invalid_json, as is a body over 64 KiB or one holding more than one JSON value.
A kind outside the four is 400 invalid_kind; a subject_kind outside the three is 400
invalid_subject_kind; a subject_ref that does not fit its kind is 400 invalid_subject.
A new verdict answers 201 with the stored verdict:
{ "id": "vrd_…", "idempotency_key": "review-brd_01-run_01-1", "kind": "review", "subject_kind": "run", "subject_ref": "superpipeline:brd_01/run_01", "judge": "prn_human01", "judge_kind": "human", "standard": "stage:draft", "value": {"decision": "approved"}, "comment": "Reads well.", "evidence_refs": [], "supersedes": null, "created_at": "…"}comment is "" and evidence_refs is [] when none were given. Recorded verdicts on the run
and its attempts appear in the run document’s verdicts, after the run’s gates.
Values
Section titled “Values”value is an object with exactly one key:
| Value | Rule |
|---|---|
{"decision": "…"} |
a non-empty string |
{"score": 0.7} |
a number from 0 to 1 |
{"label": "…"} |
a non-empty string |
{"text": "…"} |
a non-empty string |
Anything else is 400 invalid_value, as is a string holding NUL or an unpaired UTF-16 surrogate
escape.
Standards
Section titled “Standards”Every verdict names a standard. Leaving it out is 422 missing_standard; a standard in none of
these forms is 422 invalid_standard:
| Standard | Form |
|---|---|
rubric:<id>@<version> |
a rubric recorded in superwitness: id of letters, digits, _, . and -, up to 64; version a whole number from 1 |
stage:<key> |
a pipeline stage: key of letters, digits, _, . and -, up to 64 |
case:<id> |
an eval case: id of letters, digits, _, . and -, up to 128 |
A rubric standard must name a rubric version that exists, or the verdict is 422
unknown_rubric.
Rubrics
Section titled “Rubrics”Rubrics are recorded with the binary, against the same database:
superwitness rubric-add -id press -version 1 -name Press \ -scale '{"min":0,"max":1}' -body-file press.md -created-by prn_human01All six flags are required: -id, -version (1 or more), -name, -scale (valid JSON),
-body-file (a file holding the rubric text) and -created-by (the author’s principal id).
rubric-add needs SW_DATABASE_URL. It runs any pending migrations first, over
SW_MIGRATE_DATABASE_URL when that is set, then inserts the rubric over SW_DATABASE_URL and
prints rubric:press@1 recorded. Give rubrics ids that a standard can name: letters, digits,
_, . and -, up to 64.
Rubrics are append-only, like verdicts. A version, once recorded, cannot be changed or
recorded again; rubric-add fails with rubric version already exists. To change a rubric,
record the next version and judge against rubric:press@2.
Subjects and who may judge them
Section titled “Subjects and who may judge them”- A
runorattemptsubject must exist. superwitness checks with superpipeline and the AgentPod hub; a subject neither knows is 404subject_not_found, and one it cannot check is 503subject_unresolved. - An
eval_case_runsubject is not checked against any product. - A non-human judge may not judge a run or attempt it executed: 403
self_judgement. The executors are the run’s agent principal from superpipeline and each attempt’s agent principal from AgentPod. If superwitness cannot establish every one of them as a hub principal, because a source did not answer or an executor is not mapped to a principal yet, a non-human verdict on that subject is 503subject_unresolvedrather than accepted on incomplete information. - A human judge may judge any subject, including one they executed.
The judge’s kind is never taken from the request. A human principal records as human, an
agent as agent, and a service as grader.
Corrections: append-only
Section titled “Corrections: append-only”Nothing about a verdict can change once it is written. In the database, triggers refuse
UPDATE and DELETE on each row of verdicts and rubrics, and TRUNCATE on either table.
The runtime role has only SELECT and INSERT on the two tables, and is
not their owner, so it cannot drop the triggers either (Install).
A change of mind is a new verdict that names the old one in supersedes:
{ "idempotency_key": "review-brd_01-run_01-2", "subject_kind": "run", "subject_ref": "superpipeline:brd_01/run_01", "standard": "stage:draft", "value": {"decision": "changes_requested"}, "supersedes": "vrd_…"}- Only the judge of the old verdict may supersede it: anyone else gets 403
not_original_judge. - The correction must name the same subject and standard: 422
supersedes_mismatch. - The old verdict must exist: 422
supersedes_not_found.supersedesmust look like a verdict id (vrd_…): 400invalid_supersedes. - Each verdict is superseded at most once, so a chain stays a line. Superseding a verdict that
already has a successor is 409
already_superseded; supersede the latest in the chain instead.
The run document shows only the latest verdict in each chain.
Retries and idempotency
Section titled “Retries and idempotency”Every verdict carries an idempotency_key, and a key is used once across the whole store. Send
the same request again with the same key, and superwitness returns the verdict it already
stored, with 200 instead of 201, and writes nothing. Send a different verdict with a key that
is already used, by you or anyone else, and it is 409 idempotency_conflict.
So on a timeout or a retryable error, retry with the same key. The retryable errors are 503
store_unavailable and 503 subject_unresolved; both carry "retryable": true and a
Retry-After: 5 header.
Evidence references
Section titled “Evidence references”evidence_refs points at evidence; it never holds it. It is an array of at most 100 entries,
each either:
- a span id: 16 lowercase hex characters, such as
"00f067aa0ba902b8"; or - a range of session events:
{"session_id": "…", "seq_from": 1, "seq_to": 9}, withseq_from0 or more andseq_tonot below it, and no other keys.
Anything else is 400 invalid_evidence_refs.
When the database is down
Section titled “When the database is down”serve runs its migrations in the background at start. Until they succeed, and whenever
Postgres cannot answer, POST /v1/verdicts is 503 store_unavailable, run documents show
sources.verdicts as unavailable (or timeout), and the rest of the document still renders.