Skip to content

MCP tools

superwitness serves an MCP server at /mcp on its own address, over MCP’s streamable HTTP transport. It is stateless: every request stands alone, with no session to open or keep. It takes the same bearer token as the HTTP API, and every tool acts as the principal that token names.

Point an MCP client at it with the token as a header, for example:

{
"url": "https://superwitness.example.internal/mcp",
"headers": { "Authorization": "Bearer <token>" }
}

The server names itself superwitness, with the running release as its version. It has four tools, thin layers over the HTTP operations for runs, spans, logs and verdicts. The fifth route, GET /v1/runs/by-attempt/{attempt}, is HTTP-only and has no tool.

The RunDocument for one work run: attempts with fingerprints, joined traces, errors, gate and recorded verdicts, cost, and per-source status. The same as GET /v1/runs/superpipeline/{board}/{run}.

Argument Required Description
source no the run’s source system; only superpipeline, which is the default
board_id yes superpipeline board id, brd_…
run_id yes superpipeline run id, run_…

Result: the run document.

One page of a run’s spans, oldest first. The same as GET /v1/runs/superpipeline/{board}/{run}/spans.

Argument Required Description
source no superpipeline, the default
board_id yes superpipeline board id, brd_…
run_id yes superpipeline run id, run_…
cursor no next_cursor from the previous page
limit no page size from 1 to 500; default 100

Result: {"spans": [...], "next_cursor": "…"}, with next_cursor left out on the last page.

One page of a run’s log lines, matched by run.id and by the run’s trace ids, oldest first. The same as GET /v1/runs/superpipeline/{board}/{run}/logs.

Argument Required Description
source no superpipeline, the default
board_id yes superpipeline board id, brd_…
run_id yes superpipeline run id, run_…
cursor no next_cursor from the previous page
level no debug, info, warn or error; empty for all
limit no page size from 1 to 500; default 100

Result: {"logs": [...], "next_cursor": "…" or null, "trace_join": "<status>"}.

Record a verdict as the calling principal. Append-only; to correct one, record a new verdict that supersedes it. The same as POST /v1/verdicts; Verdicts has the rules.

Argument Required Description
idempotency_key yes a key you choose; a retry with the same key returns the original verdict
kind no grader, review, eval or calibration; defaults to review for humans and grader otherwise
subject_kind yes run, attempt or eval_case_run
subject_ref yes superpipeline:<board>/<run>, attempt_<id>, or case:<id>@sha256:<fingerprint>
standard yes rubric:<id>@<version>, stage:<key> or case:<id>
value yes exactly one of {decision}, {score 0..1}, {label}, {text}
comment no free text, at most 10000 characters
evidence_refs no span ids or {session_id, seq_from, seq_to}
supersedes no the id of your earlier verdict this one corrects

Result: the verdict, whether it was just recorded or returned for a retry. Unlike the HTTP body, the tool has no judge or judge_kind argument: the judge is always the caller.

A tool’s result is the same JSON the HTTP route returns, twice over: as text content, and as structured content.

A refusal is a tool result marked as an error (isError), whose text content is the same error object the HTTP API returns, with the same codes:

{"error":{"code":"run_not_found","message":"neither superpipeline nor the hub ledger knows superpipeline:brd_01/run_02"}}

There is one code the HTTP API has no use for: invalid_source, when source is anything but superpipeline.

Each tool’s input schema is published with it, with the descriptions above. Arguments are checked against it before the tool runs: a missing required argument, or one the schema does not list, is also an error result, whose text is a plain validation message rather than an error object.