COLDWORKS/ docs / audit / cli
ReviewsMemoryGuardsThe auditScoreboardDocsGitHubAbout
Menu
ReviewsMemoryGuardsThe auditScoreboardDocsGitHubAbout
Sign in
Overview
Reviews docs
Getting started
  • Quickstart
Concepts
  • Risk routing
  • Defect labels
  • The cleared band
  • What Coldworks gets wrong
Reference
  • CLI · doug-backtest
  • The report
Coming up
  • MCP · Pattern Gardenplanned
  • REST APIpreview
Meta
  • Changelog
The audit docs · preview
Getting started
  • Quickstart
  • Connect your traces
Reference
  • The audit CLI
  • The audit reportsoon
  • Guards & subtractionsoon
View as llms.txt
The audit · ReferenceDesign preview · coldworks-audit has not shipped

The audit CLI

One command with one job: read a trace export on your disk, write an audit report next to it. This page documents the designed interface — it is the contract the build is held to, published before the build so you can hold it to the contract too.

coldworks-audit run

Signature
coldworks-audit run <PATH> [--docs DIR] [--diff] [--sample] [--open] [--format NAME]

Audits the export at PATH — a file, or a directory whose .json and .jsonl exports are audited together, with span-level dedup across overlapping files. An observation that arrives twice under one run and span id is counted once, and the report says how many it dropped. Where an export carries no span ids, the report says “dedup unavailable: this export carries no span ids; overlapping files count twice.”

ArgumentWhat it doesStatus
PATHThe export file or directory. Format is detected from the first record: OTel GenAI spans, Langfuse observations (verified on the API shape), or the span-tree shape --sample writes. Override with --format otlp_genai|langfuse|span_tree.BUILT · NOT RELEASED
--docs DIRAnchor recovered judgments to lines in your own documents. Files are hashed and matched locally; anchors carry the provenance line "matched by trial, not by export-carried reference." Multi-document traces are counted unanchorable with their own reason code.DESIGNED
--diffReport what changed against the previous entry in the local ledger: new agreeing repeat groups, groups that stopped agreeing, and the change in the headline share. A first run, and an entry another reader wrote, are answers and not errors: the report says there is nothing to compare, and exits 0.BUILT · NOT RELEASED
--sampleRun on a sample span export the package writes itself — a full report with nothing of yours involved, labeled as sample data on every screen. The release-1 sample is span-only; documents arrive with --docs in release 2. It takes the place of PATH, and it is not recorded in the ledger, so your first --diff still compares your own runs.BUILT · NOT RELEASED
--openOpen the HTML report when the run finishes.BUILT · NOT RELEASED
--from langfuseSkip the manual export: page the Langfuse observations API with LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY from your environment, incremental via --since. Your credential, your machine.PLANNED · AFTER FILE PATH

What the headline counts

The report opens with one sentence: “This report is the first step of an audit, not a dashboard. Where your platform already prints this number, start from that number; what this adds is the groups that disagree and the populations it could not read.”

  • Tool calls are grouped by (tool name, arguments digest) — exact repetition, not similarity.
  • A repeat group is a tier-0 cache candidate only when its results agree across every occurrence. Disagreeing groups are reported separately as nondeterminism findings. Beside the tier-0 cache candidate label the report prints the following, where N is the number of groups whose results disagree: “On the benchmark corpus, most of this figure was cross-run: the same call in a later run, which a session cache cannot see. There, a session cache scripted from this table captured 3–20% of context and returned a stale result on 17–45% of its hits. Caching these across runs is safe only where the result depends on the arguments alone. This export cannot show that; the N groups below show the opposite, and a file read never qualifies.”
  • The headline is the share of context — the bytes fed back to the model — that sits in agreeing repeat groups. It is said as a share of context, never as a share of calls: on the one benchmark corpus measured so far it reads “about 23 percent of context,” and most of that is cross-run — the same call repeated in a later run, which a session cache cannot see.
  • It is not token-weighted. No export measured so far pairs a character count with a token count of the same bytes, so a chars-per-token ratio is not computable; where the export carries token totals they are printed as totals, and the report says which unit it is in.
  • The process share and the navigation share are reported separately, where a reader declares the tools they count, and never folded into the headline. No release-1 reader declares them. The process share is refused, with the tool list it would have used printed in its place, when more than 20% of shell calls chain commands. The navigation share is a bracket under declared rules, not a point figure.
  • What the audit could not read is printed, not dropped, each with a count: records that did not parse, calls that carried no arguments or no result, and repeat groups without the start times to order them. Spans that cannot be anchored to a document arrive with --docs in release 2.

Files it reads and writes

FileDirectionWhat it is
<PATH>readsYour export. Never modified, never transmitted.
--docs directoryreadsYour documents. Hashed locally for anchoring.
./coldworks-audit.htmlwritesThe report. Self-contained, built to be forwarded — by you.
./coldworks-audit-sample.jsonlwritesOnly with --sample: the sample export, written so you can open the shape the report describes.
./.coldworks/audit/ledger.jsonlappendsDerived figures per run, append-only: the report's own figures and its repeat groups by tool name and arguments digest. It names no file and no path. What --diff reads. Delete it any time; you lose history, nothing else.
Network calls made by run: zero. The one planned exception is --from langfuse, which calls Langfuse — your trace store — with your key, from your machine. Nothing calls Coldworks.

Exit codes

CodeMeaning
0Audit completed — including an honest zero with a field census.
1The export could not be parsed at all; the error names the first unreadable record. A record that parses and is not in its reader's shape is not this: it is counted, with the location of the first, and the audit completes.
2Usage error; help printed.
PreviousConnect your traces