The audit · ReferenceDesign preview · coldworks-audit has not shipped
The audit CLI
One command with one job: read a trace export on your disk, write an audit report next to it. This page documents the designed interface — it is the contract the build is held to, published before the build so you can hold it to the contract too.
coldworks-audit run
Signature
coldworks-audit run <PATH> [--docs DIR] [--diff] [--sample] [--open] [--format NAME]
Audits the export at PATH — a file, or a directory whose .json and .jsonl exports are audited together, with span-level dedup across overlapping files. An observation that arrives twice under one run and span id is counted once, and the report says how many it dropped. Where an export carries no span ids, the report says “dedup unavailable: this export carries no span ids; overlapping files count twice.”
Argument
What it does
Status
PATH
The export file or directory. Format is detected from the first record: OTel GenAI spans, Langfuse observations (verified on the API shape), or the span-tree shape --sample writes. Override with --format otlp_genai|langfuse|span_tree.
BUILT · NOT RELEASED
--docs DIR
Anchor recovered judgments to lines in your own documents. Files are hashed and matched locally; anchors carry the provenance line "matched by trial, not by export-carried reference." Multi-document traces are counted unanchorable with their own reason code.
DESIGNED
--diff
Report what changed against the previous entry in the local ledger: new agreeing repeat groups, groups that stopped agreeing, and the change in the headline share. A first run, and an entry another reader wrote, are answers and not errors: the report says there is nothing to compare, and exits 0.
BUILT · NOT RELEASED
--sample
Run on a sample span export the package writes itself — a full report with nothing of yours involved, labeled as sample data on every screen. The release-1 sample is span-only; documents arrive with --docs in release 2. It takes the place of PATH, and it is not recorded in the ledger, so your first --diff still compares your own runs.
BUILT · NOT RELEASED
--open
Open the HTML report when the run finishes.
BUILT · NOT RELEASED
--from langfuse
Skip the manual export: page the Langfuse observations API with LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY from your environment, incremental via --since. Your credential, your machine.
PLANNED · AFTER FILE PATH
What the headline counts
The report opens with one sentence: “This report is the first step of an audit, not a dashboard. Where your platform already prints this number, start from that number; what this adds is the groups that disagree and the populations it could not read.”
Tool calls are grouped by (tool name, arguments digest) — exact repetition, not similarity.
A repeat group is a tier-0 cache candidate only when its results agree across every occurrence. Disagreeing groups are reported separately as nondeterminism findings. Beside the tier-0 cache candidate label the report prints the following, where N is the number of groups whose results disagree: “On the benchmark corpus, most of this figure was cross-run: the same call in a later run, which a session cache cannot see. There, a session cache scripted from this table captured 3–20% of context and returned a stale result on 17–45% of its hits. Caching these across runs is safe only where the result depends on the arguments alone. This export cannot show that; the N groups below show the opposite, and a file read never qualifies.”
The headline is the share of context — the bytes fed back to the model — that sits in agreeing repeat groups. It is said as a share of context, never as a share of calls: on the one benchmark corpus measured so far it reads “about 23 percent of context,” and most of that is cross-run — the same call repeated in a later run, which a session cache cannot see.
It is not token-weighted. No export measured so far pairs a character count with a token count of the same bytes, so a chars-per-token ratio is not computable; where the export carries token totals they are printed as totals, and the report says which unit it is in.
The process share and the navigation share are reported separately, where a reader declares the tools they count, and never folded into the headline. No release-1 reader declares them. The process share is refused, with the tool list it would have used printed in its place, when more than 20% of shell calls chain commands. The navigation share is a bracket under declared rules, not a point figure.
What the audit could not read is printed, not dropped, each with a count: records that did not parse, calls that carried no arguments or no result, and repeat groups without the start times to order them. Spans that cannot be anchored to a document arrive with --docs in release 2.
Files it reads and writes
File
Direction
What it is
<PATH>
reads
Your export. Never modified, never transmitted.
--docs directory
reads
Your documents. Hashed locally for anchoring.
./coldworks-audit.html
writes
The report. Self-contained, built to be forwarded — by you.
./coldworks-audit-sample.jsonl
writes
Only with --sample: the sample export, written so you can open the shape the report describes.
./.coldworks/audit/ledger.jsonl
appends
Derived figures per run, append-only: the report's own figures and its repeat groups by tool name and arguments digest. It names no file and no path. What --diff reads. Delete it any time; you lose history, nothing else.
Network calls made by run: zero. The one planned exception is --from langfuse, which calls Langfuse — your trace store — with your key, from your machine. Nothing calls Coldworks.
Exit codes
Code
Meaning
0
Audit completed — including an honest zero with a field census.
1
The export could not be parsed at all; the error names the first unreadable record. A record that parses and is not in its reader's shape is not this: it is counted, with the location of the first, and the audit completes.