uvx coldworks-audit run traces.jsonl locally.$ uvx coldworks-audit run traces.jsonl
11.84% of context sits in agreeing repeat groups (band 11.80% to 12.02%)
On the benchmark corpus, most of this figure was cross-run: the same call in a later run, which a session cache cannot see.
214 repeat groups holding 2,871 calls · 41,218 spans · 9,406 tool calls · your tool names, kept
→ tier-0 cache candidate: 131 agreeing groups holding 1,704 calls
→ 83 groups holding 1,167 calls returned different results nondeterminism findings
not counted as findings, not hidden:
322 calls carried no result · 118 carried no arguments
report: ./coldworks-audit.htmlCandidates, never savings — a repeated judgment is a candidate for compilation, not a booked dollar. The report prints what it could not read in the same breath as what it could.
Langfuse observations carry name, input, output, and token usage — everything the audit needs. Three rungs, all running with your credentials on your compute:
| Rung | How it works | Status |
|---|---|---|
| File | Traces → Export in the Langfuse UI, then coldworks-audit run on the downloaded file. | BUILT · API SHAPE ONLY |
| Standing file | Langfuse's scheduled export writes fresh traces to your S3 / GCS bucket on its own timer; your cron line re-audits whatever lands. Two schedulers, both yours — there is no Coldworks daemon. | BUILT · NOT RELEASED |
| Direct pull | coldworks-audit run --from langfuse reads LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY from your environment and pages the observations API itself — incremental via --since. | PLANNED · AFTER FILE PATH |
No LangSmith reader ships in release 1, and --format does not list one: for most plans no export file exists until a script writes one, and no LangSmith export has been read. The reader follows the first LangSmith export in hand.
Point your collector's file exporter at a directory and audit what lands there.
Whatever the source, the headline is not token-weighted. No export measured so far pairs a character count with a token count of the same bytes, so the headline is a share of context in characters. Where an export carries token totals, the report prints them as totals, in their unit.
| Thing | Where it goes | Status |
|---|---|---|
| Your traces | Read from your disk, never transmitted. | NEVER LEAVES |
Your documents (--docs) | Hashed and matched locally, never transmitted. | NEVER LEAVES |
| The audit report | An HTML file on your disk. Forward it yourself, or don't. | STAYS LOCAL |
| Derived figures | A future opt-in --push could publish headline figures to a hosted registry page. It does not exist, and ships only under an explicitly signed exception. | DOES NOT EXIST |
This is doctrine, not a feature flag: Coldworks compiles judgments into code you own and you run. Auditing starts on the same side of the boundary the guards live on — yours.
When PATH is a directory under your working directory, the report footer prints one line for your scheduler — cron, GitHub Actions, or GitLab CI — naming the directory relative to it. The report never prints a full path, so a directory somewhere else gets no line: one that named it ./exports would re-audit the wrong place. A shell command in your infrastructure; never an app you install or a marketplace listing.
# re-audit whatever lands in ./exports, every night at 02:10
10 2 * * * uvx coldworks-audit run ./exports --diff >> audit.logEach run appends to a local append-only ledger, and --diff reports what changed against the previous entry: new agreeing repeat groups, groups that stopped agreeing, and the change in the headline share. A count of new agreeing groups that keeps rising is your agent re-buying the same decision with fresh tokens.
--from langfuse pull is the planned successor. There is no Coldworks daemon, resident process, or phone-home in any rung.