COLDWORKS/ docs / audit / connect
ReviewsMemoryGuardsThe auditScoreboardDocsGitHubAbout
Menu
ReviewsMemoryGuardsThe auditScoreboardDocsGitHubAbout
Sign in
Overview
Reviews docs
Getting started
  • Quickstart
Concepts
  • Risk routing
  • Defect labels
  • The cleared band
  • What Coldworks gets wrong
Reference
  • CLI · doug-backtest
  • The report
Coming up
  • MCP · Pattern Gardenplanned
  • REST APIpreview
Meta
  • Changelog
The audit docs · preview
Getting started
  • Quickstart
  • Connect your traces
Reference
  • The audit CLI
  • The audit reportsoon
  • Guards & subtractionsoon
View as llms.txt
The audit · Getting startedDesign preview · coldworks-audit has not shipped

Connect your traces

You already have the traces — your trace store has been writing your agent's judgments down the whole time. Coldworks reads them where they live: on your machine. A batch trace export is a file transfer, not an integration.

The three steps

  1. Export your traces from your store (about 30 seconds — every store already has this).
  2. Run uvx coldworks-audit run traces.jsonl locally.
  3. Read the report it writes next to the file.
Illustrative — no real store behind this page
$ uvx coldworks-audit run traces.jsonl

11.84% of context sits in agreeing repeat groups (band 11.80% to 12.02%)
On the benchmark corpus, most of this figure was cross-run: the same call in a later run, which a session cache cannot see.

  214 repeat groups holding 2,871 calls · 41,218 spans · 9,406 tool calls · your tool names, kept
  → tier-0 cache candidate: 131 agreeing groups holding 1,704 calls
  → 83 groups holding 1,167 calls returned different results  nondeterminism findings

not counted as findings, not hidden:
  322 calls carried no result · 118 carried no arguments

report: ./coldworks-audit.html

Candidates, never savings — a repeated judgment is a candidate for compilation, not a booked dollar. The report prints what it could not read in the same breath as what it could.

From Langfuse

Langfuse observations carry name, input, output, and token usage — everything the audit needs. Three rungs, all running with your credentials on your compute:

RungHow it worksStatus
FileTraces → Export in the Langfuse UI, then coldworks-audit run on the downloaded file.BUILT · API SHAPE ONLY
Standing fileLangfuse's scheduled export writes fresh traces to your S3 / GCS bucket on its own timer; your cron line re-audits whatever lands. Two schedulers, both yours — there is no Coldworks daemon.BUILT · NOT RELEASED
Direct pullcoldworks-audit run --from langfuse reads LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY from your environment and pages the observations API itself — incremental via --since.PLANNED · AFTER FILE PATH

From LangSmith

No LangSmith reader ships in release 1, and --format does not list one: for most plans no export file exists until a script writes one, and no LangSmith export has been read. The reader follows the first LangSmith export in hand.

From OpenTelemetry

Point your collector's file exporter at a directory and audit what lands there.

Whatever the source, the headline is not token-weighted. No export measured so far pairs a character count with a token count of the same bytes, so the headline is a share of context in characters. Where an export carries token totals, the report prints them as totals, in their unit.

What leaves your machine: nothing

ThingWhere it goesStatus
Your tracesRead from your disk, never transmitted.NEVER LEAVES
Your documents (--docs)Hashed and matched locally, never transmitted.NEVER LEAVES
The audit reportAn HTML file on your disk. Forward it yourself, or don't.STAYS LOCAL
Derived figuresA future opt-in --push could publish headline figures to a hosted registry page. It does not exist, and ships only under an explicitly signed exception.DOES NOT EXIST

This is doctrine, not a feature flag: Coldworks compiles judgments into code you own and you run. Auditing starts on the same side of the boundary the guards live on — yours.

Day 2: make it standing

When PATH is a directory under your working directory, the report footer prints one line for your scheduler — cron, GitHub Actions, or GitLab CI — naming the directory relative to it. The report never prints a full path, so a directory somewhere else gets no line: one that named it ./exports would re-audit the wrong place. A shell command in your infrastructure; never an app you install or a marketplace listing.

Report footer · illustrative
# re-audit whatever lands in ./exports, every night at 02:10
10 2 * * *  uvx coldworks-audit run ./exports --diff >> audit.log

Each run appends to a local append-only ledger, and --diff reports what changed against the previous entry: new agreeing repeat groups, groups that stopped agreeing, and the change in the headline share. A count of new agreeing groups that keeps rising is your agent re-buying the same decision with fresh tokens.

Honest limit: the cron line monitors only if fresh exports keep landing in that directory. Langfuse's scheduled bucket export, on the tiers that have it or self-hosted, supplies the files; no batch export has been read yet. The direct --from langfuse pull is the planned successor. There is no Coldworks daemon, resident process, or phone-home in any rung.
PreviousQuickstartNextThe audit CLI