Use AI to need less AI.
Coldworks reviews your pull requests, remembers what your team decided, and turns the judgments it keeps repeating into checks you own. One product, free to start. The reviewer installs on its own today; the rest carries its state on this page.
One workspace.
Your organization on GitHub is the workspace. Reviews, Memory, and Guards each work from the record of what your team decides.
Overview 11 repositories, 3 agents connected
| Pull request | Verdict | Settled by | Cited | Heat |
|---|---|---|---|---|
| Retry policy for the export worker#1842 · agent | Cleared | Guard 0412 | Retries are bounded, never infinite | |
| Move billing webhooks to the new queue#1839 · alex | Needs a reader | Model | Postgres job tables, not Pub/Sub | |
| Bump httpx to 0.28#1837 · dependabot | Cleared | Guard 0107 | — | |
| Rename tenant columns in reports#1835 · agent | Cleared | Unproven, stays with the model | Tenant id is never nullable |
Illustrative. Every figure is invented.
Works with what you already have. The audit reads the traces you already export, Memory answers over MCP, and guards load into the CI and tools you already run. Your tools stay where they are; what changes is how much the model is asked.
One record of what your team decides, read at three moments.
Before an agent starts, at the pull request, and every week after. The same decision answers all three, so a settled question stays settled and the model is asked less each time. Coldworks is not a check your agents call. What is decided is read before they start, what is compiled runs instead of them, and what is left goes to the model, with a receipt.
Agents ask over MCP whether a question is already decided and get a cited answer, or "no recorded ruling," which means unknown, not approved.
Every pull request gets a verdict: cleared, or needs a reader. Coldworks reads the record before it reads the diff, routes attention, and never blocks. It works alone, too.
Judgments that come back the same every time become guarded, tested checks. A maintainer signs each one, it runs first in your CI, and Coldworks stops asking the model about it.
Your team already decided this. Your agents should know.
Connect a repository and Coldworks derives the record from what actually shipped: merged pull requests, and the architecture documents at the head of your branch. Nothing is typed in by hand. Commitment is the decision.
- Search it, or let agents read it over MCP before they start work.
- A settled question stays settled. Recency supersedes, and outcomes grade each decision over time.
- Agents stop re-arguing closed questions and stop repeating finished work.
Most pull requests don't need a human. Coldworks works out which ones do.
Every pull request gets a verdict: cleared, or needs a reader. Coldworks routes attention. It never blocks a merge, never writes code, and publishes its own miss rate.
- Runs on your own model keys, or on the free deterministic tier with no model at all.
- Reads the record first. A change that contradicts a settled decision is routed to the person who owns it.
- Every verdict is recorded with its evidence, which is what the guards learn from.
Judgments that repeat become checks you own.
Coldworks finds the verdicts that come back the same every time, compiles each one into a guarded, tested check, and stops until a maintainer signs it. Every guard subtracts something from what the model is asked.
- Signed guards run first, in your CI. The model is asked only what the guards can't answer.
- A guard that steps outside its proof hands the pull request back to the model. Nothing guesses.
- Everything compiled is yours: plain code, tests, fixtures, receipts. It runs without us.
Running agents you didn't write? Start with the audit.
The same machinery that compiles review verdicts reads the traces of any agent you run. Export them, run one command on your own machine, and see which judgments repeat. Nothing is uploaded, and there is no account.
Findings, never savings.
The report groups tool calls that repeat with the same arguments and counts only the ones whose results agree. It prints the populations it cannot anchor in the same breath, and a scheduler line so the next export runs itself.
- Reads Langfuse, Claude Code, and SWE-agent exports today.
- One forwardable report, a local ledger, and a diff against your last run.
- When you want the repeats compiled and guarded in your own orchestrator, that is the engagement, and your merge is the signature.
How it compounds.
The order is the order the record is built in. Each step's output is the next one's input, and the loop closes back on the first.
What merges becomes the record of what was decided.
Agents read it over MCP before they start, so settled questions stay settled.
Every pull request gets a verdict, with its evidence.
Cleared or needs a reader. Each verdict is a recorded judgment, not a comment that scrolls away.
Verdicts that repeat are compiled and proven.
Each candidate is replayed against sealed history. A maintainer signs it, or it stays unproven.
Guards run first. The model is asked less.
Reviews and your agents skip what a guard already settled. Next week's verdicts start from here.
T is the share of decisions a model still makes. The rest is plain code that calls no model and holds no tools, and the registry shows what the model still decides.
The pull request proceeds either way. Coldworks only decides who has to look.
The moment a reviewer authors, it owns the authorship. Coldworks doesn't.
Frequency alone never turns a judgment into code. A person signs, with a rollback path.
When the evidence can't separate the cases, Coldworks says so and keeps that work with the model.
Plain code, schemas, tests, fixtures, receipts. It runs without us.
Free to start. Flat when you pay.
Per installation, self-serve. No figures are published while nothing is public; the model is.
- Deterministic verdict and check run, no model reads
- Memory on public repositories, over MCP and search
- Never blocks, never writes code
- Published miss rate
- Memory on private repositories
- Pooled model reads, or bring your own key
- Guards: compile, sign, and run in your CI
- Retained and exportable history, receipts, calibration
- Nothing is uploaded; no account
- One forwardable report and a local ledger
- Diff against your last run
- The engagement, when you want the repeats compiled
Paid plans are flat, not metered. You pay the same whether the model ran or a guard did. Our cost falls as guards install. That margin is the only way we make money, so we only profit by making your workflow need less AI.
Read the docs
docs/decisions.md, entry 2026-08-12.