Paladin

AI-Assisted
Architecting

Durable investigation records · Focused utility · Private

Investigation that builds on everything already established instead of restarting from zero.

Paladin deliberately does the bookkeeping, not the thinking. Humans and models investigate; Paladin keeps claims, evidence, disagreement, coverage, and current standing in the project's own files.

Where this stands

In use on our own codebases. The likeliest future is inside Arboretum, where findings and coverage join the document surface and can be read, argued with, and turned into documentation.

01 · What it is

A durable ledger in the project's own files.

Paladin is a command-line utility that gives model-driven investigation a durable ledger in the project's own files. It deliberately does the bookkeeping, not the thinking — it never runs a model itself — and its effect is cumulative: investigation that builds on everything already established instead of restarting from zero, so ten passes reach conclusions no single pass could.

A CLI, and nothing you have to adopt

It is a command-line tool with a wide set of options for examining a project from different angles — structure, failure surfaces, trust boundaries, documentation against reality — and it works with whichever agent you already use. Claude Code, OpenAI Codex, Cursor, Gemini: Paladin composes the brief, the agent investigates, Paladin keeps the record. Nothing about the record depends on which one ran.

Everything it writes stays in the project as plain, non-proprietary files: readable by a person, readable by an agent, diffable in review, and carried by the repository like any other source. There is no database to host, no account to keep, and nothing to export if you stop using it — the ledger is already in a format anything can read.

02 · The problem

Models made investigation cheap. Nothing made its results durable.

Models made investigation cheap; nothing made its results durable. Observations live in chat logs and scattered notes. The same issue is rediscovered under three names; a lead examined and rejected returns next month as new; a source change quietly invalidates last month's conclusion; and the most useful question — what has never been examined at all? — has no answer. Attention pools on busy files and loud findings, while the quiet corner nobody revisits is where surprises live.

Without a ledger
With accumulation
Separate investigations repeatedly buy similar knowledge. A durable record lets later work begin at the frontier.

[Insert — figure] Two timelines, no numbers: investigations from zero vs. investigations that accumulate. [Insert — screenshot] A findings file in the repository — ownership made visible.

How one investigation becomes durable attention.

The chart follows one investigation from human focus through selected material, AI examination, durable findings, and attention written back for the next pass.

The source material and the semantic map remain distinct. Selection is a starting point, not a boundary on where the investigation may follow evidence.
Read the chart in more detail

Faded rows sit outside the current semantic focus. Dashed rows are eligible but were not selected for this turn; solid rows form the starting set. The investigation may add findings, and every pass writes its attention back so later work can deliberately revisit neglected or changed areas.

Findings can collect automated opinions and selective human input. Synthesis can create higher-level findings without erasing the linked records underneath them.

03 · What it concretely does

The model investigates. Paladin keeps the receipt.

Composes briefs that steer attention honestly. You describe what may be inspected, the concern that matters now, and limits on scope or privacy. Paladin mixes the obviously relevant starting points with neglected areas and earlier conclusions due for another look because the code under them changed. The model stays free to follow evidence — the brief is a nudge, not a script.

Keeps findings as durable records. Each finding has stable identity: the claim, cited evidence, author, judgment history, current standing. Rejected ideas stay on file as rejected — negative knowledge, kept deliberately, so a later model cannot resell the same false lead as new. Records live in project files and move with the repository.

Issues receipts, not vibes. Every run accounts for itself: what was supplied, what was actually cited, what received attention, what changed since last time, which decisions genuinely need a person. Over runs this becomes coverage you can question — a map of examined and unexamined ground.

Holds review state across interruptions. Findings review happens in batches with explicit state, so a batch can be put down mid-way and resumed exactly where it stopped — and one considered response to an organized batch costs less, in attention and tokens, than twenty scattered replies.

Inside Arboretum

The ledger stops being a file and becomes a surface.

A command-line ledger is honest but quiet. Inside Arboretum the same records acquire a place to be read, argued with and decided on: the investigation is launched from the document surface, the findings arrive as reviewable items beside the material they cite, and coverage becomes something you can look at rather than something you have to query.

This is what the two tools do for each other. Arboretum supplies the surface, the folder scoping and the choice of assistant; Paladin supplies the discipline — bounded scope, cited evidence, attributed decisions, coverage that reports attention rather than correctness. Neither had to be redesigned for the other.

One bounded starting point, not an offer to read everything. The counters are still empty.

Consent is explicit, and so is cost. Before anything runs, the dialog names the assistant, states that it receives the project as its working directory and may use its provider quota, and requires guarded local execution to be allowed for that browser tab. The person choosing can see what they are agreeing to.

Nine observations, 41 of 351 sources examined — and the 310 never examined stated as plainly as the rest.

Coverage is the part people miss. The bars do not claim the examined areas are correct; they record where attention has actually been. The useful question — what has never been looked at? — finally has an answer, and the next investigation is chosen against it rather than against whatever is loudest.

Every decision is attributed and written to the notebook.

A decision is a record, not a dismissal. Accepted, rejected, deferred, duplicate, superseded — each is written with the name of whoever decided and their reasoning. Rejected findings stay on file as rejected, which is exactly what stops a later model reselling the same false lead as new.

One investigation, end to end

Seven screens of the examination itself, then the Arboretum setup that precedes it. Use the arrows, or ← and → on the keyboard.

04 · Where the value comes from

Turn repeated investigation from spend into an asset.

Re-covered ground is the waste: every re-derived conclusion, re-rejected lead, and re-explained context is model budget and reviewer attention spent buying something already owned. The ledger converts that spend into an asset that appreciates — later investigations inherit earlier ones, contradictions surface instead of coexisting, and human review lands on the frontier instead of the familiar.

Paladin's findings fit naturally inside Arboretum's document and annotation surface. Consolidating the record there is more likely than building a large separate product around the bookkeeping.

Connected systemArboretum

Move findings into the document surface