# Shepherd fixes the PR, Stamp approves it

By Jag Reehal, 2026-10-09. Canonical: https://arrangeactassert.com/posts/shepherd-fixes-the-pr-stamp-approves-it/

<!-- Excerpt Start -->

Shepherd and Stamp take a pull request from opened to approved.

Shepherd runs in your coding agent, reviews the change, fixes what it can and repairs the build.

Stamp runs in CI and decides whether the result earns an approval.

![A Lego shepherd with a wrench fights red bugs on a PR platform beside a board of passing unit tests, lint and security checks, then sends a green PR brick across a bridge to a CI platform, where a second figure holds a red APPROVED stamp](/_astro/shepherd-fixes-the-pr-stamp-approves-it.B-m2Ejfz_EAMdi.webp)

I keep the two apart on purpose.

The agent that changes the code never gets to approve it, and you get the decisions that need a person.

<!-- Excerpt End -->

PostHog's [list of AI coding mistakes](https://posthog.com/newsletter/ai-coding-mistakes) argues that lint, types and tests matter more once an agent writes the code. Shepherd and Stamp put that into the pull request.

An agent that writes a fix and then approves it marks its own homework. I built [Shepherd](https://github.com/jagreehal/shepherd) and [Stamp](https://github.com/jagreehal/stamp) so that authority sits in two places. Shepherd changes code and Stamp judges it, and each one's authority ends where the other's begins.

<figure>

```mermaid
---
alt: "Shepherd runs inside a coding agent and loops over a pull request: swarm reviews it, triage fixes or defers each thread, and CI repair fixes the build. When a round finds nothing new, Stamp runs in CI with deterministic gates and a model review and approves or refuses. Both leave threads, commits, statuses and verdicts on GitHub."
---
flowchart LR
  subgraph inner["Shepherd: in your coding agent"]
    S["Swarm<br/>reviews the diff"] --> T["Triage<br/>fixes or defers threads"]
    T --> C["CI repair"]
    C --> S
  end
  inner --> ST["Stamp in CI<br/>gates + model review<br/>approves or refuses"]
  ST --> GH[("GitHub<br/>threads, commits,<br/>statuses, verdicts")]
  inner --> GH
```

<figcaption>Figure 1: Shepherd gets the PR ready, and Stamp decides whether it gets approved.</figcaption>
</figure>

## How Shepherd works

Shepherd is a set of skills that runs inside a coding agent such as Claude Code or OpenCode. You install it with `bunx @jagreehal/shepherd install` and run `/shepherd 42` on a PR. One iteration works like this:

1. **Swarm** runs the repository's own lint, typecheck and nearby tests first, plus any scanners you have installed, such as `gitleaks`, `semgrep` and `osv-scanner`. A cheap model then reads the whole diff and hands the risky hunks to specialist lenses (correctness, security, simplicity, maintainability, slop, and any your team adds), which run in parallel. When the PR closes an issue, a spec lens checks the diff against what the issue asked for. A second model checks each high-severity finding against the code before it posts.
2. **Triage** works through the open review threads. It fixes the concrete ones, resolves the ones it can show are mistaken, and leaves the decisions that belong to the author. Threads another person has joined stay with that person.
3. **CI repair** brings the branch up to date and fixes what the PR broke, with every test kept at its full strength.
4. The loop repeats until a review round finds nothing new, then reads Stamp's verdict.

Wrap it in `/loop 5m /shepherd 42` and you can leave it running. Shepherd never replies to people and never approves.

## How Stamp decides

Stamp runs in CI. `bunx @jagreehal/stamp init` writes its policy and workflows into your repository. Deterministic gates go first: a deny-list of sensitive paths, size limits, a credential scan, and who wrote the PR. A model review with read-only access to the checkout judges what the gates allow, and reads the issues the PR closes to check the diff does what they ask. The model can refuse a change the gates passed, and only the gates can make a change eligible. An approval is a real GitHub review, so it counts toward branch protection.

The model review runs through the [AI SDK](https://ai-sdk.dev/), so you choose the provider with one variable: `STAMP_MODEL=bedrock:zai.glm-4.7-flash`, `opencode-go:kimi-k3`, `openrouter:moonshotai/kimi-k3`, or a Claude model on Anthropic.

The model reads the checkout through three tools, `read_file`, `grep` and `glob`, confined to the repository, and ends the review by calling `submit_verdict`, whose input schema is the verdict. Every provider I tested supports tool calling, so one verdict format works everywhere. In my tests the models read the code before deciding and returned a valid verdict through the tool. A model that answers in prose gets one follow-up call that must be `submit_verdict`.

Stamp records each decision at the point it makes it, as I argued in [You Can't Support What You Didn't Record](/posts/you-cant-support-what-you-didnt-record/), so every review leaves a run record. A Kimi K3 review of a one-line change made five tool calls, read 22,656 of its 33,580 input tokens from cache, and took 57 seconds. Limits from [autotel](/posts/logging-sucks/)'s guard bound each review: $2 of spend, three million tokens, a spin-loop rule, a tool-call cap, and a 15-minute timeout that cuts off a model call in flight. Set an OTLP endpoint and each review becomes a trace, with one span per model call and tool call.

Teams that prefer Claude Code or Codex can run those as the reviewer instead. Stamp runs them in a pinned container with a read-only copy of the checkout and only their own credential, on a private network whose one route out leads to their own model provider. The egress proxy opens a tunnel only when the TLS handshake inside it names that provider's host, so a reviewer can't reach another site that shares the provider's CDN address.

## See where every PR stands

Shepherd sets a `shepherd` commit status on the head at the end of every iteration:

- `pending` while work remains or a verdict hasn't arrived
- `failure` when only the author can move the PR, with the reason in one line, such as `stamp refused: ...` or `3 threads need you`
- `success` once Stamp has approved and nothing is open

The status lives on GitHub, so it outlasts the session that set it. You can come back to a PR hours later and read what it needs from you without opening a transcript. A push Shepherd hasn't seen arrives with no status at all, so a branch rule that requires the status keeps the PR from merging until Shepherd has looked at it.

Stamp handles description edits in a separate workflow, so the `stamp / review` check on each commit shows the verdict that counts.

## Run the loop before you open a PR

You can run the same review on your machine before GitHub sees the change. `/shepherd --local` captures your unpushed commits, staged and unstaged edits and untracked files, then runs the same lenses, verification and fix rules as the PR loop. Fixes land in your working tree. Shepherd never commits, pushes, comments or sets a status in this mode, so you review the fixes and commit them yourself.

A clean report needs `shepherd local finish` to say `complete`. That check confirms every lens finished, every changed line range reached a reviewer, and the files didn't change mid-review. It proves each line went in front of a reviewer, which falls short of proving the reviewer understood it. A complete review writes a receipt keyed to the exact content, and `shepherd local status` tells you whether the files on disk still match it. Put that in a pre-push hook if you want one.

Local mode ends at "local checks passed". Merge readiness stays with the PR loop and Stamp. To see it work, `bunx @jagreehal/shepherd local demo` builds a repository with seven planted bugs, one for each way a change reaches the review, including a deleted check.

## One rule pack for both tools

A service team cares about observability and API compatibility. A frontend team cares about accessibility and its design system. Shepherd lets your team add its own reviewers on top of the built-in lenses:

```bash
bunx @jagreehal/shepherd lens new observability --from observability --applies 'src/**'
```

That command copies Shepherd's bundled observability pack into `.shepherd/lenses/observability/SKILL.md` and registers it in `.shepherd/lenses.yml`. Drop `--from` and you get a blank skill to fill in. The `## Review` section lists rules with IDs, so a finding names the rule it applies: `[observability/error-context]` means a caught error lost its cause or the IDs you need to find the request. The pack's seven rules check for the telemetry [Instrument Before You Know the Question](/posts/instrument-before-you-know-the-question/) says you need before production breaks. They cover outbound calls with no span or log, unstructured logs, secrets in telemetry, silent data drops and unbounded metric labels. The `## Fix` section tells triage to use the logger and tracer the repository already has, and to hand you any change that adds a telemetry dependency.

Shepherd reads the lens file from the default branch, so your team decides who reviews a PR and the PR itself has no say. You can try a lens on a real PR with `/swarm --preview`, which posts nothing.

Stamp reads the same file as a rule pack. Point `.stamp/policy.yml` at it and the reviewer applies those rules to matching files, then refuses, escalates or notes each broken one:

```yaml
rules:
  observability:
    skill: .shepherd/lenses/observability
    applies_to: ['src/**']
    on_break: escalate
```

One file steers the agent that fixes the code and the gate that approves it, so the two can't drift apart. A pack can only add checks. Stamp reads it from the default branch, and a PR that edits a configured pack goes to a person, because every later review trusts its text.

## Catch an agent silencing a check

An agent chasing a green build can silence a rule instead of fixing the code. Swarm lists each suppression comment a PR adds, such as `oxlint-disable`, `@ts-ignore` or `nosec`, as a finding and reads the code it covers. Stamp flags the same comments for its reviewer. The reviewer refuses a suppression that hides a finding the change introduced and approves a narrow one that gives its reason.

A PR that loosens lint, test or scanner config gets the same treatment. A linter can't see its own config change, so Stamp's reviewer reads that diff and refuses a disabled rule, a widened ignore or a lowered coverage threshold. It approves tightening.

## Spend money where judgement matters

`.shepherd/lenses.yml` also sets the models. A `ladder` lists models from cheapest to strongest, and `models` pins a lens or a runner to one model, from any provider your coding agent can reach:

```yaml
ladder: [opencode-go/deepseek-v4-flash, opencode-go/glm-5.3, opencode-go/kimi-k3]
models:
  security: opencode-go/qwen3.8-max
```

In one swarm review through OpenCode, the router on DeepSeek V4 Flash cost $0.013, two specialist lenses on GLM-5.3 cost about $0.16 each, and the orchestrating session on Kimi K3 cost $0.71. The orchestrator ran up most of that bill, so I would tune its model first.