Shepherd fixes the PR, Stamp approves it
09 Oct 2026Shepherd and Stamp take a pull request from opened to approved.
Shepherd runs in your coding agent, reviews the change, fixes what it can and repairs the build.
Stamp runs in CI and decides whether the result earns an approval.

I keep the two apart on purpose.
The agent that changes the code never gets to approve it, and you get the decisions that need a person.
PostHog’s list of AI coding mistakes argues that lint, types and tests matter more once an agent writes the code. Shepherd and Stamp put that into the pull request.
An agent that writes a fix and then approves it marks its own homework. I built Shepherd and Stamp so that authority sits in two places. Shepherd changes code and Stamp judges it, and each one’s authority ends where the other’s begins.
---
alt: "Shepherd runs inside a coding agent and loops over a pull request: swarm reviews it, triage fixes or defers each thread, and CI repair fixes the build. When a round finds nothing new, Stamp runs in CI with deterministic gates and a model review and approves or refuses. Both leave threads, commits, statuses and verdicts on GitHub."
---
flowchart LR
subgraph inner["Shepherd: in your coding agent"]
S["Swarm<br/>reviews the diff"] --> T["Triage<br/>fixes or defers threads"]
T --> C["CI repair"]
C --> S
end
inner --> ST["Stamp in CI<br/>gates + model review<br/>approves or refuses"]
ST --> GH[("GitHub<br/>threads, commits,<br/>statuses, verdicts")]
inner --> GH
How Shepherd works
Shepherd is a set of skills that runs inside a coding agent such as Claude Code or OpenCode. You install it with bunx @jagreehal/shepherd install and run /shepherd 42 on a PR. One iteration works like this:
- Swarm runs the repository’s own lint, typecheck and nearby tests first, plus any scanners you have installed, such as
gitleaks,semgrepandosv-scanner. A cheap model then reads the whole diff and hands the risky hunks to specialist lenses (correctness, security, simplicity, maintainability, slop, and any your team adds), which run in parallel. When the PR closes an issue, a spec lens checks the diff against what the issue asked for. A second model checks each high-severity finding against the code before it posts. - Triage works through the open review threads. It fixes the concrete ones, resolves the ones it can show are mistaken, and leaves the decisions that belong to the author. Threads another person has joined stay with that person.
- CI repair brings the branch up to date and fixes what the PR broke, with every test kept at its full strength.
- The loop repeats until a review round finds nothing new, then reads Stamp’s verdict.
Wrap it in /loop 5m /shepherd 42 and you can leave it running. Shepherd never replies to people and never approves.
How Stamp decides
Stamp runs in CI. bunx @jagreehal/stamp init writes its policy and workflows into your repository. Deterministic gates go first: a deny-list of sensitive paths, size limits, a credential scan, and who wrote the PR. A model review with read-only access to the checkout judges what the gates allow, and reads the issues the PR closes to check the diff does what they ask. The model can refuse a change the gates passed, and only the gates can make a change eligible. An approval is a real GitHub review, so it counts toward branch protection.
The model review runs through the AI SDK, so you choose the provider with one variable: STAMP_MODEL=bedrock:zai.glm-4.7-flash, opencode-go:kimi-k3, openrouter:moonshotai/kimi-k3, or a Claude model on Anthropic.
The model reads the checkout through three tools, read_file, grep and glob, confined to the repository, and ends the review by calling submit_verdict, whose input schema is the verdict. Every provider I tested supports tool calling, so one verdict format works everywhere. In my tests the models read the code before deciding and returned a valid verdict through the tool. A model that answers in prose gets one follow-up call that must be submit_verdict.
Stamp records each decision at the point it makes it, as I argued in You Can’t Support What You Didn’t Record, so every review leaves a run record. A Kimi K3 review of a one-line change made five tool calls, read 22,656 of its 33,580 input tokens from cache, and took 57 seconds. Limits from autotel’s guard bound each review: $2 of spend, three million tokens, a spin-loop rule, a tool-call cap, and a 15-minute timeout that cuts off a model call in flight. Set an OTLP endpoint and each review becomes a trace, with one span per model call and tool call.
Teams that prefer Claude Code or Codex can run those as the reviewer instead. Stamp runs them in a pinned container with a read-only copy of the checkout and only their own credential, on a private network whose one route out leads to their own model provider. The egress proxy opens a tunnel only when the TLS handshake inside it names that provider’s host, so a reviewer can’t reach another site that shares the provider’s CDN address.
See where every PR stands
Shepherd sets a shepherd commit status on the head at the end of every iteration:
pendingwhile work remains or a verdict hasn’t arrivedfailurewhen only the author can move the PR, with the reason in one line, such asstamp refused: ...or3 threads need yousuccessonce Stamp has approved and nothing is open
The status lives on GitHub, so it outlasts the session that set it. You can come back to a PR hours later and read what it needs from you without opening a transcript. A push Shepherd hasn’t seen arrives with no status at all, so a branch rule that requires the status keeps the PR from merging until Shepherd has looked at it.
Stamp handles description edits in a separate workflow, so the stamp / review check on each commit shows the verdict that counts.
Run the loop before you open a PR
You can run the same review on your machine before GitHub sees the change. /shepherd --local captures your unpushed commits, staged and unstaged edits and untracked files, then runs the same lenses, verification and fix rules as the PR loop. Fixes land in your working tree. Shepherd never commits, pushes, comments or sets a status in this mode, so you review the fixes and commit them yourself.
A clean report needs shepherd local finish to say complete. That check confirms every lens finished, every changed line range reached a reviewer, and the files didn’t change mid-review. It proves each line went in front of a reviewer, which falls short of proving the reviewer understood it. A complete review writes a receipt keyed to the exact content, and shepherd local status tells you whether the files on disk still match it. Put that in a pre-push hook if you want one.
Local mode ends at “local checks passed”. Merge readiness stays with the PR loop and Stamp. To see it work, bunx @jagreehal/shepherd local demo builds a repository with seven planted bugs, one for each way a change reaches the review, including a deleted check.
One rule pack for both tools
A service team cares about observability and API compatibility. A frontend team cares about accessibility and its design system. Shepherd lets your team add its own reviewers on top of the built-in lenses:
bunx @jagreehal/shepherd lens new observability --from observability --applies 'src/**'
That command copies Shepherd’s bundled observability pack into .shepherd/lenses/observability/SKILL.md and registers it in .shepherd/lenses.yml. Drop --from and you get a blank skill to fill in. The ## Review section lists rules with IDs, so a finding names the rule it applies: [observability/error-context] means a caught error lost its cause or the IDs you need to find the request. The pack’s seven rules check for the telemetry Instrument Before You Know the Question says you need before production breaks. They cover outbound calls with no span or log, unstructured logs, secrets in telemetry, silent data drops and unbounded metric labels. The ## Fix section tells triage to use the logger and tracer the repository already has, and to hand you any change that adds a telemetry dependency.
Shepherd reads the lens file from the default branch, so your team decides who reviews a PR and the PR itself has no say. You can try a lens on a real PR with /swarm --preview, which posts nothing.
Stamp reads the same file as a rule pack. Point .stamp/policy.yml at it and the reviewer applies those rules to matching files, then refuses, escalates or notes each broken one:
rules:
observability:
skill: .shepherd/lenses/observability
applies_to: ['src/**']
on_break: escalate
One file steers the agent that fixes the code and the gate that approves it, so the two can’t drift apart. A pack can only add checks. Stamp reads it from the default branch, and a PR that edits a configured pack goes to a person, because every later review trusts its text.
Catch an agent silencing a check
An agent chasing a green build can silence a rule instead of fixing the code. Swarm lists each suppression comment a PR adds, such as oxlint-disable, @ts-ignore or nosec, as a finding and reads the code it covers. Stamp flags the same comments for its reviewer. The reviewer refuses a suppression that hides a finding the change introduced and approves a narrow one that gives its reason.
A PR that loosens lint, test or scanner config gets the same treatment. A linter can’t see its own config change, so Stamp’s reviewer reads that diff and refuses a disabled rule, a widened ignore or a lowered coverage threshold. It approves tightening.
Spend money where judgement matters
.shepherd/lenses.yml also sets the models. A ladder lists models from cheapest to strongest, and models pins a lens or a runner to one model, from any provider your coding agent can reach:
ladder: [opencode-go/deepseek-v4-flash, opencode-go/glm-5.3, opencode-go/kimi-k3]
models:
security: opencode-go/qwen3.8-max
In one swarm review through OpenCode, the router on DeepSeek V4 Flash cost $0.013, two specialist lenses on GLM-5.3 cost about $0.16 each, and the orchestrating session on Kimi K3 cost $0.71. The orchestrator ran up most of that bill, so I would tune its model first.