Would I Merge Code an Agent Wrote Alone
Claude finishes editing and says “Done.” I glance at it, looks fine, merge it. Then Monday comes and production is down.
Without a runnable check, the agent’s “done” and my “done” are two different things. The most expensive bugs aren’t crashes — they’re silent wrong answers.
The simplest check: one line in my prompt
I add this at the end of my prompt:
Run the tests after you finish. Keep fixing until they all pass.Claude will run npm test on its own, check output, fix what’s broken, and rerun. No input needed from me.
But there’s a catch — it only validates against its own understanding of “passing.” If my test coverage is thin, or “done” means more than green tests, I need more layers.
Layer up
Layer one: in-prompt check
The one-liner above. Good enough for small daily changes.
Layer two: define done
I tell Claude what “done” looks like:
After each change, run npm run build && npm test and show me the output.At least I see what it ran and what came out. Not just taking its word for it.
Layer three: stop the covering up
I add this crucial line:
Fix the root cause, don't hide the error.This stops the agent from wrapping things in empty try-catch or returning fake fallback data to make the tests pass.
Layer four: /goal conditions
I set the verification criteria as a session-level goal. Claude re-checks it after every turn:
/goal all tests pass, build output under 200KBNo need to say “run the tests again” after every change. Claude checks automatically and keeps fixing until the goal is met.
Layer five: Stop Hook
The hardest gate — I write a hook script that runs automatically when Claude stops, and blocks the turn if checks fail:
{
"hooks": {
"Stop": [
{
"hooks": [
{ "type": "command", "command": "npm test && npm run build" }
]
}
]
}
}Hooks are deterministic. Unlike CLAUDE.md instructions which are advisory, hooks guarantee the action happens. If a rule must hold every time, I make it a hook.
Layer six: verification subagent
I have a fresh model review the diff in an isolated context. The agent that did the work shouldn’t be the one grading it. Claude Code ships with /code-review that does exactly this — reviews the current diff in a separate subagent and reports gaps.
I can also write my own: “Use a subagent to review this diff against PLAN.md. Check every requirement is implemented, edge cases have tests, nothing outside scope changed.”
Common failure patterns
Kitchen-sink sessions
Running unrelated tasks in one session pollutes later reasoning. My habit: /clear between tasks.
Fixing in circles
Every failed attempt stays in context. New fixes pile on top. Eventually Claude doesn’t know what state it’s in.
I save progress to a file, /clear, and restart with a sharper prompt.
Trust without verification
Code that looks right can fail at boundaries. Review habit: I ask myself “what if this data is empty, negative, or malformed?”
Infinite exploration
Unscoped research reads hundreds of files, fills the context, and produces nothing useful. I set boundaries for research tasks, or hand them to a subagent.