Working notes on reviewing what an agent did before it becomes your problem.
What changes when the assistant can act
The speed is real, and so is the review it quietly skips.
- Commands run: installs, deletions, rewritten history.
- Twenty minutes of agent work outpaces a day of reading.
- The summary is written by the same thing that wrote the code.
- Nothing slows down enough for you to notice you stopped checking.
One task, one diff
The shape of a good session is the same whichever tool you use.
Small steps are not about the agent's ability to hold them. They are about yours to check the result.
- State one task with the finish line named.
- Ask for the plan before the edits.
- Let it work in a bounded step, not on an open mandate.
- Read the diff — every changed line, before anything else.
- Ask for proof: the output, not “should work now”.
- Commit small and often; commits are the checkpoints.
Where it helps and where it drifts
The characteristic failure is not nonsense. It is confident, well-formatted work that is wrong in one place.
The review step therefore cannot be delegated back to the thing being reviewed.
- Good at boilerplate, mechanical refactors and test scaffolding.
- Good at reading unfamiliar code and explaining it back.
- Weak on facts it cannot check in the repository.
- Weak on trade-offs nobody told it about.
- Weak at knowing when to stop and ask.