The practical answer

Check the change against a written outcome, inspect the diff, run relevant checks, and try the real user flow. Keep a record of what you verified and a way to recover if deployment fails.

1. Write down what should happen

For this guide, vibe coding means using natural-language instructions to have an AI produce or change code, with much of the implementation delegated to the agent. You can work this way and still make a deliberate release decision.

Before reviewing, state the intended outcome in one or two sentences. “The sign-in form should show an inline error for a blank email and should not send a request” gives you something testable. “Improve authentication” leaves too much room for an agent to change unrelated behaviour.

Include one example that previously failed and one existing behaviour that should keep working. Ask the agent to repeat these acceptance criteria before editing. That makes it easier to catch a plausible implementation of the wrong task.

2. Read the actual changed files

Use your Git changes view or git diff. Look for changes beyond the requested scope, deleted checks, unexpected dependencies, generated files, and configuration that affects deployment. Ask the agent to explain unfamiliar changes in plain language.

A small task that rewrites many files deserves a second look. There may be a legitimate reason, but review it before accepting the result. Do not commit every modified file simply because the agent says it has finished.

Check that no credentials, private customer data or local environment files have entered the diff. If you find a real exposed secret, removing it from the next commit alone is not sufficient; follow the provider's process to revoke or rotate it.

3. Run checks that could catch the failure

Use the repository's documented commands. Type checking can find mismatched types; tests can check specific behaviour; a build can show whether the application packages successfully. None of these alone proves the entire user journey works.

For a bug fix, ask for a regression case that fails without the fix. A test that only repeats the implementation's assumptions adds little confidence. For a small copy or spacing change, inspecting the rendered page may be more useful than adding an elaborate new test suite.

Record the command, the result and any limitation. “The validator test passes; the browser flow is still unchecked” is an honest status. Do not let an agent turn an unrun command into a claim that everything passed.

4. Try the real flow, including a failure

Open the changed page or application and perform the task a user would perform. On a form, try a valid value, a blank value, a malformed value and a failed request. For a layout change, check a narrow screen and keyboard navigation.

Pay attention to loading and empty states. A feature can look finished with sample data and still fail for a new account or an unavailable service. Check that important buttons explain what happened, and that errors do not disappear behind a spinner.

If the change affects private information, verify access boundaries with separate test accounts and appropriate test data. A working button does not prove the server has enforced permission.

5. Give a second agent a bounded review

A second agent can help find gaps when it sees the original requirement and the actual change. Ask for concrete findings: a file, a consequence and a way to reproduce the problem. Keep uncertain suggestions separate from confirmed defects.

Prevent the reviewer from silently changing the implementation during its first pass. Review findings, choose which need action, and then ask for a fix. If several agents edit at once, use separate worktrees and integrate deliberately.

You can use Claude Code and Codex together for this cycle. Changing models is optional; a clearer task and a focused reproduction often matter more than the choice of reviewer.

6. Make the release decision explicit

Before deployment, know which commit is being released, which environment it will reach, and how to return to the previous working version. Database changes may need their own recovery plan; replacing application files may not undo a migration.

Use this short checklist in your task or pull request:

  • The change matches the written acceptance criteria.
  • The diff contains only intended files and no private data.
  • The relevant checks ran, and their results are recorded.
  • The real user flow and at least one failure case were tried.
  • Review findings are resolved or explicitly accepted.
  • The release target and recovery steps are understood.

Towfu can help you keep agent sessions and branches organised while you do this work. It does not replace the decision to merge or deploy. See the workspace for the product, and keep this checklist usable even when you work from ordinary terminals.