← cd ../blog

I use AI to write code every day. I trust it more than ever.

AI makes mistakes. A good engineering workflow catches them with lint rules, tests, independent agent reviews, and human judgment.

I use AI to write code every day. I trust it more than ever.

That does not mean I expect it to be right on the first try. I don’t expect that from a human developer either.

Bugs are part of software development. The relevant question is not whether AI produces bugs. It does. The question is what catches them, how early, and at what cost.

My answer is simple: use AI to review AI, then bring in a human when human judgment is actually needed.

The human stays in the loop. Just later.

Trust is not faith

Trust means predictable behavior inside explicit boundaries.

I trust a compiler because it rejects invalid code. I trust continuous integration because it runs the same checks on every change. Neither is infallible. Both are useful because their role is clear and their output is verifiable.

The same applies to AI.

I don’t trust a raw completion. I trust a workflow that gives the AI a clear target, rejects known bad patterns, tests the result, asks independent agents to attack it, and keeps a human responsible for the final decision.

This distinction matters. If your definition of trust requires flawless output, you cannot trust any developer, tool, dependency, or process involved in shipping software.

Bugs should be inputs to the workflow

The common response to AI mistakes is to add more human review. That works, but at a great cost. Human attention becomes the default error-detection system for problems that a machine could have rejected in seconds.

That does not scale.

A bug should feed the next iteration. The workflow should expose it, give the AI useful feedback, and let the AI fix it. A human should not need to inspect every intermediate attempt.

This is how I structure that workflow.

First, make the linter opinionated

Most teams use linting for formatting and a handful of generic mistakes. That leaves a lot of value on the table.

Your linter should encode the patterns you do not want in your codebase. Forbidden imports, invalid dependency directions, unsafe shortcuts, and project-specific rules should fail automatically when the linter can express them.

Don’t leave the same comment on three pull requests. Turn it into a rule.

This gives the AI a fast and deterministic rejection layer. It also gives every future contributor the same constraint. The rule does not get tired, forget the discussion, or decide that this exception probably looks fine.

The tradeoff is maintenance. A bad rule creates noise, and noise teaches people and agents to ignore the linter. Keep rules explicit. Remove the ones that no longer protect anything.

Linting cannot tell you whether the feature solves the right problem. That is not its job. It should reject what you already know you never want.

Second, use TDD as the control loop

Test-driven development is particularly effective with AI.

Start with a failing test. The test turns an ambiguous request into an executable target. The agent can implement a change, run the suite, read the failure, and iterate until the result is green.

That loop is fast. More importantly, it is inspectable.

Without tests, the agent decides when the work looks finished. With tests, the repository decides whether the expected behavior exists.

The distinction is huge.

Tests also protect the next change. A result that works today but silently breaks tomorrow is not a good result. Regression tests make previous decisions part of the codebase instead of leaving them in chat history or human memory.

Passing tests are not proof that the implementation is correct. Tests can be incomplete. They can encode the wrong expectation. They can miss an edge case.

Which is why the workflow needs another layer.

Third, use independent agents to review the change

For material changes, I use multiple subagents with different review scopes.

Do not ask three agents to “review the code.” That is one vague task repeated three times. You will get three variations of the same generic answer.

Give each reviewer a job. One can trace correctness against the request. Another can look for regressions and missing tests. Another can challenge the architecture, complexity, or security assumptions when those concerns are relevant.

The scopes should not overlap by accident. Independence is the point.

Why use a subagent instead of asking the implementation agent to review its own work?

Context.

The regular agent carries the entire implementation conversation. It knows the reasoning, the dead ends, the compromises, and what it intended the code to do. That context helps during implementation. During review, it becomes a source of bias. The agent can fill gaps with what it meant instead of judging what the diff actually does.

A subagent can start with a clean context. Give it the requirements, the diff, and one review scope. It does not inherit the implementation agent’s explanations or attachment to the solution. It has fewer reasons to defend the result and a better chance of reading the code as it is.

A clean context does not guarantee a good review. It removes one avoidable blind spot.

AI reviewing AI is not circular. Developers review code written by other developers every day. The useful properties are independent context and a different objective, not whether the reviewer is made of carbon.

When any reviewer finds a real issue, fix it, rerun the checks, and start another full review round. Keep going until every scoped reviewer comes back clean in the same round. A review is not a ceremony. It is another feedback loop with a clear stop condition.

Fourth, review the result manually

After linting, tests, and independent reviews pass, I review the result myself.

That does not mean reading every changed line. If I still need to manually inspect most of the code, the previous gates are not doing enough work.

I run the software. I test the important paths manually. For a user interface, I look at the rendered result and search for visual mistakes at the relevant screen sizes.

If performance matters enough for a feature or a specific code path, profiling belongs in the automated tests. Define the performance budget, run the profiler or benchmark as part of the checks, and fail when a regression crosses it. Performance should not depend on a human remembering to measure it.

This is the output users will experience. A clean diff is worthless if the interface is broken, the interaction feels wrong, or the application became slower.

I read the code sparingly. Line-by-line inspection makes sense when a mistake could be expensive or irreversible, such as a security boundary, money movement, or destructive data change. For ordinary changes, observable behavior is the better use of human attention.

Code review should follow risk. It should not be a ritual applied with the same intensity to every change.

Putting the human last does not remove the human from the loop. It stops using a human as a linter and uses them for what they do best: taste and judgment.

Why my trust keeps increasing

Better models help. Better workflows matter more.

A stronger model without checks can produce plausible mistakes faster. Add clear constraints, executable tests, and independent review, and the same speed closes the correction loop faster.

That is where trust comes from. I can delegate more because failures do not have a straight path into production. They hit several gates first.

This trust is calibrated. A small, reversible change with strong test coverage can run with more autonomy. A change with a high cost of failure needs tighter boundaries and more scrutiny. The workflow should match the risk.

Blind trust would be reckless. Permanent distrust is not much better. It leaves a powerful tool stuck at the autocomplete stage because the surrounding engineering process never evolved.

I trust AI more than ever because I know how to verify it independently.

Use lint rules to reject known bad patterns. Use TDD to give the agent a target and a correction loop. Use independent agents until every review comes back clean in the same round. Then test the result manually, automate the performance checks that matter, and inspect the code only when the risk demands it.

AI does not need to stop making mistakes before we can trust it.

The mistakes need somewhere reliable to go.