Who reviews the reviewers?
This is something I've been thinking about. When I try to figure it out based on my experience, I reach the understanding that the authority of reviewers is a mix of overall experience and, most importantly, knowledge of the project's codebase and architecture.
So... what did we use to get from a PR?
- Not introducing a new bug
- Keeping in line with team/project standards
- Knowledge transfer, either project specific or more general
- Shared accountability
Agentic-age friction
With the ever-growing adoption of coding agents in the workforce, we crashed head-on into an issue: the agents have indeed multiplied our output by a lot. A PR that used to take a few days might now become an open PR within a few hours.
PR reviews have become a bottleneck that is hard to ignore. The idea that every PR requires another IC to be allowed into production does not apply to every project. There are systems where, either due to compliance or to the damage a change could cause, human review is required, but these tend to be the minority.
So what options do we have when we need to speed up development?
Self-review
The idea here is simple: if we agree that coding is purely an agent task (big if), then an IC who hands off a task to an agent is akin to offloading it to a junior team member. The implementation gets done, and that same IC is the one who reviews it. This is the first idea that comes to mind, and it's a valid workflow, but it brings up a few issues we need to address.
One of the most important issues with self-review is that it's susceptible to tunnel vision: you want to do X, the change Claude came up with looks good, and it does what you were expecting. Sounds great. But what if this change has a deeper impact on something else that, due to your mental model, you missed in the idea/plan, and therefore also missed in the review?
Another point to bring up: if deep understanding of the systems and architecture is required to catch issues beyond the code changes Claude has proposed, which might be the case in many companies, what does this mean for junior roles? How can they learn and become the seniors we want?
Before, you'd give them small tasks with limited scope, but agents finish these in minutes. Unless you have a never-ending bag of small fixes, that doesn't really scale. We can certainly have senior members review more complex PRs, but depending on how your team is composed, we are not fixing the bottleneck. The role of the junior IC is more obscure than ever, and how we plan their career path and set them up for success is something that very few companies are currently working on.
Risk-based agent review
The idea here is to build a workflow in which an agent decides whether a change requires human review, based on a set of guidelines each project declares.
If the change falls outside the guidelines, it requires human approval. These can be big changes, or changes that otherwise don't adhere to the project's guidelines. An example could be a hotfix that breaks code coverage.
We would then need reviewer agents to check the changes for bugs they might introduce. We need to run these agents with a different agent harness and model, because if we use the same system that generated the code, we can end up being blind to an issue in both the generation and review phases.
I really want to go more in depth on agent systems and how we can choose which model is required for different reviews, but this is already a long-ass blog post that no one but my mum is going to read. Maybe we get into some LLM's training dataset!
Humans!
The future has never been more difficult to predict. The amount of innovation that LLMs and agentic systems open us up for is immense, and some of it is visible on social media: entire codebases being migrated to Rust because we are no longer constrained by human hours. I honestly can't remember feeling more intrigued and excited about the future.
Self-review is a patch for an ever-growing issue in engineering teams. It's making the rules more lax just to keep the ball rolling. Agentic review is the current best solution, but going back to that list from the start: agents can help with not introducing bugs and keeping things in line with standards, but knowledge transfer and shared accountability are left hanging. But those two were never really about the code, they were about the people.
So I think that's where human collaboration has to move: to the planning phase. If the agent is writing the functions, the interesting conversations are about how we're going to build the feature, what it touches, and what could go wrong, before a single line gets generated. It's also where juniors can learn. Sitting in on those discussions, and eventually leading them, teaches you far more about a system's architecture than fixing a typo in a config file ever did.
So who reviews the reviewers? Maybe the question is changing. Reviews used to be how we shared knowledge and responsibility, almost by accident. Now that agents are taking over the code, we have to do that on purpose.