If the Work Doesn't Build Judgment, the Manager Has To
Over the last two posts I have been building toward this. First the squeeze, that the work now asks people for discernment earlier than it gives them the experience to develop it. Then the manager's version of the problem, that output no longer reliably shows you what a person can do on their own, because the tool can produce strong work on their behalf and people tend to accept what it gives them. Now, a couple of ideas on how to build judgment in tandem with these changes.
What follows are a couple actions you can take to build discernment and ferret out capability. This is a blend of my experience running a development team and what I see working with the managers I coach.
First, review the reasoning as well as the result. You are going to review the output either way, but that now mostly measures whether the person can instruct AI to produce something that looks right (which is worth knowing) not whether the person can tell when AI is overconfidently wrong.
How do you do this? Start by asking them to walk you through their thought process. What made them choose this approach over the one they set aside, and what they would check first if it broke in production. The point is not to pick apart their thinking. It is to move part of what you are reviewing from the result to the judgment that produced it, because the judgment is what you are trying to grow and it is the part the result has stopped showing you. When someone has leaned on the tool without really engaging, this surfaces it quickly, and it does so without you having to police how anyone uses AI.
In the beginning, they may struggle to explain, or their reasoning may feel shallow or frustrating. As you work with them over time, explaining your reasoning as well, you will begin to build trust and speed up your mutual decision-making. Your employee may surprise you by providing ideas or reasoning you had not considered, and you need to be open to that. None of us are perfect. The eventual goal is to have the questions you ask here echo in their mind when they come to the next big decision, so they refine their thinking at the start rather than during a review.
An important side note is that the way you frame your questions changes the response you get. This is a scenario where we want the employee to talk openly about their approach. Use open-ended questions, the ones that allow the employee to explain their thinking instead of saying yes or no. An example is changing a question like "Is this the only design you considered?" to "What other designs did you consider, or could you consider now?" Also, be careful with questions that start with "why" or create a defensive posture. Why? Because you will get a much more open response if your employee understands you are curious about their thought process instead of critiquing their approach. Instead of "Why did you design it this way?" try "What is the primary goal in this design?" or "How does the design achieve the goal?" or "What other options are there to achieve the same goal?"
One caveat on where you start. A senior engineer usually has enough prior experience to catch a confident error by reasoning about it. Someone earlier in their career often does not, and asking them to produce discernment they have not built yet mostly produces guessing. With them, start by showing the failures rather than asking them to find the failures. Move to the harder version once they have seen what wrong looks like.
The second change is to create the reps the work used to supply on its own. Before getting to what those look like, it is worth sorting what you are actually protecting against, because there are two situations here to consider.
In most work today, a person is still the gate. AI write the code, someone reviews it, and the mistakes that get through are the subtle ones. It compiled. The tests passed. It reads like something a competent engineer wrote. That is a review problem, and the skill it asks for is discernment.
In a growing amount of work, no person sits between the output and the consequence. An agent operating inside your company or inside your product acts, and the review happens afterward if it happens at all. The mistakes there are not more subtle. They are faster, and they travel further before anyone notices or expects. That is a containment problem, and the skill it asks for is knowing what a system can do when it is wrong, before you let it run.
Those are different skills. Someone can be strong at the first and completely untested at the second. It is worth being clear with yourself about which one you are building.
For the review problem, start by having someone reason a problem through before they open AI. Ask them to write down the approach they would take and what they expect to be hard about it. Then let them use the tool. The value is not in the writing down. It is in what comes after, which is sitting with them and comparing what they predicted against what the tool produced, and against what turned out to be true once the work shipped. Where were they right and did not trust it? Where were they wrong in a way they could have caught?
A second version is to hand someone AI output you know has a flaw in it. Before you say anything about where it is, ask them to write down what they think is wrong and why. Then compare. Getting it wrong is fine, and it is most of the value. What you are building is the habit of forming a position and then finding out whether it held.
Set expectations up front on both of these. Their role is to think out loud. Your role is to ask questions that show them where their thinking is thin. Timebox it so nobody is guessing about scope. Do not lead them to the answer and do not pick apart what they come up with.
The written prediction is the part that is easy to skip and the part that makes this work. Reasoning out loud after the fact quietly reshapes itself around what you already know happened. A prediction on the record does not move. It is also what lets you see progress, because you can look back across a few months of predictions and tell whether the gaps are getting smaller.
The containment problem needs a different kind of rep, and low stakes are harder to arrange. You cannot hand someone a sandbox version of an agent that has real permissions in your product.
What you can do is make the reasoning explicit before the permissions are granted. Ask the person to write down what they expect the agent to be able to do, where they expect it to overreach, and what would have to be true for a mistake to stay contained. Then bound it to that and watch what happens over the first few weeks. Come back to what they wrote.
Both of these cost time, which is the resource you probably feel you have the least of right now. If you want to reduce the time something takes later, it takes some amount of additional time now. That upfront effort pays off repeatedly, which is why developing your people matters. The executives above you are under real pressure to show output from these tools, and that pressure is fair. These practices trade some throughput today for people you can trust with throughput later. That is a real trade and you have to make it deliberately, because the default of shipping quietly spends down the judgment you are going to need.
This is a starting point in a rapidly changing landscape. Aspects of how we use AI will certainly change. What I am convinced of is the shape of the problem. For a long time the work developed people for us quietly, while we were busy with everything else, and it does not do that anymore. Whatever the right specifics turn out to be, building judgment is becoming something a manager does on purpose rather than something the job takes care of on its own. That is the change worth planning around, and I would rather start adjusting to it now than keep waiting for the old apprenticeship to come back.