The Feedback Loop No One's Coaching For: Giving Performance Reviews to an AI Employee
- Severin Sorensen

- 3 days ago
- 4 min read
Most executives can now name the AI agent handling their expense reports, drafting their first-pass contracts, or triaging their inbox. Fewer can say what a “good quarter” looks like for that agent, or what happens when it quietly underperforms. While executive teams have gotten comfortable deploying AI colleagues, they have not yet gotten comfortable evaluating them.
The scale of the shift explains the urgency. Gartner projects that by the end of 2026, 40% of enterprise applications will embed task-specific AI agents, up from under 5 percent in 2025 (Gartner, 2025). Field research from MIT’s Initiative on the Digital Economy is already documenting how teams that work alongside AI agents perform differently on real tasks, not hypothetical ones (Ju & Aral, 2025). Agents have moved from pilot projects to production, but what has not moved at the same pace is the discipline of reviewing their work the way a manager reviews a direct report’s.
Providing AI Feedback
Most leaders were trained, at some point, on a model for delivering feedback to a human being. The Situation-Behavior-Impact model, developed by the Center for Creative Leadership, is among the most widely used: name the situation, describe the specific behavior, and state its impact, keeping interpretation and personality out of the conversation entirely (Center for Creative Leadership, 2025). The model works because it forces precision. For example, “You were unreliable this month” invites an argument. Whereas, “In the March renewal cycle, you flagged four contracts as low-risk that later required legal review, which cost the team roughly a week of rework” invites a conversation about what to change.
That same discipline transfers cleanly to an AI agent, with one addition. A human behavior usually has intent behind it worth exploring, which is why CCL’s extended model adds a fourth step: asking about intent to turn feedback into dialogue (Center for Creative Leadership, 2025). An AI agent has no intent in that sense, but it has something that plays a similar role: the configuration that produced the behavior. Call it Situation, Behavior, Impact, and Root Cause.
Situation names the specific task and context, not “the chatbot” in general but the exact workflow, prompt chain, or trigger that ran.
Behavior describes what the agent actually did, in terms as observable as a transcript or an output log allows, resisting the temptation to say the agent “decided” or “chose” when what happened was closer to “produced.”
Impact states the downstream effect in the same terms an executive would use for a human report: hours saved or lost, revenue affected, risk introduced, trust gained or spent with a client or colleague.
Root Cause is where the review does its real work, tracing the behavior back to its source, whether that is the underlying prompt, the tools the agent had access to, the data it was trained or grounded on, or the absence of a guardrail that should have caught the error before it reached a client.
Root Cause is also the hardest step, and the one most organizations have not learned to do well, because it cannot be answered the way a human review answers it. With a person, you can simply ask why. An AI agent cannot answer that question in kind, so the root cause has to be reconstructed rather than asked for, which means actually having the underlying prompt, the tool logs, the data the agent was grounded on, and a clear sense of what guardrail should have caught the error before it reached a client. Most organizations running agents in production do not yet keep that record in a form anyone can review after the fact, which means the step of the model built to turn a finding into a fix is the one most reviews quietly skip.
That gap compounds in a multi-agent workflow, where one agent's output becomes another agent's input before a human ever sees the full chain. An executive should be able to answer a simple question walking into any review: if this chain produces a bad outcome, whose name goes on the corrective action. Research on agentic governance suggests most organizations cannot answer that question with confidence today (Cloud Security Alliance, 2026), the same ambiguity a coach would flag immediately in a human org chart, but one that AI's speed and opacity make far easier to leave unresolved.
The Quarterly Review
The record-keeping problem points toward a solution: review more often, in smaller batches, while the trail is still traceable. An annual AI review, the cadence most organizations default to for talent, asks a reviewer to reconstruct root cause across a year of logs and prompt changes, which is close to impossible and explains why so many of these reviews never happen at all. A quarterly AI performance review, run with the same seriousness as a talent review and folded into the same calendar, keeps the sample small enough that a root cause is still findable and the config that produced it likely has not changed twice since. For each agent operating with meaningful autonomy, the executive team or a designated owner should walk through the Situation-Behavior-Impact-Root Cause structure against a sample of its highest-stakes work from that quarter, name what changed since the last cycle, and make one of three calls: continue as configured, retrain or reconfigure with a specific root cause attached, or retire the agent from that task entirely. The point is to bring the same rigor that prevents human underperformance from going unaddressed for a year to a category of colleague that can now do a year's worth of damage in a week.
The Main Takeaway
The organizations that get this right will be the ones whose leaders treat evaluating an AI agent’s work as seriously as they treat evaluating a person’s, because the coaching skill underneath both is the same: specific, evidence-based feedback that someone is accountable for acting on. That skill was never really about humans; it was about closing the loop between what happened and what happens next, and right now, for most AI agents in most organizations, that loop is still open.
References
Center for Creative Leadership. (2025). SBI feedback model & talent development conversations. https://www.ccl.org/articles/leading-effectively-articles/sbi-feedback-model-a-quick-win-to-improve-talent-conversations-development/
Cloud Security Alliance. (2026). NIST AI Risk Management Framework: Agentic profile. https://labs.cloudsecurityalliance.org/agentic/agentic-nist-ai-rmf-profile-v1/
Gartner Predicts 40% of Enterprise Apps Will Feature Task-Specific AI Agents by 2026, Up from Less Than 5% in 2025. (2025). Gartner. https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025
Ju, H., & Aral, S. (2025). Collaborating with AI agents: Field experiments on teamwork, productivity, and performance. MIT Initiative on the Digital Economy. https://ide.mit.edu/
Copyright © 2026 by Severin Sorensen. All rights reserved.





Comments