Home > Digitalizacija > Pascal Bornet, Jochen Wirtz, Tarja Stephens, Frederique C. Cobert, Shigeki Yamaguchi, Rachel Wood, Rakesh Gohel, Helen Yu, Nima Scehi: The Human-Agent Orchestrator; Leading and Scaling AI-Driven Organizations

Pascal Bornet, Jochen Wirtz, Tarja Stephens, Frederique C. Cobert, Shigeki Yamaguchi, Rachel Wood, Rakesh Gohel, Helen Yu, Nima Scehi: The Human-Agent Orchestrator; Leading and Scaling AI-Driven Organizations

INTRODUCTION: The Perfect AI That Changes Nothing

The problem is that your organization wasn’t built to accept intelligence that doesn’t come with a title, a salary, and a place in the hierarchy. Intelligence without political capital has no power.

The real problem today isn’t the sophistication of the technology.

The technology is over-delivering. Organizations are under-adapting.

Here’s the most sobering calculation we’ve done: even if AI development stopped today, with no further improvements, it would still take organizations 5 to 10 years to fully capture the value of the technology we already have.

The Efficiency Paradox is real: speed without rethinking just creates more work. We’re treating AI as a plug-and-play software upgrade. What it actually demands is a redesign of the work itself.

Deploying AI quietly does something far more significant to every person on your team than adding a new tool. Something most of them aren’t prepared for. Something most leaders don’t even notice until the damage is done. Solid technical solutions do not guarantee success. I had brilliant AI doing exactly what I designed it to do. And I almost destroyed a team in the process. The question I forgot to answer had nothing to do with technology. Most of what you know about management still works. But there’s a subset of principles, ones you’ve relied on your entire career, that become actively dangerous when applied to AI agents. Knowing which ones, and why, changes everything. Speed without a foundation is a slow disaster. Using AI to do the same work faster feels like progress, and it fools almost everyone for the first six months. The one shift that separates organizations that transform from those that accelerate toward the wrong destination is building the management layer before the technical one.

Trust works differently here than anywhere else you have managed. With people, trust is earned through consistency, transparency, and track record. With AI agents, trust must be engineered through evidence, boundaries, and verification. The good news: the new system, once built, is more reliable than the old one because it is explicit rather than assumed. There is a trap that catches almost every leader who takes AI seriously. The more conscientious you are, the more likely you are to fall into it. It doesn’t look like a mistake. It looks like diligence.

Investigation took three forms.

  • The first was expert interviews.
  • The second was data.
  • The third was depth.

Through that research, we identified three distinct levels of organizational readiness for AI.

  • Level 1: Tool Adoption, where more than 60% of organizations currently live.
  • Level 2: Workflow Redesign, where almost 25% of organizations are starting to operate.
  • Level 3: Organizational Transformation, where only roughly 15% of organizations have reached.

Leading teams where some members are human and some are AI agents. We call this challenge Hybrid Management.

From Supervising to Orchestrating

The Promotion Nobody Asked For

The moment you start working with AI that reasons, plans, and acts, you become a manager, whether you want the job or not. Not metaphorically. Literally. You are now responsible for setting direction, evaluating output, catching errors, and deciding what good looks like. These are management activities.

AI that is not properly directed does not do nothing. It does the wrong things, confidently and at speed.

The Last Generation

“Who’s actually going to manage these AI systems? And when an AI agent makes a decision that affects my team, who’s responsible?”We gave her the consultant’s answer. Governance frameworks. Clear ownership models.

She wasn’t satisfied.”No, I mean really. If I have a team of 10 people and they’re each using five AI agents, am I managing 10 people or 60 entities? Because those agents are making decisions, taking actions, and influencing outcomes. And I have no idea how to lead them.”

Go back to 1911. Taylor publishes The Principles of Scientific Management, 17 and the world of work transforms almost overnight. Taylor’s core idea was simple and radical: break work into discrete tasks, measure everything, have managers plan, and workers execute. It sounds obvious now. At the time, it was revolutionary. Productivity exploded. The assembly line became possible. Modern manufacturing was born.

Management theory evolved around the edges. We got better at motivation with Maslow. 18 Better at delegation with Hersey and Blanchard. 19 Better at systems thinking with Deming. 20 But the core model never changed because the core assumption never changed.

Humans do the deciding. Machines do the doing.

The principles of good management turn out to be universal. They apply to any intelligent actor you’re coordinating, whether biological or digital. Clarity of objectives. Progressive delegation. Feedback loops. Trust is built through consistent performance. These aren’t human management principles. They’re coordination principles. And coordination is coordination, whether you’re coordinating people or agents.

AI agents have no psychological needs. They have no motivation to unlock. When leaders apply motivational language to agents by crafting inspiring system prompts, framing the agent’s”purpose,”or thanking it for good work, they are not just wasting effort. They’re revealing a mental model that will cause failures in less obvious places.

Communication. With humans, you communicate in context. You rely on inference, shared history, and unspoken norms. AI agents cannot read between the lines.

Error response. You are essentially bilingual. When you turn to your team member, you’re speaking one language. When you turn to the agent dashboard, you’re speaking another.

When agents join your team, three things shift simultaneously in your human workforce, and if you don’t manage them deliberately, each one becomes a scaling constraint.

  • The first is the Identity Question.
  • The second is the Boundary Problem. Once agents are handling work that used to belong to your team, role boundaries become genuinely unclear in ways they weren’t before.
  • The third is the Development Gap. If agents take over the work where junior and mid-level people used to learn by doing, those people stop developing unless you deliberately redesign the path.

The Supervision Trap

This is what we call the Supervision Trap. It is not a technology problem. It is a management problem: the mistake of applying 20th-century management logic to 21st-century systems.

The Human Crumple Zone is a harder version of this trap that most organizations do not see until they are inside a regulatory investigation or a client post-mortem.

This is what the supervision trap produces at its worst. Not just inefficiency. Not just burnout. A human being who carries formal accountability for outputs they had no realistic ability to oversee, and who becomes the proximate cause when something goes wrong. The organization’s system operated correctly, its lawyers will note. The human failed to catch the error.

The supervision trap is not only inefficient. In its most serious form, it is a liability architecture. Supervision doesn’t scale to agent speed. It never will.

The pattern in the best-performing teams was consistent. Humans defined objectives and constraints: the what and the why. Agents explored solution spaces: the how. Humans made final decisions on trade-offs: the which. And the system captured learning for continuous improvement: the next time. That’s orchestration.

True Orchestrators achieved an average productivity gain of 73%. The largest group, still operating as supervisors, averaged 28%.

The supervision instinct. It’s the part of your brain that equates leadership with oversight. Supervision doesn’t just feel responsible. It feels like evidence that you’re earning your salary. Orchestration breaks that logic entirely.

There’s a second trap. The moment your agent makes a visible mistake, your brain responds in a way that research has documented precisely: we feel losses roughly twice as intensely as equivalent gains.

And there’s a third trap We systematically overestimate how much our direct involvement actually improves things. “As a pilot, you have to do so many mandatory auto takeoffs and landings, and the plane flies itself. So basically, you’re not really a pilot anymore. You’re an in-flight technology manager. The skill is not flying the plane. The skill is knowing when to let the plane fly itself, and when to take back control.”

The skill of the modern leader is not doing or reviewing the work. It is knowing when to delegate to agents, when to trust their outputs, when to override, and how to remain genuinely in command of a system you are not directly operating.

Agents are components. The system is the multiplier.

Before you run the system, three things.

  • First: redefine what competence means.
  • Second: make letting go rule-based rather than emotional.
  • Third: build tolerance for intelligent failure.

AI generates execution. Humans generate judgment. Leaders generate coherence.

Classic management research suggested a leader could effectively supervise five to seven people.

If you’re not reviewing individual outputs but designing systems, monitoring flows, and handling exceptions, your span of control becomes theoretically limitless.

Asana’s Anatomy of Work Index, which surveys knowledge workers annually across more than 15 countries, consistently finds that 60% of professional time goes to”work about work”: meetings, status updates, alignment conversations, and email threads establishing a shared understanding of things one person already knows.

This is what we call the Coordination Tax.

That is Augmentation: the AI makes you faster at existing tasks.

What we call Amplification (or Agentic Amplification) differs in three specific ways.

  • First, the architecture: augmentation is a tool you use; amplification is a system you build,
  • Second, the direction of adaptation: in augmentation, you adapt to the tool; in amplification, the system is designed around your intent and your particular way of creating value,
  • Third, the reach: augmentation makes you faster at what you already do; amplification extends what only you can do into decisions you’re not personally making, situations you’re not personally in, at a scale you could never reach alone.

Building the System That Runs Without You

The Orchestration Design Canvas

It has six layers, each building on the last.

The framework is easy to remember as the 6S: Source, Success, Safety, Steering, Switch, and Sharpen.

Together, the six layers are broken down into 32 components. Each component has a unique code tied to its layer.

This framework applies in two contexts: Hybrid Team Orchestration, where humans and agents work side by side, and Personal Orchestration, where a team of agents works around your own intent and expertise.

The Source Layer is the foundation. It defines why the system exists, what unique value humans are meant to contribute, and which human skills must be preserved as agents take over more execution.

KEY COMPONENTS

  • S11 Intent statement: Write the mandate in one sentence, specific enough to make clear what the system is for and what it is not for
  • S12 Core value engine: 3–5 human capabilities that create value here and cannot simply be handed to agents.
  • S13 Human team foundation (hybrid team leaders only): What each person contributes that remains uniquely valuable and what they must continue to develop.
  • S14 Skills at risk of atrophy: Which human skills does this deployment reduce practice of, and the maintenance practice needed to keep each one strong.

The Success Layer defines what constitutes a good output from the receiver’s point of view, not the builder’s.

KEY COMPONENTS

  • S21 Outcome statement: What the agent needs to achieve, not just produce. Be specific.
  • S22 Audience profile: Who receives the output, what they know, what they need to do with it.
  • S23 Context paragraph: Why this outcome matters. This allows the agent to make sensible judgment calls.
  • S24 Quality bar: Concrete examples of good outputs and bad outputs. Both.
  • S25 Business outcome target: The one business metric this agent moves, by how much, over 90 days.

The Safety Layer builds the governance structure before the agent runs. It assigns clear accountability, defines which actions the agent is forbidden to take, separates real constraints from unenforced rules, and creates the decision log required for auditability.

KEY COMPONENTS

  • S31 Named owner: One person accountable for every output this agent produces.
  • S32 Constraints list: Specific actions the agent must never take without human approval.
  • S33 Quality bar: Minimum acceptable standard for outputs. Passing and failing are clearly defined.
  • S34 Validation checklist: The checklist that the system runs against every output before delivery. Logged.
  • S35 Explainability check: Named owner can explain decision logic to a client, auditor, or regulator.
  • S36 Technical vs. human gate: Each constraint classified as architecture-enforced or person-checked.
  • S37 Decision log: Inputs, output, timestamp, confidence score, reviewer identity. Recorded for every decision.
  • S38 Access scope: Authorized data sources, tools, and prohibited data. Principle of least privilege.
  • S39 Input contract: What inputs the agent can act on, and what it does when inputs are missing, stale, or conflicting.
  • S3A Retirement criteria: The specific conditions under which the agent must be retired, rebuilt, or replaced rather than patched again.

The Steering Layer defines decision authority before the agent produces its first output. It specifies which decisions the agent may make on its own, which require human approval, and which must always remain with a human.

KEY COMPONENTS

  • S41 Decision rights map: Every decision type assigned: agent decides autonomously / agent recommends and human approves / agent escalates immediately for human decision.
  • S42 Dependency map: Upstream inputs and downstream consumers: what breaks if this agent fails.

The Switch Layer defines the conditions that stop the normal flow and require immediate human involvement, no matter who would normally have decision authority.

KEY COMPONENTS

  • S51 Exception table: Trigger condition, escalation path, required response time. Every row binary.
  • S52 Default state per trigger: What the agent does while waiting for escalation. Defined at design time.
  • S53 Secondary escalation: What happens if the named person does not respond in time.
  • S54 Manual fallback protocol: Who executes the manual workflow when the agent is simply unavailable.

The Sharpen Layer is what keeps the system improving instead of quietly getting worse. It detects errors, turns them into structural learning, and shows when the system no longer fits the environment it is operating in.

KEY COMPONENTS

  • S61 Spot-check protocol: 5–10% of outputs reviewed weekly against the S21 quality bar and logged.
  • S62 Automated validation layer: The checklist that the system runs against every output before delivery.
  • S63 Human expert review trigger: Which output category always receives human expert review before release.
  • S64 Post-incident brief: Every exception, override, and caught error reviewed. One-off or gap in a layer?
  • S65 Drift monitor: Baseline, threshold, and response defined for each signal.
  • S66 Human Layer Index: Quarterly: Are humans doing more complex work than 90 days ago? Scale 1–5.
  • S67 Five governance metrics: Output volume, error rate, exception rate, override rate, Human Layer Index. Monitored from day one .
  • S68 Business outcome indicator: Monthly: Is the primary metric from S25 moving in the right direction?

The Pre-Flight Check is part of the Orchestration Design Canvas framework.

The check has three parts.

  • PART 1: DO THE LAYERS CONNECT?
  • PART 2: ARE THE HUMANS READY?
  • PART 3: HAVE YOU STRESS-TESTED IT?

Configuration Is Not a Foundation

Every agent implementation

There are three tiers.

  • Tier 3 is the configuration layer. This is the technical work: model selection, prompt engineering, API connections, infrastructure, and deployment pipelines.
  • Tier 2 is the design layer. It defines what the agent must do in terms that business owners can understand and approve.
  • Tier 1 is the foundation layer. This is where you decide, before anything is built, why the system exists, what it is allowed to do, what it must never do, and who is accountable.

Agents do not take control. Humans abdicate it.

The Space the System Creates

You’re teaching her orchestration: how to think about decision rights, how to design feedback loops, how to know when to intervene and when to trust the system.

The transition from supervising to orchestrating requires letting go of the very behaviors that made you successful.

This is the shift from industrial-age leadership to algorithmic-age leadership. From assembly line manager to ecosystem architect. From the person who watches the work get done, to the person who designs the conditions that make good work inevitable.

We are the last generation to have managed only humans.

Different Players, Different Rules

What Is My Job Now?

“I optimized the output. I forgot about the people.”

We have seen this pattern across industries and company sizes. Leaders deploy agents with extraordinary technical skill, while paying almost no attention to the humans still in the room.

71% of the leaders we studied reported making no changes to their human motivation practices after deploying AI agents.

The real productivity losses live in that gap: between leaders who assume motivation will take care of itself and people who quietly wonder why they are still there. They don’t appear in output dashboards. They appear in attrition figures and in the quiet withdrawal of discretionary effort.

The arrival of agents doesn’t create entirely new motivational challenges. It supercharges three failure modes that were already latent in well-managed teams,

The first is what we call Scope Collapse. The fix requires a counterintuitive move: expand human scope upward rather than eliminate it.

The second failure mode is what we call Mastery Vacuum. People don’t automatically migrate toward new forms of expertise. The leader prevents this by naming the new mastery frontiers early, clearly, and publicly.

Four frontiers consistently emerge as genuinely sophisticated and genuinely available to humans.

  • The first is systems thinking.
  • The second is judgment in ambiguity.
  • The third is stakeholder translation.
  • The fourth is agent orchestration itself.

The third failure mode is more subtle and arguably more dangerous. We call it Purpose Drift.

The human is not reviewing the output. They are what we call Architects of Intent.

After years of watching agents perform well and perform badly, we’ve found the answer lives in three concepts: Clarity, Context, and Correction.

  • Clarity means that an agent performs according to its objective function, and that function must be specific enough to be unambiguous.
  • Context means that agents produce outputs proportional to the inputs you give them.
  • Correction means that agents don’t grow through encouragement. They improve through precise, structural correction.

An agent underperforms not because it lacks willpower but because its instructions are ambiguous, its constraints are poorly defined, or its feedback mechanisms are broken. The managerial response is diagnostic, not motivational.

The tendency to treat agents as if they have human needs. We call this the Anthropomorphism Trap.

Agents deserve excellent management: precise objectives, well-designed workflows, rigorous review, and systematic correction.

Delegating Judgment

The central problem of delegation in a hybrid team is the gap between what an agent can do and what the situation requires.

84% of the 108 leaders we interviewed said the main cause of agent deployment problems was assuming the agent would infer intent from incomplete instructions, the way a competent human does through experience.

This is precisely what the Success Layer of the Orchestration Design Canvas exists to prevent. You don’t just say”complete the task.”You define what good looks like, in this context, for this audience, and in this situation.

The managers who actually delegate well don’t start with the task. They start with the outcome.

And they do three things consistently that most managers do only occasionally.

  • The first is to invest heavily in examples of good judgment rather than in instructions alone.
  • The second is defining constraints explicitly before the task begins.
  • The third is specifying the audience before the deliverable.

Three specific habits that make you an effective delegator with human teams become actively harmful when applied to agents.

  • The Method Trap.
  • The Relationship Shortcut. The agent has no shared context.
  • The Patience Instinct.

The Trust Architecture

  • The Reliability Question: for exactly what tasks is this agent dependable, and where does that dependability end?
  • The Conditions Question: under what changes in context, in client, in stakes, should you be reviewing its outputs more carefully?
  • The Autonomy Question: at what level of autonomy is it currently operating, and is that the level you consciously designed, or the level that accumulated while you were doing other things?
  • And the Capability Question: what human skill, in you or your team, have you kept sharp enough to catch a confident error from this agent if one arrives?

Automated systems fail at genuinely novel situations, not by expressing doubt, but by producing a confident answer to the wrong question.

That gap, between a system that is technically right and a situation that has moved outside its frame, is the Trust Architecture challenge every hybrid team leader faces.

Over-trust means relying on an agent beyond the conditions for which it was designed.

Under-trust is quieter and more costly.

Roger Mayer, James Davis, and David Schoorman identified the Three Pillars of Interpersonal Trust: Ability, Benevolence, and Integrity.

Trust does not arrive fully formed. It evolves through stages as a relationship deepens.

  • Calculus-based Trust comes first.
  • Knowledge-based Trust develops as experience accumulates.
  • Finally, Identification-based Trust emerges in the deepest working relationships.

We identified five conditions that underpin trust between humans and AI agents. The first four are design conditions: things you specify before the agent runs its first task. The fifth is a meaning condition: something you sustain through leadership over time.

  • Predictability is built through pre-deployment testing.
  • Explainability is built through choosing or configuring agents whose reasoning can be interrogated.
  • Calibrated Autonomy is built through the graduated rollout process.
  • Transparency about limits.
  • The fifth condition is different in kind from the first four. Frederique Corbett, our co-author, who researches the psychology of human-AI collaboration, names it Meaning Alignment.

The sequence matters as much as the principle.

Graduated trust is not just about extending autonomy slowly. It is about keeping the human capability sharp enough to catch the confident error when it arrives.

The Autonomy Dial

Andy Grove, Intel’s legendary CEO, developed his Task-relevant maturity framework and formalized it in High Output Management, one of the most practically useful ideas in management literature.

The Capability Axis asks: how much has this person proved they can do? The Criticality Axis asks: What happens if they get it wrong today? High skill on a low-stakes case: operate freely. Low skill on a high-stakes case: the senior surgeon stays in the room.

The Autonomy Matrix is that same judgment, made explicit and applied to agents.

The matrix is not a deployment decision. It is a management practice. Where an agent belongs at deployment is rarely where it belongs six months later. Capability changes. Criticality changes. The organization changes. What most organizations lack is not the initial placement decision. It is the systematic process for revisiting that placement as conditions evolve.

The autonomy matrix tells you how to configure the agent. The motivation practice tells you how to reconfigure the human’s sense of contribution at the same time. The two decisions are not independent.

Autonomy Calibration was essential or highly important. Autonomy should be earned through task criticality and track record, not granted for convenience. It ranked third among all orchestration principles.

Knowing When to Take the Wheel

The real challenge is not choosing between human control and system autonomy. It is designing the transfer between them, so that when the system reaches its limit, the human is ready.

Normalization of Deviance is one of three consistent failure modes in agent intervention design. The other two are equally common and equally invisible. Automation Bias is the first. Alert Fatigue is the second.

Three types of triggers cover most agent intervention needs.

  • Threshold Triggers.
  • Category Triggers.
  • Absence Triggers.

Closing the Loop

An error is not the same as a learning signal. An error is simply what went wrong. A learning signal is information turned into a clear change that the system can use next time.

With agents, behavior does not improve because of conversation. It improves when you change the instruction, the rule, or the process before the next run.

Running the Team

Before You Build

Before any agent onboards the team, four questions must be answered.

  • Question One: What Problem Are We Actually Solving?
  • Question Two: Human, Agent, or Hybrid? We have designed the Workforce Decision Matrix to evaluate this against five dimensions before reaching any deployment decision.
  • The first is input variability.
  • The second is judgment.
  • The third is error tolerance.
  • The fourth is regulatory or reputational risk.
  • The fifth is relationship centrality.
  • Question Three: Can It Say No?
  • Question Four: Is Your Team Ready?

The Org Chart That No Longer Fits

Three structural approaches bridge the gap between the hierarchy you have and the system you are building toward.

The first is Agents as Direct Reports. Each agent has a named owner, a defined scope, and a position on the chart alongside human team members.

The second is Agents as Pods. Agents are grouped by capability or workflow and treated as a shared function with a structural owner.

The third is Agents as Infrastructure, managed by a technical function and accessed by business users through interfaces.

When Things Go Wrong

The person accountable for the business outcome is the person accountable for the agent producing it.

IT has an essential role in governance, observability, and security. But IT ownership of an agent’s outputs is a category error.

Detecting drift requires a second instrument in the Sharpen Layer: one that tracks patterns in the agent’s outputs over time: average values creeping, input distributions shifting, or confidence scores changing, instead of waiting for a single output to fail.

Becoming the Orchestrator

The frameworks work. The Canvas is useful. The Accountability Map, the Failure Playbook, the Governance Structure — these are real instruments that produce real results. But none of them substitutes for the development of the human at the center of them. The Orchestrator who compounds returns from a hybrid team over three years is not the same person who started. Something has changed in how they think, what they protect, and what they are willing to let go of.

How Orchestrators Are Made

technical background had essentially no correlation with orchestration performance. What correlated was operational instinct, outcome clarity, and the willingness to have uncomfortable conversations before they became crises.

A precise statement of what the agent would handle and what it would not, and a named account of what the humans in that workflow brought that the agent could not replicate.

the Human Layer Index: whether her human team members were doing more complex, higher-judgment work than 90 days before.

Technical success and human erosion are not opposites; they can coexist for months before one of them wins.

How Orchestrators Compound

When you scale, the bottleneck is not technical capacity. It is governance capacity. You need people who can read the system and judge whether the design still matches the world, not just whether it makes internal sense. That capacity cannot be recruited. It must be developed from people who already understand your context.

Environmental Currency Reading. It is about checking whether an Orchestration Design Canvas still describes the world the agents are actually operating in.

Accountability without skill is not governance. It is the appearance of governance.

Build. Sense. Rethink. You build the system. You sense drift before it breaks trust. You rethink the design before yesterday’s logic becomes tomorrow’s failure.

After every significant system event, whether a failed deployment, a wrong decision, or a surprising result, she runs two post-mortems. The Orchestration Design Canvas post-incident brief covers the system: which layer failed, what needs to change. The second one is private. It covers her own reasoning at the time she made the design decision.

The Identity Work Nobody Does

William Bridges introduced the Neutral Zone to describe what most change managers experience as the awkward middle. Between what ends and what begins, there is a liminal space that is neither the old identity nor the new one — where the grief is processed, where old competence frameworks are relinquished, where new ones become genuinely one’s own rather than externally imposed.

Most organizations rush people past this zone.

The edge in hybrid team leadership is the willingness to do this work: to move through the Neutral Zone with enough honesty to actually arrive somewhere new rather than simply performing arrival while rebuilding the old model in the new system’s clothes.

Keeping the Edge

Automation-Induced Skill Atrophy.

It is distinct from ordinary skill erosion because the system that causes it looks like success while it is happening. The outputs are good. The efficiency metrics are strong. The team is hitting its targets. The erosion is invisible until the moment it becomes consequential,

We observe three categories of skills that show early atrophy signals

  • The first is raw Analytical Capability.
  • The second is Judgment under Ambiguity.
  • The third category is the most consequential and the least visible: Relational Skills.

The specific skill most at risk is the one that agents make structurally unnecessary: Active Listening.

The invisible erosion is not a technology problem. It is not a governance problem. It is a practice problem.

You have the architecture. You have the operational practice. You have the developmental map. What happens next depends on what you do with the space the agents create. Most leaders fill it with more execution. The ones who compound fill it with the work that only they can do. That has always been the definition of leadership worth having. The agents just make it visible.