
Insights · Blog
How to Lead Your Non-Human Team
The first wave of office AI lived in a chat window.
You asked a question, received an answer and copied it somewhere else. A better search bar with a talent for prose.
That model is already dated.
An AI agent can now retain project context, load specialist skills, search company systems, operate a browser, delegate parts of a job to other agents and continue working in the background. Across the major platforms, persistent memory, reusable skills, authenticated connectors, scheduled workflows and computer use are becoming normal parts of the stack. (OpenAI, Google, Microsoft)
If you use this capability only for drafting emails, you are leaving much of its value untouched. If you give it broad access and walk away, you are creating a different problem.
Management is therefore a useful way to think about AI, with one important qualification: these systems are not people. They do not need motivation, career plans or Friday drinks. They need clear outcomes, reliable context, appropriate access and regular inspection.
The management unit is no longer the model. It is the whole work system around it: charter, context, skills, tools, permissions, memory, evaluation and accountable owner.
The emerging division of labour
A recent Anthropic analysis of roughly 400,000 agentic coding sessions found that people made around 70% of planning decisions, while the AI made around 80% of execution decisions. Domain expertise remained a strong predictor of success. This is vendor telemetry rather than audited business-performance data, but the pattern is useful. (Anthropic)
The human defines the objective, supplies context, sets limits and accepts the result. The agent handles a growing share of the bounded execution.
That requires more than a clever prompt.
Understand the parts of an AI team
The vocabulary is becoming confusing because vendors use “agent” for almost everything. A simpler model helps.
| AI element | Business equivalent | Best use |
|---|---|---|
| Prompt | An individual assignment | One-off questions and tasks |
| Project or workspace | Shared case file | Persistent context, files and instructions |
| Skill | Standard operating procedure | A reusable method, template or body of expertise |
| Agent charter | Job description and authority schedule | Mission, scope, standards and boundaries |
| Agent | A bounded operating role | Adaptive, multi-step work using tools |
| Tool or connector | Equipment and system access | Searching, calculating, updating or communicating |
| Workflow or automation | Production process | Predictable work triggered by time or an event |
| Subagent | Temporary specialist | Independent or parallel parts of a larger job |
| Evaluation and trace | Scorecard and audit trail | Testing quality and understanding failures |
| Human owner | Accountable executive | Judgment, approval and consequences |
A skill is not an employee. It is a procedure the agent can load when required. It might contain your writing principles, due-diligence method, financial-analysis process, meeting-preparation checklist, templates and examples.
An agent combines a model with a charter, selected skills, tools, memory and operating limits.
This distinction matters. If you want every report to follow your house style, you probably need a writing skill. You do not necessarily need a “Chief Writing Officer Agent.”
Draw the process map before the org chart
The temptation is to recreate the company in miniature:
- A marketing agent
- A finance agent
- A strategy agent
- A legal agent
- A synthetic CEO to supervise them all
This can quickly become corporate theatre.
Current evidence suggests starting with one capable agent and adding specialists only when the work justifies the coordination cost. OpenAI’s own guidance recommends extending a single agent first. Separate agents become useful when tasks require different data, tools, permissions, evaluation standards or genuinely parallel work. (OpenAI)
Google Research tested 180 agent configurations and found that multi-agent systems helped with parallelisable work but reduced performance on sequential planning tasks by 39% to 70%. Central coordination also contained errors better than independent groups of agents. (Google Research)
Use the following rules:
- Keep a one-off exploration in a normal conversation or project.
- Turn a method you repeat into a skill.
- Use a conventional workflow when the steps and decisions are predictable.
- Use an agent when the job requires interpretation, adaptation and several tools.
- Add a subagent when part of the work can happen independently or in parallel.
- Create a separate permanent agent when permissions, data boundaries or accountability differ.
- Add an orchestrator when several specialists need one point of coordination.
The org chart is useful as a map of responsibilities and authority. It should not be an excuse to dress one model in five costumes.
A practical AI org chart
A small business might eventually use this structure:
- Human owner
- Coordinator: routes work, maintains context and synthesises results
- Researcher: read-only access; gathers and cites evidence
- Analyst or maker: analyses data and creates reports, models or drafts
- Reviewer: checks evidence, calculations and acceptance criteria
- Operator: uses write-access tools within limits and approval rules
- Coordinator: routes work, maintains context and synthesises results
Most companies will not need every role on day one.
The coordinator should have little or no high-risk access. Its job is to understand the assignment, route the work and assemble the result.
The researcher should be read-only and required to show its sources.
The maker can produce a proposal, analysis, presentation or campaign, but should not automatically publish it.
The reviewer should work from explicit criteria and, for important decisions, independent context or evidence. Five agents using the same model and the same source material are not five independent opinions.
The operator receives the narrowest write access possible. Sending messages, changing records, making purchases, deleting information and publishing externally deserve different approval rules.
Give each agent an operating charter
“Act as a brilliant marketing director” is a character sketch. It is not a job description.
A useful agent charter contains:
-
Outcome: What result is this agent responsible for producing?
-
Trigger and inputs: When does the work begin, and what must be supplied?
-
Scope: What is included, and what sits outside the role?
-
Approved sources: Which files, systems and data can it treat as authoritative?
-
Skills: Which procedures, templates and examples should it use?
-
Tools and permissions: What may it read, create, change or send?
-
Evidence rules: Which claims require sources, calculations or uncertainty labels?
-
Definition of done: What must be true before the agent can report completion?
-
Escalation conditions: When should it stop and ask a human?
-
Human owner: Who approves the result and answers for its consequences?
-
Evaluation set: Which representative cases will be used to test it?
-
Review date: When will its instructions, access and continued usefulness be reassessed?
Version the charter and the skills. When output changes, you need to know whether the cause was a new model, a new instruction, different data or changed access.
Grant authority in stages
Autonomy should be awarded per workflow. An agent that can autonomously prepare a morning briefing does not automatically earn the right to update the CRM or contact a client.
A useful authority ladder is:
- Level 0: Read and answer. No external action.
- Level 1: Draft and recommend. A human decides what is used.
- Level 2: Perform reversible actions. Work is limited, logged and easy to undo.
- Level 3: Prepare consequential actions for approval. The human confirms before execution.
- Level 4: Operate unattended inside narrow limits. Monitoring, stopping rules and recovery procedures are in place.
Begin with read-only access. Promote a workflow only after it has passed real cases repeatedly.
Control the authority before adjusting the personality.
Sycophancy is still real
AI systems still have a tendency to validate the user’s assumptions. A 2026 Nature paper found that training models to sound warmer could reduce accuracy and increase sycophancy. Model providers have reported substantial improvements on their own tests, but the behaviour has not disappeared. (Nature, OpenAI)
It deserves a place in the evaluation set. It no longer deserves to dominate the entire management model.
Once an agent has memory, tools and time to act, the risk surface becomes wider:
- It can misread or misrepresent a document.
- It can omit information that would change the decision.
- It can claim completion despite unfinished work.
- It can carry a poor assumption through a long chain of actions.
- It can follow malicious instructions hidden in a website, email or file.
- It can recall stale or inappropriate information from memory.
- It can exceed the intended scope of the assignment.
- It can approve its own weak work with complete confidence.
Current model evaluations increasingly track these behaviours alongside sycophancy. (Anthropic)
Telling an agent to “be sceptical” is a weak control. A reviewer should inspect sources, calculations, omissions and acceptance criteria. Sensitive actions need restricted permissions, explicit approval and an audit trail. Prompt injection also requires controls around tools and data transfer; it cannot be solved by wording alone. (OpenAI)
Run performance reviews on the work
AI output often looks convincing before it is useful. Fluency is therefore a poor performance measure.
Track:
- Accepted outputs or completed business outcomes
- Human corrections and rework
- Unsupported or incorrect claims
- Important omissions
- Cycle time
- Cost per successful outcome
- Escalations and exceptions
- Human interventions
- Permission or policy violations
- False completion claims
Keep a small set of representative test cases, including ambiguous inputs, missing data and adversarial material. Run them again whenever you change the model, charter, skill or toolset.
Review a sample of live work every week. Record failures by cause rather than merely declaring that “the AI got it wrong.” The failure may sit in the instructions, source material, permissions, workflow design or acceptance criteria.
Connect AI activity to business results
The evidence on business impact is encouraging and uneven.
A study of 5,172 customer-support agents found a 15% average increase in issues resolved per hour, rising to roughly 34% for newer and lower-skilled workers. Experienced employees saw far smaller gains. (Quarterly Journal of Economics)
In a 2026 experiment involving 791 professionals at Procter & Gamble, individuals working with AI produced innovation proposals of similar quality to two-person human teams. AI improved the ideas being generated, but it did not improve the participants’ ability to select their best idea. (Organization Science)
A separate field experiment across 66 companies and 7,137 workers found that active Microsoft 365 Copilot users spent about two fewer hours per week on email. Researchers found no corresponding change in the broader quantity or composition of their work. (American Economic Review: Insights)
The implication is straightforward: individual time savings do not redesign a business.
Management must decide what happens to the capacity AI releases. Will the company respond faster, serve more customers, conduct deeper analysis, improve quality, offer a new service or reduce risk? Without that decision, saved minutes disappear into the working day.
Choose the intended business effect before building the agent:
- More throughput
- Shorter cycle times
- Better quality or coverage
- Lower service cost
- Fewer errors
- Faster decisions
- A product or service that was previously uneconomic
Measure that result against a baseline.
Protect the development of your human team
Delegation can increase output while weakening learning.
In a small 2026 experiment, junior developers using AI scored 17 percentage points lower on a subsequent test of the material they had been working with. Those who used AI for explanations and conceptual questions retained more understanding than those who mainly delegated the work. (Anthropic)
Companies therefore need two modes.
In production mode, people should delegate repeatable execution and use AI to extend their capacity.
In learning mode, employees should explain decisions, reconstruct outputs, debug failures and sometimes complete work without assistance.
A company that removes every difficult task from its junior employees may eventually discover that nobody has developed the judgment required to supervise the automation.
The human role remains substantial: choosing the objective, integrating context, resolving value conflicts, selecting among plausible outputs and accepting responsibility.
A four-week setup plan
Week 1: Select the work
Choose one recurring workflow with a clear output and moderate volume.
Document how it works today. Record cycle time, quality problems, cost and human effort. Identify the decisions that genuinely require judgment.
Week 2: Build the first operating unit
Create one project with authoritative material and relevant examples.
Write one agent charter. Turn one or two proven procedures into reusable skills. Connect only the minimum systems required, starting with read-only access.
Keep predictable steps in a conventional workflow. Give the agent discretion only where adaptation adds value.
Week 3: Test reality
Build a test set of 10 to 20 real cases. Include excellent inputs, incomplete inputs, contradictory data and at least one hostile or misleading document.
Measure accuracy, omissions, rework, time and cost. Add a reviewer for decisions where evidence matters.
Do not let the agent grade itself in isolation.
Week 4: Operate in shadow mode
Let the agent prepare real work while the existing process remains in place.
Compare the results. Review failures. Update the charter or skills. If performance is consistent, allow one narrow and reversible action.
Sending, publishing, purchasing, paying and deleting can wait until the system has earned that authority.
Your first AI team should look rather boring: one owner, one recurring job, one agent, a few skills, limited access and a visible scorecard.
That is a healthy beginning.
On Monday, choose one job your company already repeats. Define what a good result looks like. Build from there.