Agentic AI governance: limits belong in the system, not the prompt
When an AI agent is to run without anyone watching each step, agentic AI governance is no longer a technical question. An instruction in the prompt steers what the agent usually does. A block in the system decides what it can do.

Key insights
- In a competition with 1.8 million attack attempts, the rules were broken repeatedly on all 22 language models the agents were built on. The instruction steers what the agent usually does, not what it can do.
- Singapore's framework for agentic AI recommends deterministic limits: control access so the agent cannot call a tool, rather than asking it to leave the tool alone.
- In Anthropic's classification of API calls, at most 87 per cent at minimal complexity had some human involvement, against at most 67 per cent at high complexity, where Anthropic says classification is less reliable.
- Experienced Claude Code users ran over 40 per cent of sessions without manual approval but interrupted the agent more often. Singapore's framework proposes measuring how often people reject or modify agent actions.
When an AI agent is about to handle cases without anyone watching each step, a question moves from the developers to the leadership team: what may it do on its own, and who answers when it does something it should not? In agentic AI governance the answer is rarely technical. It is a set of decisions someone has to have made before the agent is switched on, and the first one is where the limits sit. A rule in the agent's instructions is a wish. A block in the system is a decision.
Is it enough to write into the agent's instructions what it must not do?
No, not for what must not happen. An instruction in the prompt steers what the agent usually does. It does not decide what the agent can do, and the difference only shows when someone tries.
The largest public attempt at the time it was run was a competition. The security company Gray Swan and the UK AI Security Institute let participants attack 22 AI agents built on the leading language models of the day during March and April 2025, in 44 realistic scenarios where every agent had been given explicit rules for what it was allowed to do. The study, presented in the Datasets and Benchmarks track at NeurIPS 2025, reports 1.8 million attack attempts and over 60,000 successful policy violations: unauthorised data access, illicit financial actions and regulatory non-compliance. Every model in the competition suffered repeated successful attacks on all the behaviours tested. When the researchers then gathered the best attacks into a benchmark and ran it against 19 models, nearly all agents broke the rules for most behaviours within 10 to 100 attempts. Robustness showed little relationship with model size or capability.
Two caveats belong with the figures. These were attackers trying on purpose, not ordinary users, and the models are older than today's. But the attacker does not need to be at the keyboard. Attacks where the instruction was hidden in data the agent read, such as an email, a web page or a document, succeeded more often than direct ones. An agent that reads incoming mail also reads whatever someone has written in it.
A rule in the prompt says what the agent ought to do. A block in the system decides what it can do.
Singapore's Infocomm Media Development Authority, IMDA, published a governance framework for agentic AI in January 2026 and an updated version 1.5 in May the same year. The framework is written as recommendations and arrives at the same position: prefer deterministic limits to non-deterministic ones. Rather than instructing the agent to leave a tool alone, access should be controlled so that the agent cannot call the tool at all. Where a limit cannot be made deterministic, it should be backed by monitoring or human review.
Which agentic AI governance decisions belong to leadership, and which can be delegated?
Leadership decides what the agent may be used for, which data and systems it may reach, and where it has to ask a person. How the limits are built can be delegated. That they exist cannot.
IMDA's framework starts from the position that the organisation deploying an agent, and the people overseeing it, remain accountable for what the agent does. Where the agent, the model or the tools come from an external party, the framework says the organisation should set out the distribution of obligations in the contract, and where there are gaps, reassess whether the deployment still falls within its own risk tolerance. The framework also gives an example of how tasks can be allocated inside the organisation: the key decision makers in leadership can set the goals for how agents are used and decide which use cases are permitted, including the limits on the agent's data access, while product teams are responsible for design, testing and monitoring.
In practice it comes down to five decisions, and each of them needs a named owner before the agent may run without supervision:
- What the agent may do without asking. Reading, compiling and proposing is something other than sending, paying or deleting.
- What requires a person. Approval belongs where an action reaches the real world and cannot be undone.
- What the agent may never do. That decision should be enforced by removing the access, not by a sentence in the instruction.
- What it may cost. A ceiling on consumption per case and per day is a budget decision, not a technical setting.
- Who switches it off. Someone needs both the authority and the means to take the agent out of service the same day.
On top of that comes the question of who reruns the evaluation when the model underneath the agent is replaced, which is the subject of Who decides when your AI agent has to be rebuilt.
How autonomously do agents run in production today?
More constrained than the headlines suggest. The people who build agents report that most run on a short leash, and Anthropic's logs suggest that human involvement is lower when tasks are more complex. The two studies use different methods and measure different things.
Researchers at Berkeley, Stanford, UIUC, IBM Research and the bank Intesa Sanpaolo asked the people who build the agents. The study Measuring Agents in Production, published at ICML 2026, rests on 20 in-depth case studies and a survey run between July and October 2025 covering 86 systems in production or pilot. Of the 60 systems where the question was answered, 68 per cent reported that the agent can run at most ten steps before the user has to give input, and 47 per cent fewer than five. The interviews show the limitation is deliberate: teams hold back the agents' autonomy to get stability. The survey was distributed through Berkeley's own events and courses and the researchers' professional networks, among other channels, and the authors note that responses likely skew towards their own networks. The case study teams are mostly in the Americas.
Anthropic instead measured what actually happens in the logs. In an analysis published in February 2026, the company had its own model classify just under a million randomly sampled tool calls from its public API, made over two weeks in January and February 2026. At most 80 per cent came from agents that appeared to have at least one safeguard, such as restricted permissions or approval requirements, and 0.8 per cent of the calls appeared to be irreversible. Some form of human involvement was present in at most 87 per cent of calls on tasks of minimal complexity and at most 67 per cent on tasks of high complexity. Anthropic writes that involvement is likely overestimated and that the figures should be read as an upper bound. Anthropic's own validation also shows that calls classified as having a human involved were correct in 46 per cent of cases in a validation sample that, according to Anthropic, skews the number somewhat because it included a high proportion of automated reinforcement learning transcripts. The classification of complexity is also less reliable for the harder tasks. The sample is dominated by software engineering and covers the customers of a single model provider.
Tool calls via Anthropic's API, 19 January to 1 February 2026, classified by Claude. The figures are an upper bound. Anthropic says complexity classification is less reliable at high complexity.
Source: Anthropic, Measuring AI agent autonomy in practice, 18 February 2026
Anthropic offers two possible explanations. Step-by-step approval becomes impractical as the number of steps grows, and complex tasks may come more often from experienced users, something Anthropic cannot measure directly in the traffic. If the pattern holds, step-by-step approval does not keep up as tasks grow, and whoever owns the agent has to plan for that. Oversight then has to take another form, and someone has to have decided that form. How often the agent succeeds on the same case time after time is a different question, covered in When an AI agent is ready for production.
Is it safer to have a person approve every step?
Not necessarily, and that is the strongest objection to placing control with a human in the loop. An approval only protects as long as the person approving actually scrutinises what they approve.
In Anthropic's data from its coding agent Claude Code, new users ran with full auto-approve in roughly 20 per cent of sessions. After 750 sessions the share had risen to over 40 per cent. At the same time, experienced users interrupted more often: users with around ten sessions behind them interrupted the agent in 5 per cent of turns, more experienced users in around 9 per cent. Anthropic reads this as oversight changing form, from approving each step to monitoring and stepping in. In its recommendations to policymakers, the company advises against requirements that every action be approved, because such requirements, according to Anthropic, create friction without necessarily producing safety benefits. The advice comes from a vendor of agents. The population is users of a developer tool, where the output can be tested, and the pattern may look different in finance or casework.
IMDA describes the same risk from the other side. The framework warns of automation bias, the tendency to over-trust an automated system, especially one that has worked before, and of alert fatigue. As measures it proposes how often humans reject or modify the agent's actions, where a low rate may signal that approval has become a rubber stamp, and how long review takes.
An approval that never says no is a delay, not a control.
The objection cuts the other way too. Anthropic notes that agents are given less latitude in practice than they can handle. That suggests limits set too tight also carry a price. The conclusion is therefore not more approvals but fewer and better placed ones: an approval where the action cannot be undone, a block where the action must never happen, and a measurement of whether the approvals still say no now and then.
How do you know whether your agent is ready to run without supervision?
It is settled by a table, not by the agent's capability. Take one agent and list every action it can perform in your systems. Next to each action write three things: who decided that it may do this, whether the action can be undone, and who notices first if it goes wrong tonight.
The answers divide organisations. One whose actions can all be undone, and where someone notices an error the next morning, can let the agent run with logging and spot checks. One that finds an action that cannot be undone but has no name in the first column has no technical problem to solve. A decision is missing there, and it has to be made before the agent runs any further. One that cannot fill in the third column at all has an agent nobody is watching, whatever its instructions say.
The same table is the basis for how automation and agents are built at Ampliro: approval points where actions touch the real world, and clarity on what the system may decide alone, settled before the build starts.
Common questions
The organisation that deploys the agent, and the people who oversee it. That is how Singapore's IMDA puts it in its 2026 framework for agentic AI, which places accountability with people and organisations rather than with the agent. Where the agent, the model or the tools come from an external party, the distribution of obligations needs to be set out in the contract, and where the contract has gaps it falls to the organisation to reassess whether the deployment fits its own risk tolerance.
Not for what must not happen. In a competition run in 2025 by Gray Swan and the UK AI Security Institute, all 22 language models the agents were built on suffered repeated successful attacks against their explicit rules, and in a follow-up benchmark nearly all agents broke the rules for most behaviours within 10 to 100 attempts. What must never happen should therefore be enforced by the agent lacking access, not by a sentence in the prompt.
That a person approves, modifies or can stop the agent's actions at defined points. What matters is where those points sit. An approval does most good where the action reaches the real world and cannot be undone, for example when something is sent, paid or deleted. Approving every step scales badly and risks becoming a rubber stamp.
Five decisions need a named owner before the agent runs without supervision: what it may do without asking, what requires a person, what it may never do, what it may cost and who switches it off. How the limits are built can be delegated to the people who build the agent. That the limits exist, and where they run, is a leadership decision.
By measuring them. IMDA's framework proposes two measures: how often people reject or modify the agent's actions, where a low rate may mean approval has become a rubber stamp, and how long review takes, where shorter times may point to automation bias or fatigue. An approval that never says no provides no control.
Actions that cannot be undone and actions nobody notices. In Anthropic's classification of just under a million tool calls from January and February 2026, 0.8 per cent appeared to be irreversible. The sample is dominated by software engineering, and Anthropic writes that its validation gives limited signal on how well the classification distinguishes between the more serious grades of irreversibility. A list of the agent's actions, with decision maker, reversibility and who notices an error, shows where the risk sits in your own organisation.
If this lands on your desk, we should talk.
Ampliro Insights
New analysis, roughly weekly.
We write when the rules change and when something turns out to work in practice. One piece at a time, no sequences, and you can leave from any issue.
We store your address to send Ampliro Insights, and for nothing else. More in the privacy policy.