October 9, 2026

How to keep AI agents within their permissions

AI agents can use valid credentials to perform actions beyond their assigned permissions, creating risks that traditional access controls may not prevent. Token Security explains how organizations can enforce agent-specific policies without sacrificing autonomy.

How to keep AI agents within their permissions

Written by: Ido Shlomo Co-Founder and CTO of Token Security

AI agents should do more on their own. Who has the time or attention span to approve every command? And watching agents do the work they’re asked to do is mind-numbingly boring.

But agents do need boundaries, especially in corporate environments. Without a clear limit on what they can do, agentic flows tend to use all available access.

Here’s a real-world example: A developer asks their agent to figure out why a nightly export job is failing. The team’s rule for agents is simple: they work through read-only roles. But the developer also holds admin for on-call work, and both profiles sit in the same ~/.aws/config.

The agent hits AccessDenied when it tries to rerun the job, so it switches to the admin profile, assumes the role, and runs aws s3 rm against the production bucket to clear out the half-written export first.

The credential is valid. The developer may assume admin privileges, and the admin can delete S3 objects. AWS checks the signature, not who is holding the key, so as far as AWS knows, the developer did this. The violation is that an agent used a role that agents aren’t allowed to use. Unlike intent, you can check that on every request.

Who’s to blame? The developer who shouldn’t have handed the agent their credentials and let it auto-approve actions? The harness maker? That’s the wrong question. Again, agents should do more on their own, speaking both as a developer and as a manager of developers who work with AI.

Instead of assigning blame, we need to know where this deletion attempt can actually be stopped. There are several possible enforcement points, and each depends on what the control can see, what it can block, and whether the agent has another path to the same action.

We want autonomy with limits we can actually enforce

First, we need to scope the problem. There’s constant pressure to expand agentic access, and usually, the reason makes sense in isolation. A task gets stuck, for example. Or there’s a new integration for a service that holds some of the organizational context.

The result is always the same: the next task starts with more access than the previous one.

A second source of pressure is the agent itself. Agents look for new credentials when the ones they have are blocked, without asking the human in charge whether they can have that access.

This behavior is agnostic to the source of the instructions. A malicious prompt injection can cause overreach attempts, but so can mistaken assumptions during an authorized task.

Valid Key, Wrong Hands

AWS checks the signature, not who holds the key. When an agent grabs an admin profile, the request appears to come from the developer.

Token Security maps every agent to its owner, identities, and permissions, so your controls block actions outside the assigned task and allow the rest to run.

See where it stops

Be specific about what you are enforcing

Agents use tool calls to retrieve data or take action, so that’s where the enforcement matters most. That said, “control tool calls” is a very coarse rule. What if the tool is a shell that can run an SDK? It can also be a browser with an authenticated session. When you allow the tool, you also allow what’s behind it.

To properly understand what’s happening, we need the operation, its arguments, the account and resource accessed, and the identity in use. For data operations, we also need to understand what the output is and where this output is routed.

There’s a lot that goes into a proper decision about the validity of an action.

In most agent harnesses, local commands, file operations, skills, and delegation are also tool calls. They matter because of where they lead. For example, reading ~/.aws/config is how the agent finds an admin profile to use.

So enforcement at the tool call is critical, and so is deciding on the action inside it.

Where enforcement can happen

Some of the damage can be prevented through reasoning checks that assess a plan or action against the task at hand. But this is a probabilistic control, and some malicious instructions will pass through. There’s no good alternative to having proper mitigation controls that can stop an action.

Managed agent settings

Restricting tools, permissions, approval modes, and allowed integrations close to the agent.

Which clients honor the settings? Can a user, project, or agent override them? Model and effort settings can attempt to restrict access, but they’re behavioral choices by nature.

Broad when the app enforces it; little for behavioral settings

Low, except approval modes, which cost the most

Runtime hooks

Checking a supported operation before it runs, with context from the agent’s session.

Can the user or agent change or disable it? Does it cover alternate tools and subagents? Does it block before execution, including on errors and timeouts?

Per operation, for the events it sees

Low if automated, high if it asks a person

Gateways

Inspecting and blocking the requests routed through them, with a shared policy across connected agents.

Which traffic do they actually see: model requests, MCP tools, direct APIs, or network traffic? Can the agent reach the same system another way?

High for routed traffic, none for the rest

Low, since only the blocked call stops

Sandboxes

Limiting the files, processes, network routes, and credentials available to an agent.

What can the agent still do inside those limits? Are reachable services restricted to the right account and operations, or just an allowed domain?

Caps reach, not what happens inside

Medium, as tasks needing outside access break

Endpoint enforcement

Governing local agent use through endpoint software or the EDR that’s already deployed.

Can it stop a particular operation, or only a process or host? What is visible inside containers or VMs? Which hosted agents sit outside its reach?

Local agents only

High if it kills a process or host

Credential and target-service authorization

Reducing the authority granted to an agent and enforcing access at the system that holds the resource.

Are credentials specific to the agent and task? Are there alternate credentials? Can the service distinguish the agent from the person or shared account behind it?

High, wherever the request comes from

Medium, as narrow access stalls tasks and sends agents looking for more

API-based management

Changing agent settings, removing permissions, revoking credentials or sessions, and disabling access where platforms expose those actions.

When does the change take effect? What happens to existing sessions and cached tokens? Is it the prevention of future access or a response after the action?

Mostly after the action

None until it fires

These methods overlap and complement each other. A hook that calls a policy service and a gateway that consults an identity graph do a better job because the enforcement and decision points are in the locations best suited to them.

This table is far from exhaustive, and there are more considerations for some of these methods. Take managed settings, for example. Model and effort settings are behavioral in nature, but with some harnesses, permissions are much more than a prompt saying “please don’t flip out”.

In Claude Code, for example, permission rules are enforced by the application itself, and managed settings can restrict user overrides. In that case, control lives in a settings file, but that doesn’t make it less effective.

With hooks, the basic premise depends entirely on their placement in the execution chain. You can’t stop an event that’s already happened, and failure behavior varies by hook types and events.

Gateways have their own complexities, as they must base their authorization decisions on what they see. Look at AWS AgentCore, which demonstrates policy checks for connected tools. Note which calls it sees.

Sandboxes and scoped credentials solve different problems. Proper isolation can remove access capabilities, whereas a scoped credential limits what happens after access is granted. See Cloudflare’s sandbox authentication design, which shows a way to add credentials outside the sandbox, so the agent doesn’t need to possess the secret. Its operations, though, will still require authorization.

Apply the same policy to a hosted agent

Consider a support agent that’s asked to summarize open cases. Its connector uses a service account built for agents that edit and close cases. The agent decides that some cases look resolved and tries to close them.

Again, the capability comes with the credential, and the agent is within its grant. The place where the action is disallowed is the agent’s approved read-only policy.

To enforce it, the policy must be something a control can check, such as: “This agent may read cases, and any call that edits or closes one is denied, regardless of the service account it holds.”

If that agent lives in a hosted environment, an endpoint control on the employee’s laptop might have no opportunity to stop the action, since it’s all server-side.

Here, you’ll need cooperation from the SaaS platform’s own controls, rely on a tool gateway, or have a narrower service identity. And again, it depends on what the platform exposes and where the calls run.

Which brings us to coverage gaps in our controls. There’s no one perfect control here. Hooks aren’t better than endpoint enforcement, which isn’t better than a sandbox. It’s all about what you deploy and where.

Managed agent settings

Works if the client honors them

Only what the vendor exposes

Runtime hooks

On the events the runtime exposes

Only if the platform offers hooks

Sandboxes

Limit files, network, and credentials

The agent runs on the vendor’s servers

Endpoint enforcement

In reach

Out of reach

Gateways

Catch the admin calls, if AWS traffic routes through them

A tool gateway in front of the connector

Credential and target-service authorization

Keep admin out of the agent’s reach

A read-only service account for summaries

API-based management

Revoke the session after the fact

Disable the connector after the fact

Two methods can stop both of these examples before the action reaches the target: a gateway that sees the relevant traffic and credential scoping at the target. The other methods depend more heavily on where the agent runs and what the runtime exposes.

Above all, you want the agent to keep working within its approved policy. Read-only is not complete. It can still expose sensitive data, depending on where the output goes. Anthropic’s containment analysis explains why.

Choose what’s practical, then test the ways around your control

You’ll likely be deciding what’s practical based on your deployment. Managed settings require supported clients and central administration, for example. And you won’t have hooks without a runtime that exposes the right events. Gateways force you to route traffic through them, and endpoint controls mandate software installation and maintenance.

None of these is free. The controls introduce trade-offs, require maintenance, and must account for an ever-growing volume of AI-driven work.

Take the AWS example from the beginning of the article. Blocking the literal command is a start, but what if the same request comes through an SDK? With a different credential? Through delegation?

The security requirement to block deletion remains the same, but there are many different paths to it.

Then there’s the question of value, and you have to consider it as much as the control capabilities. How much delay do the controls add? How often does someone need to unblock a false positive? Are we reducing risk or introducing problems?

SACR’s ARISE report centers on runtime intervention, delegated authority, and action-level decisions. This should be brought down to a concrete request, a policy, and evidence of what happened when the agent tried it.

A short guide, based on where your agents run and what you control there:

Developer laptops and IDEs

The endpoint, agent settings, local credential files

Managed settings, plus hooks where the client supports them

A gateway for cloud API traffic, and endpoint enforcement for clients you can’t configure

Your own cloud, CI or containers

The runtime, network egress, and workload identity

A sandbox and per-agent workload credentials

A gateway on egress

SaaS platforms

Connector credentials and the platform’s admin settings

A narrow service identity per agent

Platform controls through the API, and a tool gateway in front of connectors, where the platform allows it

It probably isn’t one or the other. Instead, most organizations have all of the above, mixed together, with controls spread around different teams as well. Across these environments, the one common denominator is identity. So start with credential scoping at the target, as this dictates the reach wherever the agent is running. Then, add one policy source that every enforcement point reads from.

Do this, and you’re in a good starting position.

Where Token fits

At Token Security, we are building around the identity intelligence graph and the decisions it can support across the different enforcement points. Our current focus spans agent settings, gateways, endpoint enforcement, and APIs into identity and agent platforms.

The broader vision is to build or connect to the enforcement capability the customer needs. We don’t need to build everything ourselves.

The graph connects an agent to its owner, the identities it consumes, the permissions associated with those identities, and the resources it can access. That’s the kind of context that policy engines can use to decide what an agent may do, rather than assuming that the credential’s full authority is appropriate.

In a recent gateway demo, we demonstrated this exact idea: same developer, same session, and the agent’s call is denied on admin and allowed on readonly. Other controls can enforce read-only access, too.

Supplying the right identity context and applying the policy across the different places agents work is paramount, yet complex, as this process requires reliable attribution. If we cannot distinguish the agent’s request from a human using the same credentials, the graph doesn’t magically fix it.

What about baselines? Keep in mind that past behavior can only be used as an input, not the policy itself. The fact that an agent usually just reads data does not prove that a write is forbidden. That’s the job of the rule that pins agents to readonly. Missing or stale context also requires an explicit decision, as with hook timeouts.

I want an agent to complete the investigation or the support summary without requiring approval for every step. I also want to know exactly where an action outside that task will stop. That’s what I would ask of every enforcement method, including Token Security.

Check us out or book a demo if you’re also solving these AI security problems in your organization.

Sponsored and written by Token Security.

Originally published on BleepingComputer