TL;DR: An AI agent shouldn't hold a long-lived key to the tools it calls. Give it a short-lived token, issued at runtime and limited to one tool. This came up on ten client calls in 2026. Most teams sit on rung one or two of a four-rung ladder, and token exchange (RFC 8693) sits at the top.
Key takeaways:
Ten client calls between May 27 and August 28, 2026, across eight organizations, asked how an AI agent should log in to the tools it calls.
Most agents today hold a static API key or a shared service account. That's rung one of four.
Fetching the key from a vault at runtime is rung two. The agent still holds a live secret while it works.
Token exchange, defined in RFC 8693 in January 2020, is rung four. It issues short-lived tokens for one service and keeps the user's identity attached.
The Model Context Protocol security guidance forbids passing a caller's token through to a downstream service, so each hop needs its own token.
How should an AI agent log in to the tools it calls?
An AI agent should log in with a short-lived token that a trusted issuer hands out at runtime. The token should work for one tool and one job. The agent shouldn't carry a long-lived key. A stolen key often works until someone notices. A stolen token works for minutes. Most teams aren't there yet, and the way up has four rungs.
Most agent security talk covers inbound access, meaning who can reach the agent. This post covers outbound access, meaning how the agent proves who it is to the CRM, the ticketing system, the cloud storage bucket, and the MCP server (Model Context Protocol, the standard connector between agents and tools). Outbound is where the keys live.
What are the four rungs of the credential ladder?
Rung one is a static key or shared service account stored in the agent's setup. Rung two is the same key, fetched from a vault each time the agent runs. Rung three is a key the gateway adds for the agent, so the agent never sees it. Rung four is token exchange, where no stored key exists at all.
Here's each rung in plain terms:
Rung one, the static key. The key sits in an environment file or a config. It never expires. Anything the agent runs can use it.
Rung two, the vaulted key. A secrets manager (a locked store for keys and passwords) holds the key. The agent checks it out when it starts. Rotation gets easier, but the agent holds the secret in memory while it works.
Rung three, the gateway-injected key. The agent calls a gateway, the checkpoint every tool call passes through. The gateway holds the key and adds it to the request. The agent never touches it.
Rung four, token exchange. The agent proves who it is to an identity provider and gets a short-lived token for one tool. No downstream key is stored anywhere.
Each rung leaves less standing power in the agent's hands.
How many client calls showed this pattern?
Ten client calls between May 27 and August 28, 2026 turned on one question: how does an agent get credentials for the tools it calls? They came from eight organizations. I counted them from my own call notes. Clients are described by sector here, and none are named.
The question showed up in different forms. A direct-sales company on May 27 covered an MCP server that called downstream apps with a key broader than the user who triggered the request. A healthcare technology company on May 29 discussed short-lived credentials brokered from a vault. A life insurer on August 5 and August 11 wanted one identity pattern across three cloud platforms. A home-security company on August 6 walked through retiring its stored keys. A travel company on August 17 and August 26 weighed one broad key against a key per client. A large retailer on August 19 wanted a standard for every new integration. A holding company on August 21 raised the same broad-key problem. A lending technology company's call summary on August 28 listed a token-exchange rollout as a next step.
The details vary a lot. The question doesn't. For the buyer side of the MCP half, see the post on why security teams ask about MCP first.
Why is a vaulted key only halfway up the ladder?
A vault protects the key while it's stored. It doesn't protect the key once the agent checks it out. At that point the agent holds a live secret, and anything that can read the agent's memory or logs can read the key. A vault makes a good rung two. It's a poor place to stop.
A travel company showed me the pattern on August 17, 2026. Its agent platform used a service account to get a cloud token, took on a cloud role, and then pulled static API keys for its tools out of a secrets manager. The team called it something they "want to get away from, but is what it is." I think that's where most programs stand today. The vault step beats keys in config files. It still ends with the agent holding keys to downstream systems.
Rung three is the answer when a tool only takes an API key. The gateway keeps the key and adds it to each request. The agent never sees it, so it can't leak it or hand it to a prompt that asks nicely.
What does token exchange do that a key can't?
Token exchange lets an agent trade proof of who it is for a new, short-lived token aimed at one service. RFC 8693, published in January 2020, defines how. The new token can name the person the agent acts for and the agent doing the acting. A stored key can do neither.
The RFC 8693 token exchange standard names the pieces. A subject token represents the party the request is made on behalf of. An optional actor token represents the party doing the acting. The request also names the audience (the target service) and the scope (the access wanted). The standard separates delegation, where the agent keeps its own identity while acting for someone, from impersonation, where the agent becomes indistinguishable from that person. For agents, delegation is the one you want, because the logs show both names.
Google's Workload Identity Federation documentation describes the same idea for cloud workloads. It replaces service account keys, which Google calls "powerful credentials," with federated identity and short-lived access tokens obtained through OAuth 2.0 token exchange. Workload identity federation (WIF) means the hosting platform proves the workload's identity to the identity provider, so no key gets stored.
Two of the ten calls, on August 5 and August 6, walked through the same migration path: inventory the keys, move to a managed identity, add federated credentials, then delete the secrets.
What does the MCP security guidance say about passing tokens along?
The MCP security guidance calls token passthrough an anti-pattern and says the authorization spec forbids it. An MCP server must not accept tokens that weren't issued to it. It also shouldn't forward a caller's token to a downstream service. I read the page on October 3, 2026. The URL points at the draft version of the spec.
The page lists why. Passing tokens along lets callers bypass controls such as rate limiting. It scrambles the audit trail, because downstream logs show a different identity than the one that acted. And it breaks trust boundaries between services. The same page describes the confused deputy problem, where a server with more power than its caller does something the caller was never allowed to do.
That last one is the May 27 call exactly. In my reading, the fix follows from the rule: if a server can't pass the user's token on, each hop needs a token of its own, with the right audience and a narrow scope. Token exchange is how you get one. You can read the source yourself at the MCP security best practices page.
Which rung should each tool get?
Pick the rung by what the tool's own login can do. Modern sign-in such as OAuth or SAML gets token exchange, API-key-only gets a gateway-held key, username-and-password gets a broker, and anything else gets a retirement plan. You don't pick one rung for the whole company. You pick one per tool.
A broker is a service that logs in for the agent and keeps the password away from it. For tools that can't do any of this, the answer moves to procurement. Put the limit in the contract and set a date to replace the tool.
I gave a large retailer this four-branch tree on August 19, 2026, as a standard for new integrations and a way to sort the old ones. It worked because it turns a security argument into a lookup. Nobody has to debate the rung for a tool. They read what its login accepts.
What should you do this month?
Count the keys your agents hold, then replace the one with the widest reach first. Then move one rung at a time. Don't try to reach rung four everywhere. Start with the agents that touch customer data or can write to production.
Inventory every key an agent holds, with its owner and what it can reach.
Replace the widest key first. At minimum, give every agent its own key instead of a shared one.
Put a gateway in front of every tool that only takes an API key.
Pick one high-risk flow and test token exchange on it before you promise it everywhere.
Keys are one of five small identity jobs every agent needs. The post on why the AI mandate arrives before identity is ready covers the other four.
Frequently asked questions
Is a vault enough for AI agent credentials?
A vault is a good rung two and a poor finish line. It protects the key at rest and makes rotation easier. Once the agent checks the key out, though, the agent holds a live secret. Pair the vault with a gateway that injects the key, or move to short-lived tokens, so the agent never holds a downstream key.
What's the difference between workload identity federation and token exchange?
Workload identity federation is the setup. The hosting platform proves a workload's identity to the identity provider, and you configure that trust once. Token exchange is the call that trades the proof for a short-lived token. Google's documentation says its federation uses OAuth 2.0 token exchange to get the access token.
Does an agent need its own identity before any of this works?
Yes. Every rung assumes the agent has a record, an owner, an exit date, and its own account. A shared key makes the logs useless, because you can't tell which agent used it. Give each agent its own identity first, then climb the ladder.
What about coding agents on developer laptops?
That's a real limit. On August 5, 2026, a life insurer's call covered it: desktop coding agents run as the human user, so there's no practical per-agent identity at that layer. Apply the ladder to server-side agents first. For laptops, rely on the user's own scope limits and logging.
The best key for an AI agent to hold is none. The ladder shows you how to get there one rung at a time, and the widest key goes first.
