Agent code is the new attack surface
Teams are shipping agents through SDKs, plugins, skills, hooks, MCP servers, and CI workflows. Baz now runs a dedicated security reviewer on every Change that touches agent code.
A year ago, "agent code" in most repositories meant one service that called a model API. Today it is everywhere. It lives in SDK imports scattered across backend services, in coding-agent settings files checked into the repo root, in plugins pulled from marketplaces, in skill folders, in hook scripts, in MCP server registrations, and in CI workflows that hand a model a token and a shell. Agents are being built by platform teams, by product teams, and increasingly by other agents.
Every one of those is a different way to build an agent, and every one of them is a different way to get agent security wrong.
Why this is a bigger problem than it looks
Traditional application security assumes a fairly stable shape: untrusted input arrives at a known boundary, flows through code a human wrote, and reaches a sink a scanner knows about. Agent code breaks almost every part of that assumption.
The boundary moved. An agent's "input" is not just what a user typed. It is every web page, ticket, log line, API response, and file the agent reads while doing its job. Any of it can carry instructions. A guardrail that only inspects the user's message is guarding the front door of a house with no walls.
The sink is a decision, not a function call. In a normal service, a dangerous operation is a specific API. In an agent, the dangerous operation is whatever the model decides to do with the tools it was given. Authority is granted ahead of time and exercised later, by a component nobody can fully predict.
Configuration is code, and nobody reviews it like code. A one-line change to a hooks file, a plugin manifest, or a permission setting can quietly change what an agent is allowed to do across an entire organization. These files rarely get the scrutiny an authentication handler would, even though they often matter more.
The supply chain got shorter and looser. Agents pull in plugins, skills, container images, and helper binaries, frequently at whatever version is newest at the moment they run. The time between "someone published this" and "this is running with our credentials" can be zero.
Velocity outran review. The same tools that make agents easy to build make them easy to change. Agent code is some of the fastest moving code in most repositories right now, and it is often written by people whose expertise is the product, not the threat model.
What it looks like in practice
When we looked across real agent code, the issues did not cluster around one exotic attack. They showed up in six recurring shapes.
Guardrails and hooks that do not actually guard. Hook logic that misses cases or is never armed. We saw input normalization that stripped quoting and let shell-wrapped commands slip past a denylist, a denylist that was not updated when a new skill was added, and agent spawn paths that skipped the guarded code path entirely. The guardrail existed. It just was not between the agent and the action.
PreToolUse hook. The hook strips quotes before matching, so the shell-wrapped producer slips past. A sibling .codex/hooks.json denylist was never updated for the new skill, and a TypeScript agent.spawn() path skips the guarded wrapper entirely.
Secrets and tokens inside skills and plugins. Skills that read an observability or infrastructure token and send it onward without checking who is asking. A simple keyword pass over one set of repositories surfaced dozens of candidate findings, with a large share rated high. Skills are small, easy to write, and easy to forget about, which is exactly why credentials end up in them.
GRAFANA_TOKEN from env. The script loads a Grafana service account token. A guard limiting it to Claude Code sessions exists in one entry point and is missing in the other.
Supply chain on autopilot. A marketplace plugin set to auto-update from a default branch. A credentialed scheduled job running a floating latest container tag. Release binaries downloaded and executed with no integrity check, a pattern that appeared more than twenty times in one scan. Each of these hands whoever controls the upstream a path into your environment.
autoUpdate: true, :latest. The plugin tracks main, the image tag is :latest, and install scripts download releases/latest and execute them without a SHA-256 or Sigstore check.
Untrusted context flowing into agents. Page content and replayed API responses inserted directly into an agent's system prompt, while the guardrail scanned only the user's own input. Notably, explicit prompt-injection language almost never appears in code. The risk is structural, not textual, so searching for "ignore previous instructions" finds nothing.
System prompt template. The content is interpolated straight into the system message. The NeMo-style input guardrail scans only the user's chat message, so instructions inside the page data pass through.
Authorization gaps on agent and tool surfaces. Tool server registration that skipped the permission checks the rest of the API enforced, and agents that could be invoked by anyone able to set an identity header. Agent endpoints often get built as internal plumbing and then quietly become reachable.
FastAPI MCP router. MCP tools are registered on a router without the Depends(require_permission) guard every other FastAPI route uses, and the agent endpoint trusts the header for identity.
Customer data leaving through skills. A messaging or ticketing skill that could publish customer content before it was scrubbed, and a shared temporary file reused between an approval step and the step that acted on it. The agent did what it was told, and the data went somewhere it should not have.
Slack skill, /tmp/plan.json. chat.postMessage is called before the PII scrubber, and the plan approved by a human is written to a shared /tmp file that the apply step rereads, so it can change in between.
None of these are hypothetical, and none of them would be reliably caught by a generic SAST rule or a dependency scanner. They live in the seams between configuration, prompts, tools, and ordinary code.
The coding agents are not catching this either
It is tempting to assume the coding agents themselves have this covered. Claude Code and Codex both ship permission modes, hooks, and sandboxing, and teams lean on them. But several of the issues above lived in exactly that layer, in code built on and around those tools.
The bypassed denylist was a Codex hooks configuration that had not been updated when a new skill was added. The quote-stripping gap was in a Claude Code plugin hook meant to block a class of shell commands. The auto-updating plugin came from a Claude Code plugin marketplace. The skill that sent an infrastructure token onward was supposed to be limited to Claude Code sessions, and the check enforcing that was not applied on every path.
In each case the agent behaved as configured. That is the problem. Claude Code and Codex enforce the hooks, permissions, and plugins you give them. They do not tell you that a hook misses a case, that a denylist is stale, that a plugin tracks a moving branch, or that a skill leaks a credential. Those are properties of the code and configuration your team writes, and they change in pull requests. When the agents wrote or edited these files themselves, nothing in their loop flagged the risk.
The guardrails built into coding agents are only as good as the configuration behind them, and that configuration needs its own review.
A reviewer built for agent code
Baz now runs a dedicated agent security reviewer on Changes that touch agent code.
It starts by recognizing when a Change is agent-related at all: coding-agent configuration, agent workflows in CI, model SDK usage, direct calls to model providers, and LLM dependencies. Changes that do not touch agents are left alone, so the reviewer adds signal where it matters instead of noise everywhere else.
When a Change qualifies, the reviewer investigates it the way a security engineer who specializes in agents would. It reads the surrounding code, follows callers, and works from a curated reference of agent-specific threat patterns: injection through retrieved context, over-broad tool authority, unguarded hooks, credential exposure, unpinned agent dependencies, and unauthenticated agent surfaces. It is required to ground each finding in evidence: untrusted data that reaches a dangerous operation, a safety check that is missing where it should exist, or a dangerous setting introduced directly in the Change.
Findings arrive as ordinary review comments on the changed lines, tagged with the relevant OWASP Top 10 for LLM Applications category, and they run alongside Baz's existing security reviewers, including Advanced Security, rather than replacing them.
Why now
Agent code is being written faster than any security team can learn to review it. The patterns above are already in production repositories, often introduced in a single small pull request by someone who reasonably believed a guardrail or a config file was handling it. The right place to catch them is the same place we catch everything else: in review, while the change is still small and the author still remembers why they wrote it.