Skip to main content

Should Agent Guardrails Live in the Model or the Infrastructure?

Alex Raeburn
Alex RaeburnMarketing Manager
12 min read
Should Agent Guardrails Live in the Model or the Infrastructure?

The real question: where should an agent be stopped?

The cleanest way to think about agent guardrails is to stop treating them like a language problem. The model can talk itself into almost anything if the context gets noisy enough. The real question is simpler: where does an agent’s intent become a real action that affects files, sockets, databases, or credentials?

That boundary matters because once an agent can write to disk, open a network connection, or read a secret from the environment, the discussion stops being hypothetical. A bad answer in a chat window is one thing. A bad answer that creates a file, sends a request, or pulls a token out of a vault is a different species of problem. The moment the agent crosses into the host system, the blast radius changes.

The safest guardrail is the one the model cannot talk its way around.

That is the core shift. If enforcement lives inside the model’s own control loop, then safety depends on the same thing that produced the action in the first place: the model’s current interpretation of the prompt, the conversation, and whatever tool schema it was handed. That can work for narrow cases. It can also drift the first time the agent gets creative, the task changes shape, or a tool call lands in an unexpected branch of code. The model may still mean well. The system may still lose the argument.

A better mental model is to assume the agent will eventually do something surprising, inconvenient, or plain wrong. That is not pessimism. It is engineering. People make that assumption about human operators all the time, and agents deserve the same treatment, maybe more so because they can repeat a mistake at machine speed. If an agent decides to “just check one thing” in a directory it shouldn’t touch, or to “temporarily” use a credential it never needed, the system should already know how to say no.

This is where the location of the guardrail starts to matter more than its wording. In-model steering tries to influence what the agent wants to do. Boundary-based enforcement controls what the environment will actually allow. One lives in the text stream and can be nudged by prompt drift or tool misuse. The other sits at the point where an action becomes real, which is a much less charming place for a surprise to land.

So the practical question is not whether the model sounds careful. It’s whether the surrounding system can stop the wrong action even when the model is being clever, confused, or both. That is the lens to keep in mind as the rest of the design falls into place.

Why prompt-only guardrails break down in production

Why prompt-only guardrails break down in production

A prompt can shape behavior. It can’t promise behavior.

That distinction matters once an agent stops being a chat toy and starts touching real systems. In a sandboxed demo, a well-written system prompt and a couple of wrappers can feel reassuring. The agent agrees to the rules, remembers the rules, and even repeats the rules back to you with a straight face. Then you let it write a file, call an API, open a socket, or read from a credential store, and the whole arrangement gets less charming very quickly.

The problem is that model output is probabilistic. The same instruction does not always lead to the same next token, and that means policy embedded in a prompt is always partly a suggestion. A small context change can shift the model toward a different interpretation. A longer tool trace can crowd out an earlier instruction. A vague user request can push the agent into a workaround you didn’t think to forbid. None of this requires malicious intent. It just takes a model doing what models do: continuing the sequence in the direction that looks most likely from the context it has.

If your safety rule only lives in the text the model reads, it can be outvoted by the next few tokens.

Wrapper libraries help, but they usually help in a limited way. They can standardize tool calls, strip out obviously bad inputs, and keep the conversation tidy. That’s useful. It is not the same as stopping side effects. Once execution has been delegated, a wrapper can say, “Please don’t do that,” but the underlying action may already be one function call away. If the agent can retry, branch, or reformulate its request, the wrapper often becomes a polite speed bump rather than a hard stop.

This is where the model vs infrastructure split starts to matter in practice. The model can be told not to fetch secrets. The infrastructure decides whether the secret store is reachable at all. The model can be told not to send data to unknown domains. The runtime decides whether that network path exists. Those are different classes of control, and they fail differently.

A common failure mode is simple rerouting. The model gets blocked on one tool, so it picks another. If it can’t read a file directly, it may ask a different service that exposes the same data in a roundabout way. If one endpoint rejects a request, it may retry with slightly different parameters until it finds a path that works. None of that looks dramatic in isolation. It just looks like persistence. In aggregate, it can defeat the original intent of the guardrail without ever triggering the exact wording you tried to ban.

Another failure mode is inference. Even when an agent cannot directly access a secret, it may still infer parts of it from context that was left lying around. Error messages, logs, neighboring config, environment names, and tool descriptions can leak more than teams expect. A prompt can say “don’t reveal secrets” all day long. If the agent can read enough surrounding material, the issue is no longer disclosure in the conversational sense. It’s exposure through accessible state. That distinction gets missed surprisingly often in production systems.

The deeper the privileges, the shakier linguistic guardrails become. An agent with read-only access is easier to reason about than one that can write files, mutate database records, or call payment or identity APIs. Once the agent can commit real actions, a prompt is just one input among many. It may still reduce mistakes. It may still steer the average case. It cannot be treated as the enforcement layer.

That’s why cloud guidance around agent safety keeps drifting toward runtime controls rather than pure text instructions. Microsoft’s guidance on hosted agent guardrails and AWS’s Agentic AI security guidance both point in that direction: the place where policy bites is not the friendly paragraph in the system prompt, but the place where requests become real system calls. The wording differs, but the practical lesson is the same.

So prompt-only guardrails are still worth using. They can keep the agent on the rails when the stakes are low, and they can reduce how often the runtime has to say no. They’re just not enough once the agent has enough reach to cause side effects outside the conversation. At that point, you want the text to guide intent and the platform to enforce limits. The next step is figuring out what those limits look like when you stop trusting the model to self-police and start putting actual fences around its tools.

Put policy at the boundary: sandboxes, allowlists, and least privilege

Once the prompt stops being the gatekeeper, the runtime has to do the annoying job. That means the agent can ask, but the system decides. If the model wants to write a file, spawn a subprocess, or call out to a host on the internet, those actions should pass through something that can say no even when the model sounds very sure of itself. A sandboxed runtime is the cleanest place to start. Containers help, but they’re only the first rung. In production, teams usually want a tighter setup: read-only filesystems where possible, no host mounts, dropped Linux capabilities, seccomp or similar syscall filters, and outbound connectivity that is explicitly routed rather than assumed. If the agent doesn’t need access to the host, don’t give it access to the host. If it doesn’t need raw internet access, don’t let it improvise its own.

A guardrail only matters when the thing being guarded can’t step around it.

Put policy at the boundary: sandboxes, allowlists, and least privilege

That same logic applies to tool allowlists. A useful agent setup does not say, “the model may use tools.” It says, “the model may use these specific tools, against these exact domains, APIs, and file paths, under these exact conditions.” That’s a much smaller surface area, which is the point. A scraper agent might need to fetch from a narrow set of domains and write results into one directory. A data workflow might need access to a queue, an object store bucket, and nothing else. Even small exceptions matter, so keep them explicit. Tool allowlists are boring in the best way. They turn an open-ended request like “go fetch what you need” into a closed list of allowed moves. OpenAI’s agents guardrails guidance and Google’s Agents API docs both point in this direction: define the boundary, then make the runtime enforce it instead of hoping the model behaves.

Least privilege needs the same treatment. Credentials should be scoped to the task, not handed out like party favors. An agent that only needs to read a public API should not see a secret that can delete production resources. Environment variables should be trimmed down to the minimum set required for the current job, and sensitive values should stay out of the general process environment whenever possible. Short-lived tokens beat long-lived ones. Per-task credentials beat shared service accounts. If a job only needs read access to a bucket, give it read access and stop there. If it only needs to call one internal API, scope the token to that API and that API alone. The ugly surprise here is that many agent failures don’t look like an attack at first. They look like curiosity. The model sees a variable, a mounted directory, or an unguarded secret and treats it as fair game because the surrounding system made it visible.

The safest default is fail-closed. Dangerous actions should be explicit capabilities, not ambient permissions that happen to exist because nobody removed them yet. File writes beyond a sandboxed working directory should require a deliberate grant. Network egress control should default to denied unless a destination has been approved. Process execution should require the tool itself, not a generic shell with too much room to wander. That pattern sounds strict because it is strict. It also makes incidents easier to reason about. If a job failed, you can inspect the missing capability instead of wondering which one of a hundred loose permissions the agent tripped over.

In practice, this approach changes the shape of the whole system. The model becomes a planner and request generator. The infrastructure becomes the part that checks whether a request is allowed to become a real side effect. That separation is the whole trick. It’s also where the next layer starts to matter, because even a locked-down sandbox still leaves a trail of outbound calls and tool usage that deserves inspection.

Watching the wire: network controls, fingerprinting, and monitors

Once you’ve put the obvious policy gates around the agent, the next question is less glamorous and more useful: what can still get out over the wire? That’s where a lot of production AI security work ends up. Prompts can say “don’t touch secrets,” but a network monitor, syscall trace, or kernel-level agent can tell you whether the process tried anyway. Same for file writes and outbound requests. If the model is running in a container, VM, or locked-down host, you can watch what it actually does without asking it politely.

If the model can improvise, the monitor has to be boring, stubborn, and hard to talk over.

That sounds a little rude, but rude systems are often the ones you want in production. A hardware-backed or OS-level monitor can observe traffic outside the model’s control loop. Think eBPF on Linux, auditd, seccomp logs, hypervisor controls, or a network appliance that sees the packets before they leave the machine. These tools don’t care whether the agent had a clever rationale. They see a connection attempt to api.stripe.com, a DNS lookup for an odd domain, or a file read from a path that shouldn’t have been touched. That matters when an agent has access to credentials, external APIs, or internal services that were meant to stay out of reach.

Egress filtering belongs in the same bucket. If the agent only needs to call three services, there’s no good reason to let it spray requests anywhere else. Destination allowlists keep outbound traffic inside known-safe lanes, and per-tool routing makes the policy easier to reason about. A browser tool can go through one proxy with one set of destinations. A database tool can talk only to the database. A fetch tool can reach a narrow set of domains and nothing else. When each tool has its own path, it becomes much easier to spot weird behavior, and much harder for a rogue request to blend in with normal traffic.

That’s also why logs need more structure than a pile of raw events. Record tool calls, destination hosts, HTTP methods, file paths, and credential use in a way that a human can scan later and a detector can classify now. Was the agent reading a config file? Writing a temp file? Reaching an unexpected country-specific endpoint? Repeating a failed request with different headers? Those patterns tell you whether the system is behaving or drifting. A log line that says “tool executed” is barely useful. A log entry that says “browser tool requested PCODE_17 to a non-allowlisted host after reading PCODE_18” gives security teams something they can act on.

The better setups treat monitoring as an enforcement layer, not just a receipt printer. That distinction matters. If a monitor only reports bad behavior after the fact, you still get the leak. If it can stop the process, tear down the connection, or quarantine the session, the system gets a second chance to stay within policy. That’s especially useful when an agent handles tokens, API keys, or data that should never leave the approved path. Even a brief exposure can be enough to create a mess you’d rather not clean up at 2 a.m.

There’s a practical side to this in the tooling ecosystem too. If you’re already using guardrail hooks at the tool layer, the policy should be visible where the action happens. OpenAI’s tool guardrails reference is a decent example of putting checks around tool execution rather than trusting the model to self-police. Likewise, agent gateways such as Google’s agent gateway setup guide point in the same direction: watch and control the traffic at the boundary, not just in the prompt.

The pattern is simple enough to say and annoying enough to build: assume the agent will eventually try something odd, then make the wire tell you about it before the odd thing becomes a leak. That’s the part of least privilege for agents that doesn’t fit neatly into a prompt template, but it’s usually the part that saves you.

The production pattern: model proposes, infrastructure disposes

After you’ve watched the wire, the rest of the design starts to look less mysterious. The model can draft a plan, pick a tool, and ask for permission. It can even sound very confident about why it needs access to a file, a socket, or a credential it absolutely should not touch. That’s fine. Confidence is cheap. Enforcement is the expensive part, and it belongs somewhere else.

The clean mental model is simple: the model proposes, infrastructure disposes. In practice, that means prompts can shape behavior, but they don’t get to define the final outcome. The runtime decides whether a file write happens. The tool broker decides whether the request is well formed and allowed. The network boundary decides whether traffic leaves the box at all. If one layer misses something, the next one still has a chance to say no.

A polite agent with real permissions is still a risky system if nothing outside the model can stop it.

That layered setup sounds almost boring, which is usually a good sign in production. A prompt can steer the agent toward approved tools and sane defaults. A runtime sandbox can block file paths, process launches, and arbitrary shell access. A broker can enforce allowlists for tools, domains, and API methods. Egress controls can pin outbound traffic to known destinations. Credentials can be scoped so tightly that even a curious agent only sees what it needs for the current job. None of these pieces has to be perfect on its own. They just need to fail closed when the agent gets clever.

The real test is whether your system still behaves under bad assumptions. Don’t only check the happy path where the agent politely asks for the right API call and uses the expected parameters. Try the annoying stuff. Feed it a prompt that nudges it toward a forbidden host. Give it a malformed tool request and see whether the broker rejects it cleanly. Rotate a credential and confirm the agent can’t keep using the old one. Ask it to repeat a request in a slightly different form and see whether policy survives the detour. If your guardrails vanish the moment the model becomes inventive, they were never guardrails. They were suggestions with a nice font.

That kind of testing also catches the accidental failures that show up in real systems. A retry loop can turn a harmless request into a noisy one. A fallback path can skip validation. A tool that was meant for internal use can end up exposed to a wider agent than anyone planned. These are the bugs that make teams mutter at logs on a Friday afternoon.

The takeaway is unglamorous, which is probably why it works. Assume the agent will eventually do something clever and wrong. Then build the system so the answer is still no, even when the request comes wrapped in perfect grammar and a very persuasive explanation.

Newsletter

Stay in the loop

Join our newsletter and get resources, curated content, and inspiration delivered straight to your inbox.