Creative Input Is a Security Boundary
The odd part of the recent campaign was not the end goal. Crypto-mining, botnet behavior, and the rest of the usual malware baggage have been seen before. What made people stop and look twice was the delivery format: prompts written like poems, with line breaks and rhyme doing some of the work that plain commands usually do.
That detail matters because an internal agent is not a polite chatbot sitting in a browser tab, waiting to chat about office trivia. It can read files, query internal systems, call tools, send messages, update records, and chain those actions together. Once you give a model that kind of reach, prompt text stops being harmless prose and starts looking like untrusted input feeding a privileged workflow.
If a prompt can influence a tool call, it belongs in the same mental category as any other input that crosses a trust boundary.
That’s the clean way to think about it. A poem, a rhyme, a block of code-looking text, or a strangely formatted note can all carry instructions that steer behavior. The format may feel playful, but the model does not care whether the instruction arrives in a stern memo or a sonnet. If the content changes what the agent decides to do, then the content has security impact.
This is where teams sometimes get into trouble with LLM prompt injection. They treat prompts as conversation, when the better model is to treat them as data with side effects. The agent is part of a system that can touch production resources, customer records, internal documents, or workflow automation. That means the prompt is not just text to be summarized or answered. It can become part of an execution path.
The practical mistake is assuming the model’s own safety behavior will hold the line on its own. Sometimes it will. Sometimes it won’t. Models are inconsistent under pressure, especially when the input is crafted to be slippery, oddly shaped, or just plain annoying. Security teams already know this pattern from other places: one control is never enough when the input is hostile or simply weird in ways the system didn’t expect.
So the stance here is simple. Treat creative prompt text as untrusted input. Treat internal agents as software that can be steered into doing real work. Build around that assumption instead of hoping style alone makes an attack look harmless. The next question, then, is how the formatting itself can slip past quick review without looking obviously malicious.

How Poems Sneak Past Basic Safeguards
A plain-English malicious prompt usually trips a few alarms in a human reader’s head. A poem does less of that. Line breaks change the pace. Rhyme makes the text feel performative. Metaphor creates a little room for deniability. So a request that would look blunt and suspicious in a normal paragraph can arrive wrapped in a stanza and read, at first glance, like odd creative writing rather than a command.
That matters because a lot of review processes still depend on fast visual scanning. Logs get skimmed. Prompt histories get copied into tickets. Engineers glance at a blob of text and try to answer one question quickly: does this look weird enough to inspect more closely? Creative formatting lowers the odds of a quick yes. A block of verse can break the obvious shape of an instruction, and the eye tends to follow the rhythm instead of the payload hidden inside it.
A poem can hide a request in plain sight, but it still behaves like a request.
The trick is often boringly social. It does not require a model jailbreak in the classic sense. It relies on language shaping how the system and the humans around it interpret the text. If the prompt sounds playful, literary, or abstract, the agent may be more willing to continue parsing it as content rather than treating it as a control attempt. That is the basic idea behind creative prompt abuse and, in the more specific case of verse, poetic prompt injection. The attacker is not cracking the model open. They are nudging it.
That nudge can work in a few ways. First, rhythm and rhyme can bury imperative phrases inside softer-sounding lines. A line that reads like a stanza often feels less direct than the same instruction in a numbered list. Second, metaphor can act as camouflage. A request to reveal hidden data might be framed as a story about “showing what sits behind the curtain,” which sounds literary until you read it twice. Third, line breaks create separation. Humans are decent at reassembling intent across lines, but quick scanners often are not. The instruction survives, while the structure makes it harder to spot.
Unusual formatting also plays poorly with copy-paste habits. People copy a prompt into chat, into a ticket, or into a review tool, then skim the output line by line. A poem with short lines and odd indentation can make a malicious instruction look like a stray phrase, especially if the dangerous bit is split across several lines. You see this in a lot of prompt injection guidance from OWASP: the abuse usually works by changing how text is interpreted, not by exploiting some magic bug in the model’s internals.
Models can be nudged too. A creative prompt may give the system a reason to reinterpret what it should do. If the surrounding text sounds artistic or hypothetical, the model may soften its refusal behavior, or at least spend more tokens trying to reconcile the poem’s surface meaning with the hidden instruction. That is awkward, but it is not surprising. These systems are pattern matchers with a lot of fluency and not much common sense. If the input strongly frames itself as style, the model may treat the underlying request as part of the style.
This is why the defense question is less about “Can the model read poetry?” and more about “What does our intake logic do with weird text?” The best recent framing I’ve seen comes from the security side, where NIST’s work on agent hijacking evaluations pushes teams to test how models behave when the prompt is deliberately bent, disguised, or overloaded. That includes verse, oblique phrasing, and other formats that make intent harder to spot on first read.
The practical lesson is pretty plain. Treat the format as part of the attack surface. A poem can carry the same bad instruction as a blunt paragraph, and sometimes it carries it better because people lower their guard. If your review path only looks for rude wording or obvious command phrases, creative prompt abuse will slip through with a smile and a rhyme.
When an Agent Has Tools, the Risk Changes
A plain chatbot can say something sloppy and still remain mostly harmless. An internal AI agent with tools is a different creature. If it can read files, query a database, send a message, call an internal API, or kick off a workflow, then a prompt is no longer just text. It becomes input to a system that can take action.
That’s the part people sometimes miss when they first see creative prompt abuse. The poem is not dangerous because it is poetic. It becomes dangerous when the model behind it has enough reach to do something useful for an attacker. A weirdly formatted request that might look amusing in a chat window can turn into a file read, a search across internal docs, or a message posted to a shared channel if the agent is wired into those tools.
Once an agent can act, the prompt stops being conversation and starts being control input.
File access is usually where teams get a rude awakening. An agent that can open arbitrary documents may be able to pull configuration files, internal notes, API keys, or incident logs. If those files are indexed in retrieval, the model may fetch more than the user expected, especially if the prompt nudges it toward “relevant” context. Memory can make this messier. A malicious instruction that gets stored or echoed back later can shape a second task even after the original request has faded from view.

Internal APIs widen the blast radius even faster. A model that can create tickets, approve access, update customer records, or trigger deployments is no longer just summarizing data. It is operating parts of the business. Message sending tools create another path for abuse. A compromised agent can relay secrets, send phishing-style internal messages, or simply flood a channel until people tune it out. Database queries are obvious trouble too. Even read-only access can expose far more than a human reviewer would want a model to see, especially when joins, exports, and free-form query generation are on the table.
Workflow automation ties these pieces together. One tool call fetches a record. The next writes to a queue. Another sends a Slack message. A fourth updates a CRM field. Individually, each step looks routine. Together, they form an action chain that can move a single prompt from “please summarize this” to “please do something irreversible.” That is where chained tool calls matter. The model may not need to break anything technical. It just needs enough latitude to keep going.
Overbroad credentials make all of this worse in a hurry. A backend service account with wide read and write permissions turns every prompt into a potential admin request. Default access patterns are especially awkward here, because teams often give internal agents more scope than they would give a human user. The logic sounds innocent at first: the agent needs broad access so it can help. In practice, that means the model inherits privileges it cannot judge safely. A service account that can see everything and change a lot of things is a fine setup for automation. It is a lousy setup for untrusted input.
This is where internal trust assumptions fall apart. People often treat internal AI agents as if they are trusted coworkers in software form. They are not. They process user input, retrieve context, and choose actions on behalf of a person or service account. If an attacker shapes that input carefully, the agent may comply with the wrong instruction while believing it is being helpful. Microsoft’s guidance on prompt injection treats this as an attack technique rather than a quirky edge case, which is the right mental model, and OWASP’s LLM Prompt Injection Prevention Cheat Sheet gives a decent checklist for thinking about the problem in practical terms. Their advice lines up with the ugly reality here: the agent is part parser, part planner, part actor, so the attack surface is bigger than the chat box.
The implication is pretty simple. If a model can only answer questions, a strange prompt is annoying. If it can touch files, APIs, and workflows, that same prompt may become an execution path. That is why prompt normalization, permission scope, and tool boundaries matter so much for internal AI agents. The next step is figuring out which controls actually hold up when the model is allowed to reach into production systems without getting cute about it.
Layered Defenses That Hold Up in Production
Once an agent can read internal data, call tools, or send anything on its own, you have to stop treating prompt text like a harmless chat message. Creative prompt abuse sits in the same family as prompt injection, the term NIST uses for inputs that try to steer a model away from its intended behavior. The format can be a poem, a code block, a fake policy memo, or a neatly wrapped set of bullet points. The shape changes. The job for defenders does not.
Treat prompts like code before the model sees them. That means normalizing text so the weird stuff is plain to your filters. Collapse repeated whitespace. Canonicalize Unicode. Strip zero-width characters and odd formatting where your application can afford it. If your pipeline accepts Markdown, HTML, or rich text, decide what you actually need and reject the rest. A lot of abuse hides in formatting tricks, not in exotic model behavior. Policy checks should run at this layer too, before the prompt reaches the model. If the text contains attempts to override instructions, invent new roles, or smuggle in conflicting directives, you want that caught early, not after the agent has already started planning.
A prompt is not a suggestion box when the model can touch systems, send messages, or change records.
The next control is boring in the best possible way: narrow tool permissions. Give the agent only the calls it truly needs, then trim again. If it only needs read access to one internal search endpoint, do not hand it a general-purpose service account. If it only needs to draft an email, do not let it send one without a second step. Least privilege sounds obvious until an internal agent is built with the same permissions as the team’s most trusted admin account, which happens more often than anyone likes to admit. Separate credentials per tool. Use allowlists, not broad API access. Keep scopes short-lived where possible.
For higher-risk actions, put a human or a separate execution path in the loop. Payment approvals, account changes, destructive updates, external sends, and credential access should not happen just because the model reached a confident conclusion. A review step can be manual, policy-based, or routed through a smaller service that checks the request against business rules before anything happens. The point is not to slow everything down. The point is to make sure the model cannot turn a weird prompt into an irreversible side effect by itself.
Output checks matter just as much as input checks. A model can look calm and still hand back something that should never be executed. Validate outputs before they trigger side effects. If the agent produces a database query, inspect it. If it produces JSON for a downstream job, parse and verify the fields against a schema. If it produces a tool plan, compare that plan to what the caller requested. This is where output monitoring earns its keep. You’re looking for odd tool sequences, repeated retries against the same endpoint, sudden changes in language, or a plan that jumps from harmless lookup to privileged action without a good reason. Those patterns are often easier to spot in logs than in the moment the prompt is being read.
Regression tests close the gap between “seems safe” and “survives contact with reality.” Build a small adversarial set and keep it around. Include poems, oddly formatted prompts, oblique instructions, nested quotes, fake system messages, and long blocks of text that bury the actual ask near the end. Add examples that try to split one instruction across lines or disguise intent in a polite request. The goal is not to win a poetry contest. The goal is to see whether the agent still obeys your policy checks, respects tool restrictions, and refuses to take forbidden actions. Teams often test happy paths until the workflow looks solid, then learn later that a single stylized prompt can send the agent off script.
Logging and alerting tie the whole thing together. If a prompt trips a policy rule, log it. If an agent asks for the same tool three times in a row, log it. If a request produces an approval denial followed by a retry with different wording, log that too. Keep enough context to reconstruct what happened without storing more sensitive data than you need. Then alert on the patterns that matter, not every minor refusal. A good alert tells you when an agent starts behaving like it’s been nudged by a crafted input, and it does so quickly enough that you can cut off the blast radius before the workflow spreads.
Anthropic’s work on trustworthy agents makes the same general point from a different angle: agents need guardrails around what they can read, decide, and do. That’s the practical posture here. Build the checks at the edges, keep permissions tight, and assume the prompt text is untrusted until proven otherwise.
Build for Abuse, Not Trust
The cleanest lesson here is a slightly annoying one: weird formatting does not mean safe input. A poem, a rhyme, a block of oddly spaced lines, or a prompt that reads like a school assignment can still carry the same malicious intent as blunt prose. If anything, the creative wrapper can make a bad request easier to miss during review. So the right default is simple. Treat prompt text as untrusted, no matter how playful, literate, or harmless it looks at first glance.
Novel wording is not a safety signal. It is just another way to smuggle the same request past a tired reviewer or an overconfident system.
That mindset matters even more once an internal agent can do real work. A plain chatbot can annoy you. A tool-using agent can read files, query systems, send messages, call APIs, and chain those actions together before anyone notices the drift. At that point, the question is no longer whether the model “understood” the prompt. The question is what it can reach if it takes the bait. That’s where least privilege stops sounding like a textbook line and starts paying rent. Give the agent only the permissions it needs for the job it is actually doing, not the ones it might need on a heroic day in an imaginary future.
Teams often look for one control that will save them from prompt abuse. That usually ends badly. Model-side refusal checks help. Input normalization helps. Output validation helps. Sandboxing LLM agents helps too. None of them should carry the whole load. A safer setup comes from stacking boring controls that each limit a different slice of damage. If one layer misses the trick, the next one catches part of it, and the blast radius stays smaller than the attacker hoped.
That also means every new tool, permission, and workflow deserves a suspicious read before it goes live. If an agent can open a ticket, modify a record, fetch internal data, or post to Slack, assume someone will try to steer it with language that sounds less like a command and more like a puzzle. Review the path from prompt to action, not just the model output. Ask what happens if the request is disguised, repeated, split across messages, or tucked inside creative formatting. If the answer makes you wince, trim the scope before shipping.
The habit that lasts is pretty boring, which is usually a good sign in security work. Keep adding adversarial tests. Keep tightening permissions. Keep checking logs for shifts in tool use, phrasing, retries, and odd sequences that didn’t show up last month. Models change. Prompts change. Attackers certainly change. Your guardrails should change with them. That way, when the next poetic payload shows up, it lands in a system that was built to expect trouble instead of admire the formatting.




