Untrusted context
Web pages, tickets, documents, code, and messages can become instructions through indirect prompt injection.
The useful guidance is scattered across standards bodies, security teams, cloud vendors, and research papers. This guide turns it into a practical reading path and a control plan you can use.
Traditional software waits for inputs. Agents interpret context, choose tools, and take sequences of actions. The security question is no longer only “is the model safe?” It is “what can this identity see, decide, and do, and can we intervene?”
Web pages, tickets, documents, code, and messages can become instructions through indirect prompt injection.
Individually safe tools can form an unsafe chain when an agent combines access across systems.
Errors propagate faster when agents can write files, change infrastructure, send data, or approve workflows.
Endpoint controls see processes. SaaS logs see API calls. Neither alone explains the agent’s intent and full action path.
Selected for authority and practical value, not volume. Start with the first three, then choose the material closest to your role.
The clearest current view of the agent-security problem space: identity, authorization, data, monitoring, and human oversight.
A practical threat model for planning reviews, controls, and red-team exercises around autonomous systems.
Implementation guidance for the protocol layer connecting agents to tools, data, and privileged actions.
Concrete development patterns: constrain tool arguments, isolate execution, validate outputs, and design explicit approvals.
A useful browser-agent model covering untrusted page content, tool exposure, confirmations, and data boundaries.
Operational guidance for permissions, irreversible actions, tool chaining, and treating agent behavior as part of the trust boundary.
A shared language for adversary tactics and techniques against AI-enabled systems, useful for test plans and incident mapping.
Research on capability attestation, origin authentication, and trust propagation in multi-server agent environments.
A systematic analysis spanning skills, tools, and protocol ecosystems, with implications for architectural defenses.
A resource list is only useful if it changes the system. These are the control layers we think every organization deploying agents needs.
Inventory agents, models, MCP servers, tools, credentials, owners, and the environments where each can act.
Use task-scoped identity, least privilege, short-lived secrets, sandboxing, egress controls, and explicit action boundaries.
Capture prompts, tool calls, policy decisions, approvals, data movement, and outcomes in one reviewable timeline.
Block unsafe behavior in the moment, preserve evidence, revoke access, and turn incidents into stronger policy.
If any answer is “we don’t know,” you have found the next piece of work.
Discover agent activity, understand intent, and enforce policy across the full automation lifecycle.
Request a demo →