Loading...
Loading...
Loading...
Blog Topic
Security and trust controls for tool-connected agents and MCP systems.
24 metadata-ranked posts in this topic
Ranked for relevance, freshness, and usefulness so readers can find the strongest Armalo posts inside this topic quickly.
MCP, A2A, ANP, and related protocols are moving faster than the trust models around them. The window to shape secure defaults is now.
When websites expose tools to browser agents, trust moves from page content to tool manifests, side-effect labels, and receipts.
Cross-agent work needs delegation receipts, counterparty trust checks, tool boundaries, and recertification after material change.
WebMCP is exciting because it gives browser agents structured tools. It is risky because side effects become easier to hide behind normal UI actions.
An MCP server you connect inherits your agent's authority. The blast radius of one bad server. The boundary patterns and a Trust Boundary Spec you can implement.
Most skills run with the agent's full credential set. They should run with capabilities scoped to the smallest task they need. The spec, the runtime work, and a manifest you can write today.
AI teams are accumulating permission debt every time an agent keeps access after its evidence, scope, owner, model, or tool boundary changes.
Trust oracles are public by design. That same publicness gives attackers a free reconnaissance layer. This is the security essay on read-side probing, and the controls that turn an oracle from a target map into a defensive asset.
Indirect prompt injection is usually framed as input filtering. For consequential agents, it is a planning and authority failure.
Agentic red teams should probe authority ladders, tool receipts, memory provenance, recursive promotions, and incident recovery.
The move toward OS-level agent workspaces changes the security conversation: the boundary is no longer just the model, it is the workspace around action.
Browser agents will not stay in harmless browsing mode. They need labels that distinguish reading, drafting, submitting, buying, exporting, and deleting.
Permission receipts make agent authority inspectable: who granted it, what evidence supported it, when it expires, and what narrows it.
Agentic incident response needs mission context, tool receipts, permission history, and recursive rollback in one command surface.
Zero trust for agents means every tool, memory, mission, and improvement request proves scope before authority moves.
Authority-security analysis of Agentic OS Mission Control, Armalo Agent recursive self improvement, governed autonomy, trust evidence, and real-world AI operations.
We scanned public agent skill catalogs and found 824 skills with adversarial behavior. Here is the taxonomy, the dominant patterns, and the audit checklist that catches them.
Before importing a new skill, diff its declared capabilities against your existing skill set. What's new? Why? Required permissions? A reviewer template you can use today.
A tool's provenance is a signed manifest binding source repo, build SHA, and signing key. Here is the audit pattern and a manifest schema you can adopt today.
Three sandbox modes for agent skills: process, container, microVM. When each is appropriate, how each fails, and a Sandbox Mode Selector you can run today.
Skill v1.2 was clean. v1.3 added a tool that talks to an attacker server. The trust scope of a skill must include version range. A Skill Version Pin Policy you can adopt.
Most agent runtimes import skills the way npm imported packages in 2015 — by name and by trust. The path forward is attestation at import time, with a checklist worth running.
MCP and tool protocols are making action easier. That makes tool governance the border-control layer for agents that touch data, money, code, and customer systems.
Agent identity matters, but identity without delegation receipts cannot prove who authorized what, for which scope, and with what recourse.
Safety Research
A public roadmap for calibrated workspace research across eight evidence gates: calibration, behavior, specificity, entanglement, sparse features, agent telemetry, self-monitoring, and adversarial robustness.