As organizations aggressively shift from using static Large Language Model (LLM) chatbots to fully dynamic, autonomous AI agents, traditional compliance and governance frameworks are hitting a breaking point. By 2027, Gartner predicts 40% of enterprises will demote or decommission their autonomous AI agents – not because the technology failed, but because their governance did. The culprit, according to Gartner Senior Director Analyst Shiva Varma, is “binary governance”: treating agents as either completely locked down or fully trusted, with nothing in between.
The key takeaway is this: AI agents cannot be governed uniformly. Right now, most organizations are oscillating between two deeply flawed extremes: a total corporate ban or a Wild Wild West. But the truth is, most organizations today are oscillating between exactly those two extremes. Even when IT enforces a total ban on outside AI tools, Shadow AI still happens. Developers naturally want to build and solve problems; if governance feels overly bureaucratic, they’ll bypass controls and use untested tools to keep things moving. A total ban doesn’t eliminate risk – it just pushes it underground.
The real problem isn’t the use of AI agents – it’s applying static, industrial-era controls to machine-speed software and expecting that to work.

Governing Agentic Software Development Environments
Real control means effectively managing all the software elements that go into the agent. Many teams mistakenly view these agents as “magic,” overlooking the standard elements that influence their functionality and security.
Every autonomous AI agent is built from a recognizable stack: a foundation model, MCP servers (Model Context Protocol) that connect it to company data and external systems, plugins that extend its capabilities, and skill files that instruct how it behaves. Each component is, at its core, a software artifact – and each carries its own supply chain risk. Thus, IT leaders need to keep in mind:
- AI agents aren’t magic: They rely on a standard setup that includes a Model, MCPs (Model Context Protocol servers), tools, plugins, and skills.
- MCP servers can be risky: They connect agents to company data and can be downloaded like any other software. If they’re not vetted, they could let hackers into your system. For example, recent JFrog Security Research highlights these exact software supply chain risks within mcp-run-python, including aServer-Side Request Forgery (SSRF) flaw (CVE-2026-25904) and alack of isolation vulnerability (CVE-2026-25905) that can lead to a complete MCP takeover.
- Plugins can add more risk: Different agents need different plugins (e.g., an OpenAI plugin won’t work in Microsoft Copilot), meaning companies often juggle multiple variations where a single bad plugin could steal sensitive information.
- Skills define agent behavior: These instruction files tell the agent how to work. Because developers pull them from community sources, a corrupted skill can quietly change how an agent operates.
- Visibility is key: Many teams don’t know these issues exist, leaving their agents highly vulnerable. Organizations are still managing software dependencies based on trust instead of thorough checks.
The Role of a System of Record in Proportional AI Governance
The race to deploy enterprise AI agents cannot be won with a binary mindset. Imposing monolithic, heavy-handed approval chains across every workflow stifles innovation, while leaving armies of agents unmonitored invites systemic disaster.
To safely scale AI, organizations must adopt a layered approach to governing and securing software components, including packages, binaries, MCPs, skills, and dependencies. By establishing a system of record, enterprises can curate incoming components, scan everything prior to use, enforce policies at the boundary, and maintain the ability to roll back instantly. So when the next vulnerability is disclosed, the answer to “are we exposed?” comes in minutes, not days.
However, AI agents introduce a fundamentally different problem: they present a wider and significantly faster-moving attack surface. Because high-autonomy agents are capable of making hundreds of decisions a minute, traditional passive registries and manual approval gates simply cannot scale to secure them. These traditional methods might work for low-autonomy tasks, but AI agents require dynamic policy enforcement at runtime, rather than just at intake.
High-autonomy agents are a different problem. When an agent can make hundreds of decisions per minute – chaining tool calls, accessing data, executing code – static intake policies cannot scale. These agents require dynamic enforcement at runtime: the ability to detect anomalous behavior the moment it occurs and intervene before the damage spreads. Think of it as a circuit breaker: revoking access to a compromised skill, disconnecting an untrusted data context, or blocking a specific tool call the moment something goes wrong – rather than waiting for a manual review process that machine-speed agents will always outpace.
Moving Beyond “One-Size-Fits-All”
As enterprises move from static chatbots to highly autonomous agents, rigid uniform governance is a guaranteed path to failure. The window for getting ahead of this is closing and governance debt compounds as agent deployment accelerates. Thus, here are four steps security and engineering leaders can take now to help shore up their enterprise systems:
- Take inventory of what’s already running. Most enterprises can’t fully account for the AI agents in their organization, let alone the MCP servers, plugins, and skill files powering them. Before you can govern anything, you need to know what exists.
- Treat AI components as software artifacts. MCP servers, skill files, and plugins aren’t configuration — they’re code, with the same supply chain risks as any open-source dependency. Apply the same vetting, scanning, and provenance tracking you’d apply to any third-party library.
- Classify agents by autonomy level and calibrate controls accordingly. Define what “high-autonomy” means for your organization and build a tiered policy framework around it. High-autonomy agents need runtime enforcement, not just intake checks. Start with your most sensitive systems and work outward.
- Make governance fast enough that developers won’t route around it. Shadow AI is a friction problem, not a defiance problem. If approving a new MCP server takes two weeks, developers will find another way. Automated scanning with clear policy enforcement is the only model that scales at the pace AI development actually moves.
The enterprises that will scale AI safely aren’t the ones with the strictest controls or the most permissive cultures. They’re the ones that match governance to the actual risk of each agent – with a system of record that enforces policy at every boundary, full visibility into every component, and the ability to act the moment something goes wrong.
Authored by Prasanna Raghavendra, Senior Director of R&D, JFrog
