
Willow Joins Okta Cross App Access: Identity-Governed AI Agents for the Enterprise
Your teams are deploying AI agents faster than any governance framework was built to handle. The bottleneck isn't ambition. It's the secure connectivity and AI access control those agents need to work across the tools your business runs on.
Today Willow is joining Okta's Cross App Access (XAA) ecosystem. Every agent connection that flows through Willow can now inherit your existing Okta identity and policy, so AI moves forward without the security reviews, custom governance builds, and unmanaged risk that stall most enterprise deployments. Okta is the access layer for people. Willow is the access layer for agents. Cross App Access is where the two meet, and it's what makes enterprise AI governance a standards-based problem instead of a custom build.
The AI agent security risk is already here:
88% of organizations report confirmed or suspected AI agent security incidents, and 92% of organizations that experienced an AI-related breach lacked proper AI access controls. The pattern isn't that agents are inherently unsafe. It's that most are deployed on static API keys that never expire, create no audit trail, and give IT no visibility into what the agent is doing. This is the shadow AI problem, and it grows every week your teams add another tool.
That leaves security teams with a bad choice: accept unmanaged risk, or stall AI adoption entirely. Willow was built to remove that choice. XAA makes the removal standards-based.What Cross App Access (XAA) actually isCross App Access, formally the Identity Assertion Authorization Grant, is an open protocol built as an extension of OAuth and incorporated into MCP as an authorization extension. It's vendor-neutral, designed to work across any cloud, framework, or SaaS ecosystem, which makes it the practical standard for MCP authorization at enterprise scale.Instead of static API keys, XAA issues dynamic, identity-based tokens scoped to exactly what the agent needs for each task.
This is least privilege for non-human identity: each token is issued in real time through the user's active Okta identity, revocable at any point, and logged for a complete audit trail. Every agent connection flows through your central identity policy, not around it.
AI agent identity and access management, built on OktaAI agent identity and access management is the layer most stacks are missing. On-prem had Active Directory. SaaS had Okta. AI agents get Willow, and with Cross App Access, agent SSO now extends the same identity standards that govern your human workforce to every autonomous worker in your environment.What Cross App Access changes for your teams
- Automatic enterprise compliance. Agents follow the identity-based rules you already enforce. No custom AI governance to build, no new policy framework to stand up.
- Faster deployment. Skip the security reviews that have been blocking automation from going live. Teams move from pilot to production.
- A better experience. Fewer consent prompts, less approval fatigue, no loss of control.
- Secure connectivity across your stack. Agents connect to the tools your teams use under the same AI access control standards that govern people.
Why Willow adopted XAA for AI agent governance?
Willow is the control plane for every AI agent, tool, MCP, and skill in the enterprise. Agent requests pass through Willow's gateway, MCP servers, and orchestration layers, and Willow gives each one a real identity, scoped access, runtime guardrails, and an audit trail tied to a human.Cross App Access lets Willow route and govern that agent traffic through Okta's identity and policy engine directly. Developers building on Willow plug into Okta's central policy engine through XAA and inherit enterprise-grade AI agent access governance, without building custom infrastructure.This is the same model already running in production. At Wix, roughly 5,000 people use Willow every week across about 600 governed tools and MCPs, generating more than 300,000 governed tool calls a week. Innovid and Riskified run on it too. XAA extends that governance to any Okta customer through an open standard."Organizations shouldn't have to choose between adopting AI tools and maintaining visibility over enterprise access," says Aaron Parecki, Senior Director of Identity Standards at Okta. "By connecting apps with the Cross App Access protocol, security teams get more control and users get a better experience without all the OAuth consent prompts
What this means for Willow customers?
AI agents can now interact with Willow without static keys or repetitive manual consent prompts. Access is governed by your existing Okta policies, auditable in real time, and scoped to exactly what the agent needs for each task. Security defines policy once. Employees self-serve safely. That is enterprise AI governance and end-to-end enablement from the same place.Get started with Cross App Access
- Cross App Access is included in your Okta workforce SSO plan. Agent SSO brings XAA into Okta workforce plans at no extra cost and treats agents as first-class identities: okta.com/solutions/cross-app-access
- See Willow's integration in the Okta Integration Network (OIN): https://docs.withwillow.ai/docs/admin/build/mcp-servers/connectors/built-in/okta#authentication-types
- Read Okta's announcement: okta.com/blog/product-innovation/cross-app-access-enterprise-ai

Willow's 17 G2 Badges
AI Agent Identity & Governance: Willow's 17 G2 Badges
Your identity stack knows every employee. It knows their role, their groups, and which apps they can open.
It has no idea that three of them connected an MCP server to Jira last week. Each one used their own token, with write access, from an agent nobody approved.
That gap is why an AI governance platform just landed on G2's Identity and Access Management grid.
In G2's Fall 2026 reports, Willow earned 17 badges across 16 reports in three categories: AI Governance Tools, AI Orchestration, MCP Gateway and Identity and Access Management (IAM). The badges include Best Support in two categories, Easiest to Do Business With, and High Performer in all three. The ratings come from verified users, not analysts or a vendor scorecard. Willow holds a 4.9 out of 5 on G2 with 98+ reviews.
Here's what the badges mean, why the IAM placement is the part that matters, and what users actually said.
One product, three G2 grids
Willow is rated in three G2 categories that rarely overlap: AI Governance Tools, AI Orchestration, and Identity and Access Management.
Each category measures something different:
- AI Governance Tools (611 products) measures whether you can oversee AI with policy, risk controls, and compliance evidence.
- AI Orchestration (925 products) measures whether you can coordinate models, agents, and tools across the business with governance built in.
- Identity and Access Management (302 products) measures whether you control who gets access to what. It's the grid shared with Okta, Microsoft Entra ID, and JumpCloud.
Most AI tools show up in one of these. A gateway sits in one grid, a shadow-AI scanner in another, an agent builder in a third. Getting rated in all three from the same verified users is the clearest signal we have that Willow runs as one control plane, not seven tools stitched together.
Why an AI governance platform belongs on the IAM grid
AI agents are users. They're non-human, they act with delegated authority, and they work at machine speed across dozens of systems.
Traditional IAM answers one question well: can this person sign in to this app? Agents raise harder ones:
- Which tools should this agent see for this task?
- Can it read this Jira project, or also delete from it?
- Whose authority is it acting under?
- Can you prove afterward what it touched?
Answering those takes identity and governance at the moment of action. That's the layer Willow adds. It doesn't replace your identity provider. It inherits from it. Willow sits on top of Okta, Entra, Active Directory, or JumpCloud, and gives every agent a real identity tied to a human, scoped access to exactly the tools it needs, runtime guardrails, and a full audit trail.
Every infrastructure era gets its access layer. On-prem had Active Directory. SaaS had Okta and SSO. Agents need one built for them. G2 just added a Non-Human Identity Management category for the same reason. The market sees the gap.
Analysts grade the program. Users grade the runtime.
Gartner published its first Magic Quadrant for AI Governance Platforms in June 2026. It's a useful map of the program layer: AI inventories, risk tiers, regulatory mapping, and evidence workflows. Every enterprise needs that layer, and the MQ evaluates it well.
Runtime is a different question. What happens when an agent with a token hits a production system at 2 a.m.? Independent reads of the MQ note that its methodology doesn't formally evaluate runtime blocking or sandboxing, and treats MCP governance as an emerging capability.
Peer reviews fill that gap. The people writing them run the platform every day. They don't describe policy documents. They describe the tool call that got scoped, the connection that got caught, the ticket they never had to file.
What verified users actually said
Four themes dominate the reviews. Each one maps to a real problem enterprises hit when AI adoption meets security.
1. Access without the IT ticket
The most common theme across Willow reviews is integrations that were already there on day one.
"Slack, Google Drive, Rocketlane, HubSpot, all connected, no ticket to IT to get any of it working."Adi K., Payments Operation Specialist, Mid-Market
A payments ops specialist isn't a developer, and she'll never write an MCP config. That's the point. Willow ships 1,000+ pre-built connectors and wraps any API as an MCP, so employees connect approved tools themselves and IT sets the rules once. Security stops being the bottleneck and becomes what makes the rollout possible.
2. Least privilege before the call runs, not a report after
Most guardrails try to catch a bad action after it happens. Reviewers describe the opposite.
"Handles governance at the point where the agent actually gets its tools, not as a report after the fact."Eric S., Backend Developer, Enterprise
Eric also named the problem it replaced: agents connecting to internal systems "with way more access than they needed, and no real way to prove after the fact what touched what." Willow generates tools for the task at runtime, so an agent only sees what the job requires. Every call lands in the audit trail and can stream to Splunk or Loki. Guardrails that hope to catch mistakes, versus tools that can't make them.
3. Shadow AI, found and governed
One reviewer titled their review "From Shadow AI to CISO Approval in Two Weeks." Another wrote this:
"Shadow AI detection has caught more unauthorized connections than I expected."Gopal A., Senior Software Engineer, Mid-Market
Most gateways secure the connections you already know about. The real risk is the ones you don't: MCP servers, skills, and agents installed by individual developers. Willow's endpoint and browser sensors find them, then bring them under the same policy as everything else. Visibility plus the control to act on it, from one place.
4. It passes security review, then it scales
"Willow had SOC 2, audit logging, and SSO through our existing IdP in place from day one."Verified User, Real Estate, Mid-Market
Another reviewer's title says it plainly: "Passed our security review faster than any vendor we've onboarded." After approval, scale takes over. At Wix, around 5,000 people use Willow weekly, more than the entire engineering org, across HR, legal, finance, design, and R&D. That's 300,000+ governed tool calls every week, all behind Okta SSO with full audit.
The support badges are the adoption story
Best Support in two categories. Easiest to Do Business With. For an infrastructure vendor, those badges say something specific.
Agent governance is a moving target. The MCP spec changes, new agents ship monthly, and every new use case brings a new edge case. When your control plane has a question, you need an engineer who answers today, not a ticket queue. Reviewers named it again and again: "Seamless Integrations, Zero Downtime, and Same-Day Support." "Fast-Moving Product with Genuinely Responsive Slack Support." "The only vendor with true, great and instant support."
It's also why Willow goes live in seven days, not pilots.
How to evaluate AI agent access platforms on G2
Whether or not you shortlist Willow, here's how to read G2 for this category:
- Check which categories a vendor is rated in. A product rated only in AI Gateways is likely routing traffic. Ratings across governance, orchestration, and identity point to a platform.
- Filter reviews by role. Security, platform, and business users should all show up. If only developers review it, expect a dev tool.
- Look for runtime language. "Scoped," "blocked," "audit trail," "least privilege" describe enforcement. "Dashboard" and "visibility" alone describe observation.
- Weigh recency. This market changes every quarter. A review from 2024 describes a different product.
- Read the Relationship Index. Support and ease of doing business predict how fast you'll get to production.
What High Performer means, and what's next
High Performer means top-tier customer satisfaction, with market presence still growing. That's the honest read for a platform that came out of stealth this year. Our users rated us as highly as anyone in these categories. Market presence is the part we're building, one governed rollout at a time.
Every era gets its access layer. Choose it on purpose now, or assemble it by accident after an incident.

The Willow September 2026 digest blog
The Willow August Digest: Let Agents Act on Their Own. Govern Every Step.
Most of what shipped in Willow this August points the same direction. Agents are doing more without a person in the loop, and the controls had to grow up to match. A background agent that reconnects its own integrations, a model call that gets checked before it runs, a risk score that tightens access the moment it climbs: each one lets you hand off more of the work while keeping the same answer to "who approved this?" A few of these started as requests customers sent us directly. Here's what changed.
Guards now run on the model call itself. They used to stop at the tool call, which left the reasoning step as a blind spot: the model got to think before your policies got a say. That gap is closed. Every Anthropic inference request is now evaluated against your guards before it runs, so the model call is governed like everything else instead of sitting outside the perimeter. You turn it on from the Guards page, and it applies to every inference request from that point on. We ran it across our own workspace first. Enforce guards on inference
.png)
Access to each MCP server is now something you scope, not something you assume. Not every assistant needs to touch every system, but until now connecting a server exposed it to everything by default. You can now set an allowlist of AI tools per MCP server, and the gateway rejects any call from a tool that isn't on it before the request ever lands. One connector, only the tools you meant to expose, configured per server in seconds.That principle, agents get exactly what they need and nothing more, runs through the next two releases as well. Scope an MCP server
.png)
Background agents can hold their own credentials. The re-auth loop is over. Instead of borrowing a person's connection and stalling the moment it expires, an agent now authenticates per integration on its own identity, so no one has to reconnect a tool on its behalf. Point it at the MCP server default, a connected account, or its own custom key, including custom API tools, set the identity once, and it runs unattended from there. Give an agent its own credentials
.png)
Identity risk now decides what an agent can do. Access should react to risk in real time, not wait for a manual review that happens after the fact. User risk scores now sync directly from CrowdStrike and feed your guard conditions, so a flagged identity gets tighter limits the moment its score climbs. The signal your security team already trusts becomes the signal that gates the agent working on that user's behalf. Set your first risk-based guard

Token spend is now a ceiling you set, not a number you find on the invoice. Runaway usage used to surface after the bill did. Now you define per-model and per-conversation budgets once, and the Willow Usage Hooks plugin enforces them client-side across Claude Code, Cursor, and Codex. The Tokens page splits into three views, Dashboard, Analytics, and Optimize, so you can see the spend and cap it from the same place. Set a usage policy
.png)
Radar is now a scorecard, not a spreadsheet exercise. Tracking your AI security posture shouldn't mean assembling the picture by hand every time someone asks. Radar now scores five categories against a target and against your peers, and hands you a start-here list ranking whatever's furthest off track. Drill into any entity for the detail behind its score, and coverage caveats flag where the underlying data is still thin, so the number never claims more certainty than you actually have. Open your Radar scorecard
.png)
Also shipped this month: Cursor Cloud Agents now run as background agents, so their branch pushes and pull requests fall under the same governance as every other agent, with no separate lane for cloud-run work. Willow also got a design refresh, including a redesigned admin dashboard that brings adoption, usage, and security onto one screen.
New to watch and read
A few things worth your time if you haven't seen them yet:
- Eyal on why token spend is a security signal, not just a line on the invoice (6 min read).
- Shalev on the shadow AI already inside your org, and what to actually do about it (9 min read).
- Eyal's field note from Black Hat 2026, on why governing AI agents beats banning them (1 min).
- And Eyal on how to stop paying for the tokens your agents waste (1 ).
The through-line
The pattern under all of it is the same one from July, pushed a step further. The old tradeoff was block a capability or trust it. This month's work is about letting agents run on their own without that being the same thing as letting them run unchecked. More autonomy on one side, more control on the other, set as policy rather than left to hope.
Questions on anything above? Just reply, we read every one.
Keep building. The Willow Team

Grok Bot vs Other Autonomous AI Agents: How It Works, How It Compares, and How to Run It Safely
On August 11, 2026, xAI launched Grok Bot: a team of always-on AI agents that have their own computer, sign into the tools you already use, and work 24/7 until a job is done. It is the clearest sign yet that the market has moved past the chat assistant. The new unit of AI at work is not a prompt box, it is a teammate you hand a task to.
This guide explains what Grok Bot actually is, how it compares to the other autonomous agents enterprises are evaluating, and the question every one of them forces a security team to answer: how do you let an agent sign into your systems and act on its own without losing control of what it can touch.
What is Grok Bot?
Grok Bot is xAI's autonomous agent product, launched in early beta on August 11, 2026. Instead of answering inside a chat window, each Bot gets a computer of its own in the cloud, signs into your applications, and does multi-step work end to end, coming back only when something needs your approval.
The defining features, per xAI's launch:
- A computer of its own. Bots run on a shared cloud computer, so work does not stall when you close your laptop. They sign into apps, tools, and websites, including platforms with no clean API or MCP, and drive the interface the way a person would.
- Message it like a teammate. You hand off work in a chat thread from desktop or iOS. There are no workflows to build first.
- Teams of Bots. People run several Bots in parallel with a chief-of-staff Bot on top. Bots message each other, pass work, and pull a human in only for judgment calls.
- Learns by watching. Ask a Bot to follow along while you do a task once. It saves the steps as a routine and runs it on its own next time.
- Availability and pricing. Beta is open to SuperGrok Heavy ($300/month), Cursor Ultra ($200/month), and Cursor Premium Teams ($120/seat/month) subscribers, distributed through Cursor. Enterprise access is a waitlist as of launch.
The examples xAI showcases are ordinary enterprise work: a sales Bot updating the CRM from call transcripts and drafting follow-ups, an operations Bot seating new hires and processing invoices from Gmail, and an engineering Bot reproducing a bug in the UI, filing the ticket, and handing the fix to another Bot.
How Grok Bot works
The mechanism is what makes Grok Bot powerful and what makes it a governance question. A Bot logs into your tools once, then uses your apps and websites just like you would, including the ones that are hard to navigate or have no API. Because it runs on its own always-on cloud computer, it keeps working unattended, and because Bots can coordinate in group chats, one task can fan out across several agents acting at the same time.
Read that as a security leader and the picture sharpens. An autonomous, non-human worker is signing into your systems, often with a human's credentials, acting across apps around the clock, and driving interfaces directly where no API gate exists to check it.
Grok Bot vs other autonomous AI agents
Grok Bot enters a crowded field. Every major lab now ships an agent that plans and acts, not just answers. They differ most in where they run, how they reach your tools, and how much oversight they build in.
| Agent | Runs on | How it reaches tools | Autonomy model | Oversight built in | Enterprise availability (Aug 2026) |
|---|---|---|---|---|---|
| Grok Bot (xAI) | Its own always-on cloud computer | Signs in and drives apps and websites directly, including no-API tools | Teams of Bots, parallel, 24/7, learns by watching | Approval only when a Bot decides it is needed | Beta on consumer and Cursor plans; enterprise waitlisted |
| ChatGPT Work / agent (OpenAI) | Virtual computer and browser | Browser control plus connectors and tools | Hours-long multi-step projects (launched Jul 9, 2026) | Step approvals and takeover on sensitive actions | ChatGPT Enterprise |
| Claude Cowork + Computer Use (Anthropic) | Cloud and desktop harness | MCP, connectors, and screenshot / mouse / keyboard computer use | Task and background agents, permission-gated | Oversight-first: permission gating, human-in-the-loop, Compliance API | Claude Enterprise |
| Gemini agents (Google) | Google cloud, Workspace, Antigravity | Workspace, connectors, computer use, deep research | Research and coding agents (Jules), enterprise agent platform | Workspace admin controls | Google Cloud / Workspace |
| Cursor background agents | Cloud dev environment | GitHub and repo access, triggered from Slack or IDE | Coding tasks that return a pull request | Pull-request review by a human | Cursor Teams |
| OpenClaw (open source) | A server or cloud instance you host | Messaging apps (Slack, Discord, WhatsApp, Signal) bridged to any LLM; writes its own skills | Scripts autonomous workflows in plain language, self-extends, keeps long-term memory | Whatever you build in. Self-hosted, so oversight is on you | Open source (Apache 2.0), self-hosted |
| Hermes Agent (Nous Research) | Anything from a $5 VPS to a GPU cluster | 20+ platforms (CLI and messaging) to any LLM provider; self-authored skills | Closed learning loop: curates its own memory, writes and improves its own skills | Self-hosted, so oversight is on you | Open source (MIT), self-hosted |
The honest reality check across all of them: autonomous agents are impressive and still imperfect. On the OSWorld 2.0 benchmark of realistic long-horizon computer-use tasks, a leading model with maximum reasoning fully completed only about 21% of them (OSWorld 2.0, 2026). That is not a reason to avoid these agents. It is the reason they need guardrails, approval gates, and an audit trail, because an agent that is right most of the time still acts on your systems the rest of the time.
Two contrasts matter most for an enterprise buyer. First, access model: Grok Bot's willingness to drive any interface, even without an API, is its superpower and its blind spot, because interface-level action is the hardest kind to govern centrally. Anthropic's Claude leans the other way, trading some raw autonomy for permission gating and a compliance trail. Second, enterprise readiness: several of these ship inside an enterprise plan with admin controls today, while Grok Bot's enterprise tier is still a waitlist, which means early adopters are running it on individual or team plans, outside central IT.
The governance gap this class of agent creates
Grok Bot is not uniquely risky. It is a clear example of a pattern every autonomous agent shares, and the pattern is what security teams have to govern.
- Borrowed identity. When a Bot signs in with an employee's credentials, its actions are indistinguishable from that person's. There is no separate, revocable identity for the agent, and no clean way to answer "which agent did this, on whose behalf."
- Standing access, 24/7. An always-on agent holds live sessions into your systems around the clock, long after the human who set it up has logged off.
- Interface-level reach. An agent that drives the UI directly, with no API in the path, slips past the API gateways and connectors most governance is built on.
- Fan-out. Teams of agents acting in parallel multiply every one of these exposures at once.
- Shadow adoption. With enterprise tiers waitlisted, employees adopt these agents on personal or team plans first. The open-source, self-hosted ones make this sharper still: an employee can stand up OpenClaw or Hermes on a $5 VPS, point it at any LLM, and let it write its own skills, entirely outside IT. Security often learns about them after they are already working inside the business.
This is the same root issue behind the OWASP Excessive Agency risk (LLM06): an agent with more access and autonomy than its task requires, and no boundary enforcing the difference. The productivity is real. So is the exposure, and it does not show up until someone asks what a specific agent touched last Tuesday.
How to run autonomous agents safely: Willow Background Agents
The answer is not to block this class of agent. Blocking pushes it onto personal devices and removes your visibility entirely. The answer is to run the always-on, does-the-work-for-you model through a layer that gives every agent an identity, scopes what it can do, and logs every action.
That is what Willow Background Agents provide. A background agent fires on a trigger, does the job, and logs every step, scoped and identity-backed from the start. Set one up to own a recurring task like prospect research, support triage, or reconciliation, and it carries the work forward with the same governance that covers the rest of your agents:
- A real identity per agent, inherited from your existing identity provider (Okta, Entra ID, JumpCloud) and tied back to a named human, so every action is attributable and access is revoked the moment the person leaves.
- App-aware permissions that scope not just which tools an agent can reach but what it can do inside each one, read versus write versus delete, on which data.
- Runtime guardrails that inspect prompts, tool calls, and outputs before they execute, so a risky action pauses for human approval instead of running unseen.
- A full audit trail tied to a real employee, streamed to your SIEM.
- Shadow-AI discovery through endpoint sensors and a browser extension, which is how you find the ungoverned agents already running, including a Grok Bot, an OpenClaw instance, or a ChatGPT agent an employee installed on a personal plan.
Willow does not replace the agent you choose. It is the control plane the agent runs through. If you want the Grok Bot experience, an always-on teammate that finishes the work, Willow Background Agents give you that model with governance built in, and Willow's discovery surfaces the ungoverned autonomous agents that adopt themselves across your org before central IT ever approves them. This is the same control plane that governs roughly 600 tools and about 5,000 weekly active users at Wix, every action tied to a real identity.
Should your enterprise use Grok Bot?
Grok Bot is a genuinely strong product and a signal of where work is going. For individuals and small teams on the eligible plans, it can take real work off your plate today. For an enterprise, two facts should shape the decision. It is in beta with the enterprise tier still waitlisted, so central controls are limited. And like every agent in its class, it needs an identity, permission, and audit layer around it before it touches regulated systems.
The question is no longer whether autonomous agents belong at work. They are already here, adopting themselves one download at a time. The question is whether you can see them and govern them. Choose the agent that fits the job, then run it through a layer that makes it accountable.

Top 6 AI Agent Observability Platforms for 2026
AI agent observability means tracing what your agents do (every tool call, model inference, sub-agent handoff, and action taken inside a connected system) across the full execution path.
It is not application performance monitoring, which watches infrastructure and response times.
And it is not LLM tracing alone, which records model inputs and outputs and stops at the tool-call boundary.
Agent observability requires going past the API boundary (which app is being accessed) to see everything an agent does while inside an app or tool like Linear, Salesforce, or PostHog.
KPMG's Q4 2024 survey found 51% of respondents exploring AI agents and 37% piloting them, for a combined 88% (KPMG, Q4 2024). A 2026 LangChain survey found that 89% of organizations had implemented some form of agent observability, but only 62% had detailed tracing for individual steps and tool calls (LangChain, 2026).
AI agents are off to a flying start within organizations attempting to capitalize on the capabilities of this new technology.
But often, when something goes wrong at scale (and when working with AI agents it is a matter of when, not if it does!) there is no trace, no audit record, and most often, no way to say who authorized the action.
TL;DR
- Choose Phoenix or Langfuse for self-hostable tracing and evaluation; add Braintrust when production traces must feed CI/CD evaluations.
- Choose Fiddler for combined ML and agent monitoring; choose Zenity for discovery, security posture, and inline response.
- Add Willow when agents need employee-bound identity, action-level permissions, runtime enforcement, and authority-linked audit records.
Why AI Agent Observability Matters in Production
In PwC's 2025 survey, 79% of surveyed executives said their companies were already adopting AI agents (PwC, 2025).
An AI agent in production is not the same problem as an AI model in production. A model produces an output.
An agent takes an action. Often, a lot of actions. Autonomously. And unpredictably.
The mistake some make is thinking of AI agent monitoring in the same manner as model monitoring.
The model layer covers latency, cost per token, input and output logs.
That is necessary, and adds a valuable layer of visibility.
But the model layer stops being sufficient the moment an agent has read or write access to a production database, app, or tool.
Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, driven in part by inadequate risk controls (Gartner, June 2025). And getting controls for AI agents right is no simple task.
Agents act inside tools, and multi-step workflows branch into execution trees. Prompts alone cannot guarantee that an agent will stay within the intended path, especially when prompt injection or poisoned content influences a later step. One important way to reduce the blast radius is to enforce least-privilege, action-level permissions alongside approvals, sandboxing, network controls, and monitoring.
Those controls need permissions infrastructure. A practical solution is to connect agent permissions to the IdP and access policies the organization already maintains, such as Okta or Entra ID.
The AI Agent Production Gap
of organizations are exploring or piloting AI agent initiatives
KPMG, 2024
of organizations have already adopted AI agents across business functions
PwC AI Agent Survey, May 2025
of enterprise software applications will include agentic AI by 2028
Gartner, October 2024
What to Look For in an AI Agent Observability Platform
The criteria that separate a real agent observability platform from a repurposed APM or LLM-tracing tool come down to five questions.
(1) Can it trace a full multi-step run as a single linked tree?
A span is one recorded step in a workflow. The root span represents the complete run, while child spans represent the model calls, tool calls, and handoffs inside it.
For example, a support agent might read a Zendesk ticket, check Salesforce, query a billing system, and issue a refund. One linked trace should nest those four steps so an incorrect refund can be traced back to a faulty entitlement returned earlier in the workflow.
(2) Does it see inside the tool, or only that the tool was called?
Most platforms stop at the MCP (Model Context Protocol, the open standard that lets agents connect to external tools) gateway. That tells you which app or tool was connected to.
Action-level visibility means capturing which record was read or written, which query ran, which file changed. In short, it captures the bulk of actions AI agents complete, those that happen past the gateway and inside the tool itself.
(3) Can it name the employee behind each action?
This asks whether each action traces back to the human authority under which it executed, using the IdP (the directory that manages employee access through systems such as Okta, Entra ID, or JumpCloud).
A service-account log may show that salesforce-bot changed an opportunity. An authority-linked record should additionally show that Priya's agent requested the change, which permission allowed it, and whether the action stayed within her approved scope.
(4) Does it read your existing IdP and CI/CD, or bolt on a second system?
If a platform copies users and permissions into a separate credential store, a role change may update in the IdP but not in that specific copy. A transferred employee could retain agent access that the current role no longer permits. This is credential drift.
The same split applies to CI/CD. A platform that hooks into the pipeline you already run gates every agent change on tests that fire automatically, so a prompt or model tweak that drops quality or safety scores gets caught before it ships. One that makes you stand up a separate evaluation system leaves those checks outside your release process, where they get skipped under deadline and stop matching what actually reaches production.
(5) Can it export an audit trail an auditor can follow?
SOC 2, SOX, and the EU AI Act want evidence that specific actions occurred. The export needs structured fields, such as action type, tool, data touched, identity, timestamp, and outcome.
| Platform | Multi-step tracing | Action-level visibility | Authority attribution | IdP + CI/CD fit | Audit export | Shadow AI discovery | Runtime enforcement |
|---|---|---|---|---|---|---|---|
| Arize Phoenix | Yes | Span-level | N/D* | Yes | Partial | N/D* | N/D* |
| Langfuse | Yes | Span-level | N/D* | Yes | Partial | N/D* | N/D* |
| Braintrust | Yes | Span-level | N/D* | Yes | Partial | N/D* | N/D* |
| Fiddler AI | Yes | Model / behavioral | N/D* | Partial | Yes | N/D* | Yes |
| Zenity | Execution-path | Step-level | N/D* | Partial | Partial | Yes | Yes |
| Willow | Partial | Record-level | Yes | Yes | Yes | Yes | Yes |
These tools work at different layers and often run side by side, so the table above compares capabilities rather than ranking rivals. We checked each against the vendor's public documentation as of July 2026. N/D* means we could not find public confirmation of a capability, not proof that it is missing.
Keep in mind here that IdP + CI/CD fit means the platform reads your identity provider for its own single sign-on and role-based access and hooks into your CI/CD pipeline.
Langfuse, for example, offers enterprise SSO and RBAC, with SCIM and audit logs on its Enterprise tier.
That is separate from delegated human-authority attribution, which asks whether each production tool action is recorded against the employee authority under which it executed, including the effective permission scope at that moment.
The Top 6 AI Agent Observability Platforms for 2026
The six products occupy different but complementary layers, from open-source tracing and evaluation, to enterprise monitoring and security, and identity, access, and runtime governance.
1. Arize Phoenix: Open-Source LLM Tracing and Evaluation
Arize Phoenix is self-hostable with native support for OpenAI Agents SDK, LangGraph, and CrewAI (among other frameworks).
Tracing and quality assessment live in one place. Multi-step traces record each part of a run, and automated evaluation can use another LLM to grade those recorded steps against defined criteria.
The core project is source-available under the Elastic License 2.0. It's not a permissive OSI license like MIT, but you can still self-host it, modify it, and run it inside your own commercial product for free (Phoenix license documentation).
What the Elastic License forbids is taking Phoenix and offering it to third parties as a hosted or managed service. You cannot resell Phoenix itself, but your org can run it in production for its own purposes.
Phoenix documents user authentication and identity-provider integration, but as of mid-2026 the identity it records is the login or service account, not the employee whose permissions an agent action ran under (Phoenix authentication documentation).
2. Langfuse: Open-Source LLM Engineering Platform
Langfuse covers similar ground to Phoenix. Its core observability, evaluation, prompt-management, and related APIs are MIT-licensed and free to self-host. However, enterprise features such as SCIM, audit logs, and data-retention controls require a commercial license (Langfuse licensing).
Where it pulls ahead of Phoenix is in prompt engineering.
Prompt versions are tied to the traces they produced and the eval scores those traces earned. Practically, it means you can measure the quality impact of a single prompt change across a live deployment.
Strong CI/CD teams like it as an evaluation gate before an agent change ships.
Langfuse is SOC 2 Type II and ISO 27001 certified and captures trace user metadata and Enterprise audit logs, but that attribution stops at the session level. As of mid-2026 it has no documented way to tie a downstream tool action to the authorizing employee's live permissions.
3. Braintrust: LLM Evaluation and Tracing
Braintrust connects pre-production evaluation with production traces and failures.
Its shared data model lets a production trace be replayed as an eval input, and an eval failure be traced back to the run that triggered it.
That makes it the pick for teams whose dev-time eval scores and production behavior keep diverging.
It also supports a bring-your-own-cloud deployment, where Braintrust's data plane runs inside the customer's AWS, Azure, or Google Cloud environment. That gives teams with data-residency requirements control over where sensitive trace and evaluation data is stored.
Braintrust can tag traces with application-provided user metadata, but a trace still shows what happened without recording who was accountable for it. As of mid-2026 it has no documented way to bind an action to the employee who authorized it.
Braintrust is SOC 2 Type II compliant and supports HIPAA requirements, with business-associate agreements available (Braintrust security documentation). Those controls secure the platform itself. Tying an action to the employee who authorized it is a separate job, and not one Braintrust does.
4. Fiddler AI: Enterprise ML and AI Observability
Fiddler AI approaches agent observability from the enterprise-risk side.
It monitors predictive ML models and agentic AI under one governance dashboard, tracing an agent's full run over native OpenTelemetry, from the top-level request down to each step.
On top of that sit real-time guardrails for faithfulness, hallucination, toxicity, jailbreaks, and PII or PHI detection. Fiddler documents sub-100ms guardrail latency and deployment inside the customer's environment (Fiddler Guardrails documentation).
Fiddler maps audit evidence to frameworks such as NAIC for insurance and SR 11-7 for bank model risk. It also documents air-gapped deployment options, enabling banks, insurers, and government teams to run it inside networks that do not touch the public internet.
Fiddler is a strong fit for teams that want predictive ML and agentic AI monitoring in one enterprise environment.
Fiddler documents agent traces, decision context, and governance evidence, but its visibility stays at the model and behavioral layer. As of mid-2026 it does not reach inside a connected tool to show which record moved, or bind that action to an employee identity (Fiddler monitoring documentation).
5. Zenity: Enterprise Agent Security and Observability
Shadow AI discovery is a central Zenity capability, and the company raised a $38M Series B to build it out.
Zenity discovers and inventories agents across SaaS copilots (Microsoft 365 Copilot, Salesforce Agentforce, and ChatGPT Enterprise) and homegrown frameworks built on AWS Bedrock or Google Vertex AI.
It then maps agent behavior to attack paths.
For many enterprises the first observability problem is simply knowing which agents exist. Zenity covers that well. Its research found upwards of 79,000 apps built per organization on copilots and low-code platforms, with 62% of them carrying security vulnerabilities.
Its AI detection-and-response layer blocks unsafe agent actions inline, before they land. Most monitoring tools only flag a risky action after it has already run.
Zenity documents agent ownership, permission mapping, and step-level monitoring, and can see that a Copilot agent requested a file. However, as of mid-2026 it stops short of tying that action to the employee whose authority it ran under (Zenity AI Observability).
6. Willow: Identity and Access Layer for AI Agents
Willow is not a pure-play observability tool. It is the agent identity and access platform, and it is on this list to show the layer the other five leave open: identity, access scope, runtime enforcement, and authority-linked audit records.
Observability records the action. Willow records both the action and the authority behind it.
With Willow, every agent action is bound to a real employee identity from your existing IdP (Okta, Entra ID, JumpCloud), scoped to app-aware permissions at the action level, and written to an immutable audit trail.
The audit trail shows whose identity the agent inherited, whether that person's live permissions covered the action, and whether it stayed in scope.
At Wix, Willow now governs roughly 5,000 weekly active users across about 600 connected tools and MCPs, and more than 300,000 governed tool calls a week (Wix case study).
Willow also discovers agents, MCP servers, and AI tools already running in the organization, including unapproved ones. Endpoint- and browser-level detection helps teams identify shadow AI and either bring it under existing governance controls or block it.
Willow's governed MCP gateway serves each agent only the tools and actions its identity and task permit (Willow Tools and Skills). An agent cleared to use Jira for one workflow does not automatically receive the tool's full range. The policy is enforced at runtime instead of only being detected after the action.
A one-click kill switch can revoke an agent or tool across the organization when something looks wrong (Willow product overview).
Separately, every audit event exports to Splunk, Loki, or Grafana, whatever SIEM you already run (Willow July 2026 product digest).
How to Choose the Right Platform for Your Production Stack
Where you start depends on where your agents already are.
If you are still building or evaluating models, start with open-source tracing. Arize Phoenix and Langfuse both self-host, speak OpenTelemetry (the open standard for distributed tracing), and show each step of a run.
Add Braintrust when you need a CI/CD evaluation loop that turns production traces into reusable test cases. Its bring-your-own-cloud deployment is relevant when the data plane must remain inside the customer's cloud environment.
Once agents are live under real enterprise risk, though, you need to move up to an enterprise monitoring layer.
Fiddler fits teams watching predictive ML and agentic AI on one compliance dashboard, or anyone who needs an air-gapped install.
Zenity fits teams whose main exposure is SaaS copilot sprawl and who need agent discovery, shadow AI detection, and attack-path analysis first.
Once agents act on production systems, observability alone leaves a gap. It records what happened, but not who was allowed to do it. Willow fills that gap. It ties each agent action to a real employee's authority, enforces scoped permissions as the action runs, and records who authorized it.
The Identity Attribution Gap Most Platforms Leave Open
Varonis found that 99% of the environments it analyzed had sensitive data exposed in ways AI could surface. Separately, 98% contained unverified applications, including shadow AI (Varonis, 2025). In Proofpoint's U.S. survey, 80% of CISOs reported concern about potential customer-data loss through public GenAI platforms (Proofpoint, 2025).
That is the exposure Willow exists to close: agents acting without enforced identity, scope, and runtime controls.
When an incident involves an agent action, the observability log says "the action happened." The identity layer says "Sarah's agent, scoped to read-only access in Project Alpha, took an action that exceeded her authorization; here is the audit-trail entry, the scope at the time, and the timestamp."
The two layers are additive. The observability platform provides the operational record, and the identity layer provides permissioning, efficient access management, and audit logs linked to human identities. Together they provide a foundation for scaling agentic AI rollouts while maintaining control and visibility.
The Identity Attribution Gap
of organizations have sensitive data already exposed to AI tools, including tools employees added without IT approval
Varonis State of Data Security, 2025
of U.S. CISOs report concern about customer data loss via public GenAI platforms
Proofpoint Voice of the CISO, 2025
can reliably trace agent actions to a human or system across all environments
Cloud Security Alliance / Strata Identity, February 2026
Further Reading
- Willow Platform, Identity and Access Layer for AI Agents: https://withwillow.ai/platform/identity-access
- Willow Wix Case Study: https://withwillow.ai/blog/wix-case-study
- Arize Phoenix: https://arize.com/phoenix/
- Langfuse Documentation: https://langfuse.com/docs/observability/overview
- Braintrust Platform: https://www.braintrust.dev
- Fiddler AI Agentic Observability: https://www.fiddler.ai/agentic-observability
- Zenity AI Observability: https://zenity.io/platform/ai-observability
- Varonis 2025 State of Data Security Report: https://www.varonis.com/blog/state-of-data-security-report
- KPMG AI Quarterly Pulse Survey, Q4 2024: https://kpmg.com/kpmg-us/content/dam/kpmg/corporate-communications/pdf/2025/q4-pulse-deck.pdf
- PwC AI Agent Survey, 2025: https://www.pwc.com/us/en/tech-effect/ai-analytics/ai-agent-survey.html
- Gartner press release, June 25, 2025: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027

The Willow August 2026 Digest
We're sending this month's digest a few days early, before August 1st, because the team is heads-down at Black Hat next week. Here's what changed: the releases and fixes that actually change how you run the platform day to day, including a few that started as requests you sent us.
The shift this month has one shape. AI that does more on its own, with the controls to match. Agents that run without a person driving every step, and safeguards that keep pace as adoption grows.
Come meet us at Black Hat. Booth #5905, the calmest spot on the show floor, packed with demos, games, prizes, and outdoor SWAG. We're also hosting a private dinner for security leaders on Wednesday, August 5. Limited seats. Register and pre-book your time.
Stop babysitting the work that doesn't need you: Background Agents are live
Background Agents are now generally available. Until now, a recurring task like prospect research or support triage still needed a person to kick off every run. Now you set an agent up once to own the task, and it carries the work forward and reports back, with the same identity, scope, and audit you already have everywhere else in Willow. Autonomy without a blind spot. We run these ourselves. Set up your first background agent.
Put your security policy inside the coding agent: Guard Hooks
Guard Hooks run your Willow policy directly inside Claude Code, Cursor, and Codex, inspecting every prompt, tool call, and output before the model acts. Malicious prompts get blocked, risky commands need approval, and secrets or PII are masked on the fly. No proxy, no code changes. It deploys in minutes as a one-click plugin, so the policy travels with the coding agent instead of stopping at the network edge. Turn on Guard Hooks.
Oversight without the wait: Slack Ranger
Oversight doesn't have to mean a bottleneck. When an agent is about to take an action that needs sign-off, the request lands in Slack with the full picture, and anyone can approve or decline in seconds. Risky moves still get a human yes. It's just a fast one, and it's how our own team signs off on agent actions every day. Connect Slack Ranger.
One line of control, wherever agents run: AWS Bedrock AgentCore
Agents don't stay in one place, so governance shouldn't either. Agents built on Amazon Bedrock AgentCore now run through Willow with the same oversight and controls as everything else. Where an agent lives stops deciding whether it's governed. Bring in your Bedrock agents
The data-residency blocker is gone: EU hosting
If data residency requirements have been holding back your rollout, that blocker is gone. Willow now runs entirely on European servers, so teams with strict requirements can move forward with confidence. Nothing to reconfigure, just where it runs. Teams are already live on it. Move to EU hosting.
Connect any tool in about two minutes
Setup used to mean someone learning each integration one at a time. A guided, step-by-step flow now gets any tool connected in about two minutes, and anyone can run it themselves. It's the same flow we use internally. Connect your first tool.
Also shipped this month
Agent Memory Discovery surfaces the notes and instructions your agents rely on. Integration Conditions let you set rules for when an automated action should or shouldn't run. Teams can now set their own login session limits. Safety checks now catch hidden or disguised text. A new On-Prem page centralizes self-hosted setup. POV Management tracks every pilot in one view. Realtime Alerts send safety warnings straight to Slack. Group Admin Roles give each team its own owner, viewer, and lead. And Run Impersonate lets an admin sign in as another user for support.
New to watch and read
A few things worth your time if you haven't seen them yet:
- Shalev introduces Governed Background Agents with Willow (1 min watch).
- Wix x Willow: The Access Layer Behind 1M+ AI Tool Calls a Week (1 min watch).
- Shalev on Willow x Claude's Compliance API: closing the audit gap for agents (2 min read).
- Shalev's framework: How to Deploy AI Background Agents in the Enterprise (5 min read).
The pattern, if you're tracking it
Last month the releases were about making the capabilities you already had safe enough to turn all the way on. This month the capability itself grew. Agents that run on their own, without a person driving every step. What did not change is the rule underneath it: every bit of new autonomy shipped with a matching control. A background agent that owns a task, and the same audit trail behind it. A coding agent that acts, and a policy running inside it. A faster human yes, not a removed one. That's the bet behind Willow. The fastest way to give AI more room to work is to build the controls precise enough that giving it that room stops being a risk.
Questions about anything above? Reply or reach out, the team would love to discuss.

The 8 AI Agent Security Tools Enterprise Teams Should Evaluate in 2026
Most security stacks were built to watch people. But the fastest-growing actor in your enterprise now is an AI agent that no security tool was built to watch.
Agents log in with credentials no human owns, reach into tools like Jira, GitHub and Salesforce, and act faster than any human reviewer can follow.
The market has responded with a sprawling set of security tools for AI agents, each covering different areas, from catching risky tool calls to controlling what each agent is allowed to do.
Heading into 2026, 48% of cybersecurity professionals expected agentic AI to become the number-one attack vector (Dark Reading, 2026). Yet most enterprises still lack AI-specific security controls.
Below are 8 AI agent security tools spread across those five core categories. Each is strongest at one job, whether that's spotting shadow AI, screening prompts for injection, or scoping what an agent can do inside a tool.
TL;DR
- AI agent security splits into five categories, and no single tool covers them all.
- Match the tool to the gap you actually have.
- Identity and app-aware permissions are the layer most stacks still miss.
| If your gap is... | Start with |
|---|---|
| Agent identity and action-level permissions | Willow |
| You can't see what agents are running | Zenity, Operant AI, Willow |
| An unusually wide GenAI surface in one platform | Prompt Security |
| Behavioral intent across execution traces | Lasso Security |
| The live inference path (injection, leakage) | Lakera |
| Model supply-chain and pre-release scanning | Protect AI |
| Government and regulated inference defense | CalypsoAI (F5) |
| Endpoint, runtime, and gateway in one | Operant AI |
What does "AI agent security" mean in 2026?
AI agent security means controlling what an autonomous, tool-using agent is allowed to do.
Doing it well takes real engineering work. Policy documents alone won’t satisfy regulators or security needs.
A human has to be able to see what the agent is doing and step in to stop it.
It is broader than older "LLM security," which mostly inspected prompts and responses within AI chatbots.
An AI agent that can call tools, build its own memory, chain steps together, and even collaborate with other agents fails in ways a chatbot never could.
The top new risks, as identified in the OWASP Top 10 for Agentic Applications 2026 (released December 2025), are:
- Tool misuse
- Rogue agents
- Agent goal hijacking
- Cascading failures across multi-agent systems
The OWASP agentic list addresses what the older LLM Top 10 never covered: what an agent does with its access (deleting a production database, approving a payment, or changing another user's permissions).
The potential attack surface for agentic applications is both broad and deep. Non-human identities outnumber human identities 45:1 on average, and up to 144:1 in cloud-native environments (Cloud Security Alliance, May 2026). That’s the breadth. The depth comes from how far agentic AI can reach into your systems.
Each of those machine identities is a live login that can reach tools and data, and most carry more access than the job needs.
In IBM's 2025 report, 13% of organizations had a breach of an AI model or app. An additional 8% were uncertain whether they had been compromised, which is what happens when the identities involved were never visible in the first place.
Strikingly, of organizations who were compromised, 97% of them did not have AI access controls in place (IBM, “Cost of a Data Breach Report 2025”).
The most common issue we see, and the one we built Willow to cover, is that AI agent identities were never set up through formal access management, permissions are not tied to real human identities, and there is no clear process to review or revoke permissions when a project ends or a team member changes roles/leaves the company.
The guiding principle for AI agent security is that you cannot secure an agent you cannot see or control. This leaves you two main problems to solve: visibility and control.
What are the categories of AI agent security tools?
| Category | Lead tool | What it secures | Where it stops |
|---|---|---|---|
| MCP gateway & tool-call governance | Operant AI | Every MCP tool call, inspected and enforced | Agent permissions not inherited from your IdP |
| Runtime detection & response | Lasso Security | Agent behavior across live execution traces | No in-tool read/write permissions |
| AI security state management & discovery | Zenity | Agent inventory, shadow AI, attack-path mapping | Discovery and detection, not in-app permissions |
| Inference firewall & red teaming | Lakera | Prompts and responses on the live inference path | No agent identity or in-tool governance |
| Identity & access governance | Willow | Every agent action tied to a real employee's in-app permissions | No model scanning or pre-release red teaming |
The market can be broadly divided into five core categories. A tool that is excellent in one category of work may do little for another. So it’s important to build a stack that covers everything your organization’s AI agent workloads truly need.
1) MCP gateway and tool-call governance
Such tools sit between every agent and tool. They inspect and enforce each call.
MCP (Model Context Protocol) is the connection standard agents use to reach tools. The gateway sits on that path and checks every call.
2) Runtime detection and response
Watches the agent's live execution (tool calls, memory reads, retrievals) and flags or blocks abnormal behavior.
It is the agentic equivalent of endpoint detection.
3) AI security state management and discovery
Finds every agent and AI tool already running, maps how they connect, and scores the risk. This is how you surface shadow AI before it surprises you.
4) Inference firewall and red teaming
Inspects prompts and responses for injection and data leakage, and attacks your own agents on purpose to find holes first.
5) Identity and access governance
Gives each agent a real, human-linked identity and app-aware permissions (granting an agent a specific set of actions inside a tool (read, write, or delete) scoped to what it needs).
This is the layer most stacks still skip.
Every category matters, though few teams need all five on day one. Start with the category that matches your biggest exposure and read those tools first.
The 8 AI agent security tools enterprise teams should evaluate in 2026
Several were acquired by larger security vendors in 2025, which changes how you buy them.
Tools were selected based on (1) category leadership in at least one of the five AI agent security disciplines, (2) documented enterprise production deployments, and (3) publicly verifiable security controls. Emerging research projects without enterprise deployability were left out.
1) Prompt Security: broadest GenAI surface coverage
Prompt Security covers a wide GenAI surface, with broad visibility into how AI enters a company.
You’ll find employee GenAI usage, homegrown apps, code assistants, and agentic AI in one platform.
It discovers shadow AI through a browser extension and network-level visibility, and inspects every prompt and response for injection and data leakage (Prompt Security).
Gartner named it a Cool Vendor in AI Security.
SentinelOne acquired Prompt Security in September 2025 for about $180 million, so it now ships inside a public-company security portfolio.
For agent control, it runs an MCP gateway that enforces allow and block policies on each tool call, stopping a disallowed or shadow-MCP call in real time before it reaches the server.
What it does not do, though, is bind each action to a directory-provisioned employee identity, or enforce read-versus-write permissions inside the connected tool.
2) Lasso Security: behavioral intent detection
Lasso Security leads on behavioral intent. It reads what an agent is trying to do across its full execution trace. Something prompt scanners never see.
Its Intent Security engine claims sub-50ms behavioral analysis at a vendor-stated 99.83% detection accuracy, and Lasso ships the first open-source security gateway for MCP (Python, MIT-licensed), so teams can audit the code and self-host it (Lasso Security).
Lasso is SOC 2 Type II certified and was listed by Gartner as a Cool Vendor for AI Security in 2024.
Beyond detection, Lasso enforces inline, stopping or quarantining a risky or hijacked action at the proxy layer before it executes (rather than merely alerting after the damage is done). It also covers discovery, risk assessment, and red teaming.
As of mid-2026, though, it does not enforce granular read-versus-write permissions inside a tool, and agent identity is not tied to a human directory.
3) Lakera: real-time inference firewall
Lakera is built for the inference path. Its Guard API is a real-time firewall that catches prompt injection, jailbreaks, and data leakage before they reach the model.
It pairs that with Gandalf, a public AI red-team community with 1M+ users and 80M+ adversarial prompts.
The Lakera team is known for its sharp research, too, including a zero-click remote-code-execution exploit through MCP and agentic IDEs (Lakera).
Its API-first, low-latency design suits teams hardening live request paths, because a sub-50ms REST call adds negligible latency to production traffic and drops in front of any LLM without forcing you to re-architect the app.
Lakera holds both SOC 2 Type II and HIPAA attestations.
Check Point announced its acquisition of Lakera in September 2025, and the platform now anchors Check Point's Global Center of Excellence for AI Security.
The platform's gaps are agent identity and MCP action-governance on live production traffic.
As of mid-2026 it provides no agent-identity model. Additionally, while it screens MCP interactions for injection risk, it does not run an action-governance gateway on live MCP traffic. This means it flags a poisoned tool description but cannot stop an agent from invoking a tool it should not. An important distinction. Teams still need a separate policy gateway to allow or block each call.
4) Protect AI: deepest model supply-chain security
Protect AI’s Guardian scanner reads 35+ model formats for backdoors and deserialization attacks.
Its Layer product adds runtime tracking of conversation flow and tool calls, letting teams catch multi-turn prompt injection, jailbreaks, and data leakage as they unfold and block unsafe actions at runtime. Alerts can be routed to Splunk or Datadog.
One unusual aspect of Protect AI's capabilities, though, is its Recon product. Recon runs automated red teaming from a 450+ attack library.
Additionally, its huntr platform is the first AI/ML bug bounty, with 17,000+ researchers and 2,520+ CVE submissions (Protect AI).
For a poisoned model or a compromised training pipeline, Protect AI is among the strongest options here.
Palo Alto Networks completed its acquisition of Protect AI in July 2025, and it now lives inside Prisma AIRS.
But while Layer blocks malicious actions inline, as of mid-2026 it does not scope permissions to a specific action inside a connected SaaS tool. It can stop an injected or clearly malicious call but does not prevent a normal-looking agent that reads, writes, or deletes beyond what its task actually requires.
It adds no agent-identity or self-serve provisioning layer, either, so buyers lack (1) an accountable human behind an agent, and (2) the ability to grant agents scoped access.
However, Palo Alto's broader Prisma AIRS platform, where Protect AI now sits, has since added agent-security capabilities of its own.
5) Zenity: agent-native discovery and monitoring
Zenity is purpose-built for agentic AI security, and Forrester included it in its AI Governance Solutions Landscape for Q2 2025.
Observe discovers agents across SaaS (Salesforce Agentforce, Microsoft Copilot Studio), homegrown stacks (Bedrock, LangGraph, Vertex AI), and endpoints (Cursor, Claude Desktop).
Govern applies secure-by-design policy to agent configurations, permissions, and memory before deployment.
Defend then analyzes an agent's full execution path at runtime (tool calls, memory access, and data flows) to catch prompt injection and intent hijacking. Blocking unsafe actions inline rather than only alerting on them (Zenity).
Zenity has SOC 2 Type II, ISO 27001, and ISO 27701 attestations.
As of mid-2026, Zenity's own materials describe its attribution as agent- and application-level.
Zenity blocks unauthorized tool calls and API invocations at runtime, but does not bind each action to a directory-provisioned employee's in-app permissions the way an identity and access platform does.
6) CalypsoAI (F5 AI Guardrails): government and regulated inference defense
CalypsoAI brings deep national-security pedigree. It has worked with U.S. federal agencies including the Department of Defense and the Department of Homeland Security, and reaches FedRAMP and IL5 environments through Palantir's FedStart, though it is not itself FedRAMP-authorized. It holds SOC 2 Type I and Type II attestations.
It defends the inference layer in real time against injection, data exposure, and policy violations. It also runs agentic red teaming and centralizes audit logging for compliance (F5 AI Guardrails).
F5 acquired CalypsoAI in September 2025, and it now ships as F5 AI Guardrails.
As an inference firewall, as of mid-2026 it has no agent-identity model and applies role-based (not app-aware) permissions.
It surfaces AI usage inline through the network rather than through endpoint agents. That catches every prompt and response routed through the inference gateway, but it is blind to AI a user reaches outside of that path, like a public model called straight from a personal or off-network device.
7) Operant AI: runtime, endpoint, and gateway defense in one
Operant AI is the rare tool combining a dedicated MCP gateway, agent runtime protection, and endpoint discovery in one platform.
Gartner names Operant a Featured Vendor across five AI-security research notes, none of them a Magic Quadrant.
- Endpoint Protector finds shadow AI and MCP servers on employee machines.
- Agent Protector adds action-level tracing, inline blocking, and automatic redaction (PII, PCI, and PHI) across roughly 100 data types for cloud agents.
Operant provisions its own platform users through Okta and Entra ID (SSO, SCIM), but as of mid-2026 its agent permissions are governed by Operant's own controls rather than inherited from your identity provider the way an access platform does.
It scopes which tools and intents an agent can use, although it cannot be verified through its public materials (as of mid-2026) whether it allows for per-operation read-versus-write limits inside a single tool.
8) Willow: agent identity and app-aware permissions governance
Willow is the Agentic Access Platform for AI agents. It enables scaling AI agents across your organization, with agent identities tied to real employees and app-aware permissions enforced at the moment of action.
It powers companies like Wix, Innovid, and Riskified. At Wix, nearly 5,000 weekly users are enabled with secure, governed AI agent use across ~600 connected tools and MCPs, amounting to 300,000+ tool calls every week (Wix case study).
Willow’s agent identity implementation means that every AI agent inherits its identity from a real person. All through your existing identity provider. Roles and permissions (down to granular in-app action-level permissions) flow through to every action the agent takes (Willow Identity & Access).
This works through the Okta, Entra ID, and JumpCloud identity providers your org already runs.
While other tools only control which apps an AI agent can access, Willow adds an additional layer, enabling control of what an agent can do once inside apps like Jira (read tickets, create them, reassign, or delete) and in which projects (Willow Identity & Access).
Willow also ships with:
- Native shadow AI discovery
- Integrations with Splunk, Loki, and Grafana
- Immutable audit trails (a tamper-proof log of every agent action)
- One-click revocation across every agent touching a system
(Willow Governance & Compliance).
It is SOC 2 Type II certified and can be deployed as SaaS, self-hosted, or on-prem with full isolation for regulated industries.
It’s the best fit for teams that need least-privilege enforcement (each agent gets only the access it needs) on what agents can do.
When something goes wrong and security asks, "Could that agent have touched our customer data?" you pull the Audit Trail.
Willow does not cover model supply-chain scanning or pre-release red teaming. Pair it with a model-security platform like Palo Alto's Prisma AIRS, which now includes Protect AI, to fill that gap.
How these AI agent security tools compare by category
Compare them one category at a time.
The table below places each tool in the category it leads and names what it primarily secures.
Use it to spot which categories your current stack already covers and which it leaves open.
| Tool | Category it leads | What it secures | Key gap | Best for |
|---|---|---|---|---|
| Prompt Security (SentinelOne) | GenAI surface coverage | Employee AI, prompts, code assistants, MCP traffic | Agent identity; action-level permissions | Wide single-view GenAI surface |
| Lasso Security | Behavioral intent detection | Full execution traces, behavioral intent | Action-level permissions inside tools | Behavioral analysis across execution traces |
| Lakera (Check Point) | Inference firewall | Prompt injection, jailbreaks, data leakage (pre-model) | Agent identity; MCP gateway ops | Hardening live inference paths |
| Protect AI (Palo Alto) | Model supply-chain | 35+ model formats, ML pipelines, AI bug bounty | Agent identity; action-level permissions | Pre-release model scanning |
| Zenity | Agent discovery and monitoring | SaaS agents, homegrown stacks, endpoint AI | Directory-bound identity; native in-app permissions | Mapping agent sprawl across SaaS |
| CalypsoAI (F5) | Government and regulated inference defense | Inference, policy violations, agentic red teaming | Agent identity; app-aware permissions | Federal and regulated enterprise |
| Operant AI | Runtime + endpoint + gateway | MCP gateway, cloud agents, shadow AI on endpoints | IdP-inherited agent permissions; in-app read/write | Endpoint + runtime + gateway in one |
| Willow | Agent identity and app-aware permissions | IdP-linked identity, action-level permissions, audit trails, MCP gateway, shadow AI on endpoints & in browser | Model scanning; red teaming | Identity, permissions, and audit for agents |
Most tools cluster around detection, inference defense, and discovery. The layers that watch and react.
Identity and action-level permission enforcement is the thinnest column, which is why a full stack usually pairs a detection tool with an identity and access layer like Willow.
How should an enterprise team choose an AI agent security tool?
Choose by your loudest, highest-risk gap.
Run your stack against the five categories and buy for the empty column.
(1) If you don't know what agents are running
If you want to surface an agent your team spun up six months ago, still running on a developer's personal API key, start with discovery and monitoring (Zenity, Operant AI, or shadow AI discovery with Willow).
(2) If agents act dangerously inside tools they can reach
If the idea of "delete config" and "drop the database" sitting behind the same open door worries you, then you need identity and app-aware permissions. For that, choose Willow.
The damage almost always comes from an authorized agent doing an unauthorized thing inside a tool it was allowed to reach.
(3) If your exposure is the live request path
This is where zero-click attacks like EchoLeak land. EchoLeak (CVE-2025-32711) exploited Microsoft 365 Copilot to exfiltrate data through a crafted email, with no user interaction.
An inference firewall or runtime detection tool (Lakera, Lasso, CalypsoAI) hardens prompts, responses, and execution traces in real time.
(4) If you ship homegrown models or AI apps
Catch poisoned models and injection flaws before release, before a bug bounty researcher or an attacker does it for you.
For model supply-chain scanning and red teaming before release, choose Protect AI, now delivered through Palo Alto's Prisma AIRS.
Several leaders here were acquired in 2025, so "buying the tool" increasingly means buying into a larger security platform.
Weigh that cost against a best-of-breed specialist for the one category you most need.
Most teams end up with two: a detection or discovery tool, plus an identity and access layer like Willow that controls what agents can do.

Best MCP Gateways for Enterprise AI Teams in 2026
An enterprise standing up AI agents this year runs into the same wall: dozens of agents, each wired to internal tools, with no single place to say what any of them is allowed to do or to see what they already did.
The fix has a name now, it’s called an MCP gateway.
If you’re in the market for one, this is a buyer's guide for you.
Willow builds an identity and access layer for AI agents, so we have a stake here, and we will say plainly where Willow fits and where it does not.
But the field holds strong tools built by serious teams.
The aim is to help an enterprise pick the right MCP gateway for its specific situation, even when that turns out not to be Willow.
| Your priority | Best pick | Why |
|---|---|---|
| Overall governance depth | Willow | Action + context layer, identity-native |
| MCP threat detection | Runlayer | ToolGuard/ListGuard, 18k+ vetted servers |
| Fastest compliance approval | MintMCP | SOC 2 Type II + HIPAA, BAA available |
| Already on Cloudflare | Cloudflare | Native edge, FedRAMP Moderate |
| Prompt injection defense | Archestra | Lethal Trifecta guardrail, AGPL-3.0 |
| Full open-source lifecycle | Obot | Kubernetes-native, GitOps, built-in chat |
| MCP + LLM gateway in one | Lunar MCPX | Dual traffic, SOC 2 at Enterprise tier |
| Multi-server routing | MetaMCP | One entry point, aggregates any MCP server |
If you’re new to the terminology or the solution, an MCP gateway is a single control point that sits between every AI agent and every tool that agent connects to.
MCP stands for Model Context Protocol. It’s an open standard that lets an agent discover and call tools in a uniform way. Similar to an API.
It’s built to move data between an AI agent and a tool. But nothing more. The base protocol enforces no authentication or authorisation by default. HTTP transport carries optional OAuth, but most MCP deployments skip it.
This leaves significant governance gaps, and a large space for things to go wrong in your org. The risk compounds because AI agents can make sweeping changes, edits, and removals across an entire tech stack in seconds.
An MCP gateway is the layer that:
- Authenticates each connection
- Decides what the agent is allowed to do
- Records every call
Enterprises deploying AI agents at any meaningful scale need some form of this to deliver AI safely, within the limitations of their regulatory environment.
What is an MCP gateway, and why does an enterprise need one?
An MCP itself simply standardizes the wiring between AI agents and apps. By default, the protocol authenticates nothing and keeps no record of what ran. HTTP transport carries optional OAuth, but most deployments skip it. A malicious tool has no checkpoint.
In short, it adds all of the capability to AI systems, but little of the governance. That left enterprises with full capability and no governance, which is the problem MCP gateways were built to solve.
An MCP gateway is the control panel for every connection between your AI agents and their tools. It decides who can connect, what they can do, and keeps a record of everything that ran.
Without a gateway, every agent-to-tool connection is wired with credentials scattered across config files. There is no central place to see or stop anything.
The sprawl is already large, and growing.
Non-human identities (the software accounts that agents, service accounts, and API keys run as) already outnumber human ones 45 to 1 on average. Up to 144 to 1 in cloud-native environments (Cloud Security Alliance, May 2026).
Wire ten agents to ten tools by hand and you have a hundred ad-hoc connections, each with its own auth and its own blind spots.
A gateway collapses those hundred connections into a hub-and-spoke model. Every agent connects to one governed entry point, and the gateway connects to the tools.
Enforcing policy and maintaining reliable audit trails and logs is now possible.
Two main forces are driving MCP gateway adoption
While the core reason for deploying an MCP gateway is to bring some level of control to AI agent workloads and their connections to external tools, there are two other driving forces compelling enterprise adoption.
One is shadow AI, the other, regulation.
Shadow AI is when your employees are running agents and tools nobody approved. One in five organizations has already suffered a breach caused by it (IBM Security / Ponemon Institute, 2025).
Even so, 97% of organizations hit by AI-related breaches lacked proper AI access controls (IBM Security / Ponemon Institute, 2025).
Shadow AI looks like your employees using AI in browsers, on their personal devices, or linked to personal credit cards for company work.
Without the right tooling, you may never know it is happening.
Hence the name, shadow AI.
Companies with high shadow AI use had breach costs averaging $670,000 more than organizations with little or no shadow AI use (IBM Security / Ponemon Institute, 2025).
Regulation is the second driving force, and only magnifies the need to gain visibility and control over shadow AI use.
The EU AI Act's transparency obligations (Art. 50) apply from 2 August 2026. Record-keeping, logging, and human oversight for Annex III high-risk AI systems apply from 2 December 2027 (Gibson Dunn, 2026).
That means audit-grade logs are no longer a nice-to-have for any AI agent touching regulated data.
And you have two categories of agents touching that data. One you know about. And the other, in the shadows.
For most enterprises the gateway is now a must-have. The open question remains which one, and based on what criteria.
How to evaluate an MCP gateway: the criteria that separate them
Six key things decide which MCP gateway is right for your organization.
The decision rests on your needs and existing tooling across:
- Identity
- Access-control depth
- MCP-specific threat handling
- Audit and compliance
- Deployment flexibility
- Open source versus commercial
What separates vendors more than any other aspect is how deep their control goes and what they secure against.
Does the gateway give each agent a real, accountable identity?
The strongest gateways tie every agent to an accountable identity. No anonymous keys.
An agent running on a shared API key or a free-floating service account that nobody can trace back to a person is a governance issue. There is no accountability. Nor is there any possibility to decouple an agent from its owner when they leave the company, because you simply don’t know who is responsible for it.
Secure implementation looks like an AI agent that inherits its identity from a real employee through the identity provider (the system like Okta or Microsoft Entra ID that already answers "who is this and what are they allowed to do") your organization already runs.
Inherited identity and permissions is what makes accountability and clean offboarding possible. It is the most useful question most buyers never ask a vendor.
How deep does the access control go?
There are levels to access control. You can block certain apps, or block certain actions within those apps. The first level results in two kinds of errors (1) employees and AI agents lack access to tools they need, or (2) employees and AI agents have too much access, enabling them to cause serious irreparable damage.
Many providers stop short at the first level here, failing to enable granular permissions within individual apps.
Ask your vendor which category they fall into:
- Connection-layer control answers "can this agent reach this tool at all?" Almost every gateway does this, allowing or blocking a whole tool, like granting "Jira access."
- Action-layer control answers "what can this agent do once it is inside the tool?" with app-aware permissions granting "read tickets in Project X" instead of blanket Jira access.
Action-layer control is rare, and it is the layer that matters most. The costly incidents are almost always authorized agents doing something inside a tool they should never have been allowed to do.
Does it secure against MCP-specific attacks like tool poisoning?
The signature MCP attack is tool poisoning, when malicious instructions are hidden in a tool's description rather than its output.
Your AI agent connects, loads each tool's metadata into its context window (the working memory the model reads before it acts), and inadvertently ingests a poisoned description.
It happens before any tool even runs. All you (or your AI agent) need to do is install the tool.
MCPTox benchmark tested 45 live MCP servers and 353 real tools against poisoned descriptions (MCPTox, arXiv 2508.14925). They found:
- The highest refusal rate was under 3% (that’s the ceiling for how good AI agents are at refusing contaminated instructions)
- More capable models are more susceptible to attacks because they read them as legitimate instructions
- Such attacks had up to a 72.8% success rate in poisoning AI agents
Ask whether a gateway inspects tool definitions, not just prompts.
What does a real audit output look like?
Most gateways log tool calls at the connection layer and surface them in a dashboard. For a compliance audit, you need a record tied to a named employee, showing who authorized the agent, what it was permitted to do, what it actually did, and when.
The right setup streams tool-call events to Splunk, Loki, or Grafana and ships pre-built exports for SOC 2, GDPR, HIPAA, and ISO 27001.
With EU AI Act Annex III logging requirements set to apply from December 2027 for high-risk AI systems, months of unattributed agent activity is a gap worth closing (Gibson Dunn, 2026).
What deployment options does the platform support?
If your regulatory environment requires data to stay in a specific jurisdiction, some SaaS-only options may be off the table.
Lunar MCPX runs on shared cloud for the free and Pro tiers and self-hosted Kubernetes for Enterprise (Lunar pricing). Archestra and Obot are fully self-hosted. Willow supports SaaS, self-hosted, and full on-prem or air-gapped (Willow platform).
Open source or commercial: which fits your requirements?
Open source gives full code visibility and no vendor lock-in. But patching, scaling, incident response, and more all fall to your team.
Four tools in this list offer open source solutions. These are Archestra (AGPL-3.0), Obot (MIT), Lunar MCPX (MIT), and MetaMCP (MIT).
AGPL-3.0 requires any modifications be contributed back to the project. MIT does not. If you are building internal tooling on top of whichever gateway you pick, that governs what you can do with the code.
None of the four in this list have published SOC 2 Type II, HIPAA, or ISO 27001 certifications for their open-source tiers. Lunar Enterprise is an exception (SOC 2 Type II, HIPAA, and PCI DSS) but that tier is paid.
The best MCP gateways for enterprise AI teams in 2026, compared
| Tool | Action-layer control | MCP threat detection | SOC 2 Type II | HIPAA | Self-hostable | Open source |
|---|---|---|---|---|---|---|
| Willow | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ |
| Runlayer | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ |
| MintMCP | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ |
| Cloudflare | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ |
| Archestra | ✗ | ✓ | ✗ | ✗ | ✓ | ✓ |
| Obot | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ |
| Lunar MCPX | ✗ | ✗ | ✓† | ✓† | ✓ | ✓ |
| MetaMCP | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ |
† Enterprise tier only.
The best MCP gateway depends on your priorities, existing tech stack, and AI agent workloads. As well as your industry and specific regulatory requirements.
No single tool wins on governance depth, MCP-specific threat detection, open-source control, and platform breadth at once.
Willow: the Agentic Access Platform
Willow is an Agentic Access Platform. The MCP gateway is one component of that. Alongside it sit shadow-AI discovery, runtime guardrails, and a self-serve portal for the whole org.
Pick Willow when the priority is governing both which tools agents can reach, and what agents do inside a permitted tool.
Willow’s MCP gateway:
- Sits between every agent and every tool
- Handles authentication at the connection layer
- Enforces permissions at the action layer
- Logs every tool call
With it, you can route any MCP-compatible agent (Claude, Cursor, ChatGPT, and more) through one governed entry point (Willow Tools & Skills). The catalog covers 100+ governed connectors for enterprise tools like Jira, Salesforce, and GitHub, with 1,000+ integrations in the broader marketplace.
A self-serve portal means the whole org can connect their agents to approved tools without opening a ticket.
Shadow-AI discovery is native through endpoint sensors and a Chrome extension for browser AI use, covering important surfaces some other MCP gateways miss.
Browser use in particular is a common way shadow AI shows up in organizations, with employees using personal subscriptions to popular AI tools.
While most gateways control the connection, Willow enforces a three-level model, enabling you to define and control permissions at the connection level and the action level.
It gives you the power to decide not just "can this agent access Jira" but also "what can it read, create, reassign, or delete, and on which data" (Willow Identity & Access).
Most enterprises running AI agent workloads at any meaningful scale need some form of this granular permissions infrastructure.
Setting it up is made easy, because each AI agent inherits identity from a real employee through your existing identity provider (Okta, Entra ID, or JumpCloud).
SCIM (the standard that auto-provisions and offboards accounts) is built in, meaning you can revoke agent access the moment the employee leaves. That avoids one of the biggest pain points in AI enabled enterprises today.
The employee leaves, but their AI agents and MCP connections keep running, and workloads accumulate errors. Credentials stay live, and Ex-employees can sometimes re-enter production systems through connections that were never closed.
Willow’s audit logs integrate with Splunk, Loki, and Grafana. Pre-built exports cover SOC 2, GDPR, HIPAA, and ISO 27001.
It is also one of the most flexible in terms of how you deploy it within your company. Choose from SaaS, dedicated cloud, on-prem, or air-gapped (Willow platform).
Wix runs Willow across ~5,000 employees and ~600 tools, producing 300,000+ tool calls per week, every one tied to a real identity. Asaf Yonay, Head of AI Core at Wix attributes Willow to their successful AI adoption, saying “We are six to ten months ahead of most companies in AI adoption. More code to production, fewer incidents.” (Wix case study)
Willow is not without its limitations, though. The core platform is proprietary, not open source. Though it does have a free tier, teams that require a fully inspectable, self-owned stack will look at the open-source entries below instead.
Runlayer: MCP-specific threat detection
Runlayer is the strongest choice if real-time MCP threat detection is your top concern.
Its ToolGuard and ListGuard detectors run semantic analysis on MCP server metadata and tool definitions to catch:
- Tool poisoning
- Command injection
- Supply-chain attacks at the connection layer
- Prompt injection through tool schemas
The Runlayer team curates a catalog of 18,000-plus vetted MCP servers.
The platform is SOC 2 Type II, HIPAA, and GDPR compliant. If the threats you’re seeking to protect against are tool poisoning and supply-chain attacks, Runlayer is the right pick. It controls which tools an agent can reach but not what the agent can do once inside an app or tool like Linear, Datadog, AWS, or Supabase (Willow vs. Runlayer).
MintMCP: the compliance-first choice
MintMCP’s Virtual MCP architecture enables role-based permissions and curated tool sets per team.
When it comes to onboarding (and offboarding) employees, it is synced by SCIM meaning that when an employee leaves, their AI agents can be decommissioned smoothly (MintMCP).
MintMCP’s Agent Monitor traces every tool call made, and can be connected to coding agents like Claude Code and Cursor (MintMCP Agent Monitor).
The founding team includes Jiquan Ngiam and Vijay Vasudevan, who helped build TensorFlow and Coursera at Google (MintMCP About).
MintMCP is both SOC 2 Type II and HIPAA compliant (MintMCP Trust).
If SOC 2 Type II and HIPAA certification is what procurement needs, and you don't yet need to control what an agent does inside a specific tool like HubSpot or Stripe, then MintMCP is worth evaluating (Willow vs. MintMCP).
Cloudflare: the platform-scale option
For enterprises already running on Cloudflare's global network, extending into MCP governance is a natural addition to an existing deployment.
It is a strong fit for teams already building on its edge compute tooling.
Its AI Gateway, MCP servers, and Agents platform run on its powerful global network spanning 330-plus cities.
Managed OAuth enables agents to authenticate on behalf of users without insecure service accounts (Cloudflare Agents).
Cloudflare ships with SOC 2 Type II, ISO 27001, PCI DSS 4.0, and FedRAMP Moderate (Cloudflare compliance).
What it is not, however, is an agent-governance layer.
Cloudflare's MCP portals include mandatory IdP authentication and per-tool allowlisting, but stop short of per-action permission scoping inside tools.
Because Cloudflare is gateway-rule-based rather than endpoint-based, shadow MCP detection is limited.
If you’re already shipping with Cloudflare, it will probably account for a few of your requirements. Though you’ll likely need to layer in an additional solution or two to get the full coverage you need (Willow vs. Cloudflare).
Archestra, Obot, Lunar, and MetaMCP: the open-source field
Open-source gateways are the right call when self-ownership and inspectability matter most, outranking the benefits of managed compliance.
The four leading open source MCP gateway projects differ significantly in maturity and focus. There is also risk of projects being sunset, and no longer maintained. But if you’re already building on top of open source, you’re accustomed to this.
Archestra
Archestra stops prompt injection at the environment level before the model ever responds.
Its guardrails target what it calls the Lethal Trifecta, the three preconditions that combine to turn a prompt injection into a data exfiltration event (1) private data in scope (2) untrusted external content in the context window, and (3) an active outbound channel.
When all three line up, the guardrail fires before any model evaluation happens.
OAuth on-behalf-of flows run each tool as the authenticated user so every tool call traces back to a real employee.
The code is AGPL-3.0 licensed with full source visibility (Archestra). This means any modifications to the code must be released publicly (the project can't be quietly forked into a proprietary product.) For this reason, many enterprise companies flag AGPL as requiring legal review before use. Google bans its use outright.
For teams seeking prompt injection and data exfiltration prevention (and who can run the infrastructure themselves) Archestra is worth exploring.
Concept Ventures backed it with a $3.3M pre-seed so they have funding to continue development (Concept Ventures).
Obot
Obot covers the full MCP lifecycle in one platform (hosting, registry, gateway, and a built-in chat client.)
It is MIT-licensed and Kubernetes-native, built with GitOps workflows by the founding team behind Rancher Labs (acquired by SUSE) and Cloud.com (acquired by Citrix). Mayfield and Nexus Venture Partners put $35M behind them at seed (Obot).
Production deployments require Kubernetes, and access control is server-level RBAC. Teams already running Kubernetes get a platform that fits into existing GitOps and CI/CD pipelines without additional infrastructure overhead.
The built-in chat client extends the platform to reach non-engineers without handing anyone direct access to platform configuration, so governance covers every team running agent workflows (Willow vs. Obot).
Lunar MCPX
Lunar pairs two layers in one product. A tool-governance gateway for MCP traffic and an AI gateway for LLM traffic.
Its MIT-licensed free tier runs on a shared cloud for up to 50 users. Being MIT licensed means enterprises can modify and deploy without open-sourcing changes.
Access control at the tool level, rather than the in-tool action level. That means no control of what agents actually do inside apps like Figma, ServiceNow, Workday, HubSpot, or Stripe.
The Enterprise tier adds self-hosted Kubernetes, SSO, full RBAC, audit trails, and HashiCorp Vault integration (Lunar).
Boomi announced its intent to acquire Lunar in May 2026 (BusinessWire, May 2026). If the deal closes, Lunar becomes part of a larger integration platform vendor.
MetaMCP
MetaMCP aggregates multiple MCP servers into a single interface. Agents reach any server through one entry point, without per-agent server configuration.
It organizes connections using a Servers-to-Namespaces-to-Endpoints model and even ships its own GUI for managing server connections (MetaMCP).
MetaMCP provides routing infrastructure. It solves the multi-server aggregation problem with one entry point connecting agents to any MCP server, without per-agent configuration overhead. Identity management, audit depth, and compliance coverage will require a separate governance layer on top (Willow vs. MetaMCP).
The gap most MCP gateways share, and why it matters
The pattern across every vendor above is the same: connection-level control is standard; action-level control is rare. But very few control what the agent can do once inside.
Most costly agentic incidents trace to authorized agents doing something inside a tool they were never meant to touch.
If access is connection-level and both "read config" and "drop the database" sit behind the same open door, your governance strategy is simply hoping the AI agent doesn’t do anything it shouldn’t. Or an employee doesn’t direct it to do something it shouldn’t.
App-aware permission is the only way to keep the door open (and benefit from AI productivity gains) but remain protected against high-risk actions.
When defining your agent access permissions, ask three questions:
- Which tools can the agent connect to? The connection layer, the one almost every gateway already handles.
- What can it do inside each one? The action layer: read versus write versus delete, scoped to the task.
- Under what conditions, on which data, with whose approval? The context layer, where high-stakes actions pause for a human sign-off before they run.
And find an MCP gateway and AI agent governance layer that enables you to implement the permissions you need. Rather than letting a platform’s capabilities dictate how secure your AI agent systems can be.
Below is a quick decision-making rubric to help you narrow in on the right solution for you.
| Your situation | Best pick |
|---|---|
| MCP threat detection is the priority: tool poisoning, supply-chain attacks | Runlayer |
| Already on Cloudflare edge, need managed OAuth without adding a vendor | Cloudflare |
| Compliance checklist gates procurement: SOC 2 Type II + HIPAA required | MintMCP |
| Full self-ownership and code inspectability matter most | Archestra / Obot / Lunar / MetaMCP |
| Governance depth: what agents do inside tools, provable to auditors | Willow |
If you want to see app-aware permissions in practice, Willow's platform overview walks through the three-level model end to end (Willow platform).
.png)
How a HubSpot Super Admin Runs Claude on Her CRM Every Day
Connect an AI agent to HubSpot: How a SalesOps Lead Fired Her BI, Commissions, and Forecasting Tools
It’s Monday morning. Hila, the Sales Operations Manager at Agora (a real estate investment management software), just turned on her computer. Ping. Ping Ping.
Her Slack channel fills up with exactly what she needs. Every closed-lost deal from the last seven days, the exact explanations why (with actual data, not just the Account Executive’s generic reasoning), and overall trends. The report shows just the deals her team owns. No sensitive data was left exposed in the making.
She did not have to dig for answers. An AI agent did everything while she was sleeping. All she did was connect an AI agent to Hubspot.
Her chosen agent? Claude.
Her connector? HubSpot.
The result? She nixed her BI, commissions, and forecasting tools. Getting rid of the forecasting tool alone has since saved the company $15,000/year. She has everything she needs to make informed business decisions, without having to go find it.
Want to know how she did it? She tells all.
Most importantly, she shares how she did it without compromising the company’s entire CRM. The Claude HubSpot connector is scoped to the records she owns.
What RevOps teams actually do with HubSpot and AI agents
"I use Claude every day for everything,” Hila tells us from the jump of our conversation. With a Claude HubSpot connector, she is able to manage real workflows. Leading Sales Operations at a top-tier real estate investment company means she needs to make sure all data is secure. She can confidently say it is.
If you’ve ever asked, “What can AI agents do in Hubspot?” let’s take a look at a few of her most pivotal use cases:
- Closed-Lost Analysis in Hubspot with AI
Cadence: Scheduled weekly
Running as a background agent, Claude pulls all closed-lost deals from the last 7 days via a secure Hubspot AI integration. All context is provided, including: internal/external labels, stages, sources, owners, and team hierarchy. Instead of one agent doing everything (a costly time suck), the work is split across sub-agents to avoid a back and forth loop. One agent gets deals, one aggregates lost reasons, and one runs the analysis. At the end, she gets a cohesive summary on Slack with deals, reasons, and trends.
- AI Deal Hygiene
Cadence: Scheduled Weekly
The prompt reads: "Give me all the deals of every AE that needs cleanup." Claude pulls each Account Executive’s deals and checks against her defined parameters (close data wrong, stage drift, amount mismatch, etc.). Next, each AE receives a tailored Slack message with the necessary action. Her team receives evidence-based nudges. Managers no longer have to micromanage. Clients get the service they deserve. It’s a win-win-win.
- Analyze Gone Dark Deals HubSpot
Cadence: Ad hoc
The Account Executive label deals as “Gone Dark,” with no extra information provided. Hila wants to understand the reasoning. In her words, "Just because people write ‘gone dark’ as the loss reason doesn't mean that's true."
Did the AE stop responding? Did the prospect stop responding? Was it both or something else entirely? Claude pulls the deals, cross-references every engagement, and classifies the actual reason why. The reason for the loss has evidence now. There is no BI tool or analyst required. With a CSV file and Claude, the truth becomes clear in seconds. This information prompts action.
- Self-Service Pipeline and Forecasting for Finance
Cadence: On-Demand
She no longer needs to ever ask an AE to "send me the pipeline.” Instead, she built a role-specific HubSpot skills for finance that knows exactly what information to ask for. It interrogates the requester (time period, scope, pipelines), queries HubSpot, and returns an exportable CSV.
As she explains, “"I taught it exactly what properties are important for bookings, for closed won, for open, and the Skill asks them, what's your time period, what's the scope, what are the pipelines, and then it pulls them the deals from HubSpot and makes an exportable CSV." She Vibe-coded a forecasting tool, canceled the SaaS they were using, and, “Now, we get to save $15,000 a year.”
- Subscription Management
Cadence: On-Demand
Rather than having to manually track subscription renewals for all of Agora’s clients, she taught Claude how to map the existing line items to the correct subscriptions, identify the missing renewals, and create the assumed auto-renewal records.
Being able to update HubSpot records through an AI connector saved her countless hours of manual work. She shares, “That connector is also scoped specifically to me as an Operations user, since record-editing access is limited to only a select group of people.”
- Pro Tip: The CSV Shortcut
When she already knows the deal set, she exports a HubSpot list to CSV and uploads it, instead of making the agent crawl the CRM. "Claude reading a CSV file is so much more efficient than Claude going into HubSpot and trying to narrow down everything."
The Takeaway
Every critical workflow has been automated by giving the AI agent HubSpot access, securely. All results are pushed directly to Slack, removing the need for people to ask questions. The answers are already waiting. Best of all, every action is governed and traceable.
The truth is that every department needs different levels of access in HubSpot. With a HubSpot connector, managers and leads can set up specific scopes for each respective team and said access will only apply to its relevant users.
She’s a Sales Operations Llead that has indirectly transformed into a BI tool, coder, apps developer, all without ever having to write a line of code. As she says, "I'm trying to surface information to people without them having to look for it." Using a Claude HubSpot connector made it doable, easy, and scalable.
Where it stalls (the honest part)
Using HubSpot AI automation for RevOps has real results. But, it doesn’t always go as planned.
Agents on HubSpot can break or get blocked.
Here are a few lessons she learned (and shared), so you don’t have to learn the hard way:
- Reaching Limits
This was a use case where Hila’s Hubspot MCP server was failing. Running daily, the connector was supposed to cross-reference three data sources: Avoma meetings from the last day, HubSpot emails, and a list of partner domain emails. If a meeting participant matches the partner domain, it is to log the AE/Customer/Partner trio and scan new-business deal emails for the partner domain activity. If the trio hadn’t been logged in 30 days, it is to push an alert to the Partnership Team’s Slack channel to take action.
However, Hubspot returned “results too large” on a 7-day email pull. "It actually just failed. The HubSpot results are too large when I'm trying to look at the emails from the last seven days." Along with that constraint, the Avoma meetings are named by the AE, so there are naming inconsistencies.
- Choosing the Wrong Tool for the Job
Bulk writes are not the right job for a HubSpot connector. Her team was using Claude to standardize hundreds of records to fix date formatting inconsistencies (US/EU data flip). Claude fixed the CSV file and then pushed updates to the HubSpot connector. But, it took hours and burned heavy token usage as it batched 10 records at a time.
Instead, a direct CSV import to HubSpot could’ve taken 30 seconds. The connector is not the right tool for bulk writes. Had she had token visibility, she would’ve known this right away and been able to prevent wasting resources.
- Exposing Sensitive Data
The connector may be the right key, but who you give the key to matters most. A single Hubspot API key exposes the entire CRM and marketing database to the AI agent. In this case, it is Claude.
Nobody should be able to grant an agent organization-wide CRM access. It puts sensitive data at immense risk. In most cases, organization leaders don’t even know when it’s happening (this is the problem of Shadow AI).
How to connect an AI agent to HubSpot safely
To connect AI agent to HubSpot securely, make sure it is per-user and per-object scope at runtime. Every action should be audited to a human. Every prompt should be behind a guardrail.
Hila did it with Willow, and so can you.
Willow is a robust Agentic Access Platform for every AI agent. It’s a control plane. Not a tool. Not an MCP gateway.
With Willow, you get secure HubSpot AI integration:
- Least privilege at runtime: The agent touches only the deals and contacts the rep already owns.
- Scoped: Instead of giving the AI agent a blanket token with full access to read/write/delete, the agent is granted on-demand permission that is restricted. This makes it so it can only work on the specific prompt. Permissions expire once the job is done.
- Secure architecture: People receive only the information they need, when they need it. To adjust access, manual permission is required. It is read-heavy by default (optimized for data consumption). It is write behind an approval guard (any data creation, deletion, or modification must be approved before it takes effect).
- Auditable: Skills pulled from the Internet can carry API keys or prompt injection, leaving your entire organization vulnerable. The worst part is that this shadow AI can be running without anyone ever knowing. Willow provides a full audit trail of every action. And, every action is tied to a human user.
With Willow, you can shrink the AI attack surface, adopt AI safely, at speed, and prove AI is actually working.
Connect HubSpot the governed way
Hila is the hero of this story. Want to be the hero in your organization?
Rev up your RevOps with a single click that leads to a scoped, audited, and security-approved HubSpot connector for your chosen AI agent.
With Willow, you can make Monday mornings feel like Friday evenings.

Willow at Black Hat USA 2026: Find the AI Basecamp at Booth #5905
Black Hat is where the security industry comes to see what is actually coming next. This year, one of the loudest conversations on the floor will be the one security teams have been having in private all year: how do you let AI agents into the enterprise without handing them the keys to everything?
That is the question we built Willow to answer. And this August, we are bringing the answer to Las Vegas.
The fastest way to put AI to work is to govern it
AI agents are already in your organization. They are connected to Jira, GitHub, Slack, your databases, and your internal APIs, often on personal keys, with no audit trail and no one watching. Security teams cannot approve what they cannot see, and employees will not wait weeks for a ticket.
Willow is the control plane in between. One gateway, any agent, every tool, with the identity, least-privilege access, guardrails, and audit trail your CISO signs off on. Wix runs on Willow and calls themselves six to ten months ahead of most companies on AI adoption, with more code shipping to production and fewer incidents.
Find us at the Willow AI Basecamp, Booth #5905
We did not build another booth with a screen and a bowl of mints. We built a base camp: the calmest place on the show floor, and the one where you leave knowing exactly how to govern the agents already loose in your org.
Here is what is waiting at Booth #5905.
Live demos
Sit down, and in one click watch agent access get granted, scoped, or shut off across an entire organization. This is the part security leaders have been asking for.
The guardian game
A live AI guardian will be holding a flag, and it is not planning to lose. Talk it into giving the flag up, then defend it against everyone who comes after you. Fair warning: the guardian is sassy.
Top players each day walk away with a basketball signed by an award-winning baller.
An off-the-record dinner for security leaders
We are hosting a private dinner during the week for a hand-picked group of CISOs and security leaders. No agenda, no slides, no pitch. Just great food and the conversations that never happen on the expo floor. Seats are limited and the guest list is curated.
Meet the team
Eyal, our CEO. Shalev, our CTO and co-founder. Naor, marketing. Roi, GTM. Bring your hardest question about agent identity, least privilege, or what is really connected across your org. They can answer it.
Why this matters now
Most agents today are over-permissioned by default, and most guardrails try to catch problems after they happen. That gap is where the next incident lives. Governing AI agents is no longer a nice-to-have on a roadmap. It is the thing standing between fast AI adoption and a postmortem.
Black Hat is the right room to talk about it, with the people making these calls every day.
Come find base camp
Willow will be at Black Hat USA 2026, Booth #5905, Mandalay Bay, Las Vegas.
Pre-book a demo: https://withwillow.ai/events/blackhat-2026/
Request a seat at the dinner: https://luma.com/willow_basecamp_dinner

AI Security Posture Management: What It Is and How It Applies to AI Agent Fleets
An agent doesn't log in once. It calls APIs continuously, across dozens of tools, often without a human in the loop. It might be running a workflow in Jira, pulling data from GitHub, and sending a message in Slack – all within the same minute. And in most enterprises right now, nobody has a complete picture of which agents are doing what.
That's the gap AI security posture management is designed to close.
What AI Security Posture Management Actually Means
AI security posture management (AI-SPM) is the ongoing practice of identifying, assessing, and controlling the security risks that AI systems introduce into your environment. It borrows the "posture management" framing from cloud security (CSPM) and applies it to a new category of non-human actors: AI agents, LLM-connected apps, and the tool access they carry.
The core questions AI-SPM tries to answer:
Traditional security tooling doesn't answer these questions well. Agents don't authenticate the way humans do. They often run on long-lived API keys or shared credentials that sit entirely outside your identity provider. They can be spun up by any employee with a credit card and a browser.
Why Agent Fleets Create a Distinct Posture Problem
A single AI agent connected to a few tools is manageable. A fleet of agents – some sanctioned, some not – is a different challenge.
According to the Gravitee State of AI Agent Security 2026 report (n=750), 54% of organizations have already experienced a security incident tied to AI agents. Separately, Onyx Security reported in March 2026 that 93% of enterprises run agents with excessive permissions, and 80% expose sensitive data through them.
These numbers reflect a structural problem, not a configuration mistake. Most enterprises don't have a governed path for deploying agents – so agents get deployed ungoverned.
The Shadow AI Problem
The agents your security team knows about aren't the whole picture. Employees are connecting Claude, Cursor, ChatGPT, and other agents to internal tools without IT approval. They're building quick automations – sometimes called vibe-coded apps — that touch production systems. None of this shows up in your identity provider. None of it has an audit trail.
This is shadow AI, and it's already running in production at most organizations. Any serious posture management program has to account for it, not just the agents that went through formal procurement.
The Permissions Problem
Even sanctioned agents often hold far more access than they actually need. An agent authorized to read Jira tickets might also have write access to close them, reassign them, or delete them. An agent connected to GitHub might have repo-level access when it only needs to read a single branch.
This happens because most tools don't expose action-level permission scoping. You grant access to the app, and the agent inherits everything that comes with it. Posture management means identifying where that over-permissioning exists and enforcing tighter scopes.
The Identity Problem
Agents need real identities – not just API keys, but governed identities that tie back to your organization's identity provider. Without that, you can't enforce consistent policy, you can't cleanly revoke access when an employee leaves, and you can't answer an auditor's question about who authorized what.
Most enterprises don't have this today. Agents run on credentials stored in a spreadsheet, a loosely controlled secrets manager, or a developer's local environment.
The Core Components of AI Security Posture Management
A practical AI-SPM program covers five areas:
1. Agent inventory and discovery
You need a complete, continuously updated list of every AI agent active in your environment – including the ones IT didn't approve. That means monitoring for new MCP connections, new OAuth grants, and new API key issuances tied to AI tools.
2. Identity and authentication
Every agent should have an identity that connects to your IdP – Okta, Entra ID, or JumpCloud – not a shared service account. Identity should be provisioned and deprovisioned through the same lifecycle management that governs human users.
3. Permission scoping
Access should be granted at the action level, not the application level. An agent that needs to read GitHub pull requests doesn't need to merge them. An agent that creates Jira tickets doesn't need to delete them. Posture management includes auditing current permissions and enforcing least privilege.
4. Runtime monitoring and audit trails
Every agent action should be logged – not just "agent X connected to tool Y," but the specific calls made, the data accessed, and the outcomes. This is what makes compliance audits survivable and incident response possible.
5. Guardrails and response
Posture management isn't just visibility. It includes the ability to block policy-violating actions, flag PII exposure in real time, and route sensitive operations through a human approval step before they execute.
How This Applies in Practice
Consider a realistic scenario. Your engineering team has been using Cursor with a GitHub MCP connection for six months. Your sales team spun up a Claude-based agent connected to Salesforce last quarter. Someone in operations built a quick automation that touches Gmail and Attio. None of these went through IT.
From a posture standpoint, you have:
An AI-SPM program surfaces all of this. It doesn't require you to shut the agents down – it gives you the visibility and control to govern them without blocking the work.
Where Willow Fits
Willow is built specifically for this problem. It connects to your existing identity provider – Okta, Entra ID, or JumpCloud – and gives every AI agent in your environment a real, governed identity. Permissions are scoped at the action level inside each tool, not just at the application level. Every MCP call and agent action is logged automatically.
Shadow AI discovery is built in. When an employee spins up an unsanctioned agent, Willow surfaces it – security teams get visibility without having to go looking for it.
Guardrails run at runtime. PII protection, Slack-based approval workflows, and automatic credential revocation on offboarding are all part of the platform. Willow is SOC 2 Type II certified and supports SaaS, self-hosted, and on-prem/air-gapped deployments.
It's live in production at Wix, Innovid, and Riskified – not a pilot program.
What Good AI Security Posture Looks Like
A mature AI-SPM posture doesn't mean agents can't run. It means they run under the same governance standards applied to human users.
Security defines policy once. Employees self-serve safely. The audit trail is automatic. When something goes wrong – or when an auditor asks – you have answers.
That's the goal. Not to slow down AI adoption, but to make sure it doesn't create a security debt that compounds quietly until it does.

6 Best AI Governance Platforms for Enterprise Compliance 2026
"AI governance" now means two different jobs. Most buyers find it the hard way.
One job is governing the models your data science and risk teams build and buy. That means proving that:
- A credit model isn't biased
- A use case maps to the EU AI Act
- An auditor can trace a decision
The other job is governing the AI agents your employees are already running today.
A chatbot answers prompts. An AI agent uses a large language model to take actions on your behalf.
The difference is material. Because an AI agent can:
- Call and use tools
- Read, exfiltrate, or delete data
- Make irreversible changes
Agents can reach into your apps (Jira, GitHub, Salesforce, and Snowflake) through a growing pile of MCP (Model Context Protocol, the open standard that lets an agent connect to a tool) and API connections, or into your server or GitHub repository.
The platforms below are good at different jobs. This maps which job each was built for.
Platforms were selected based on their standing as top players in the enterprise AI governance market, with reference to the Forrester Wave: AI Governance Solutions (Q3 2025) and Gartner AI Governance Platforms Magic Quadrant (June 2026).
We evaluated publicly documented capabilities, and weighed them against the governance layers that enterprise compliance and security teams are audited against in 2026 (such as the EU AI Act, NIST AI RMF, SOC 2, and ISO 27001.)
Model governance and agent governance solve different problems

Model governance and agent governance answer fundamentally different audit questions, and the gap between them is where most compliance programs fail.
Model/policy governance asks "is this AI system fair, documented, and compliant?"
Agent governance asks "did this agent have the right to take that action, and can I prove who it was acting for?"
A platform built for one rarely covers the other well. Buying the wrong one is about more than just feature coverage, it leaves a real gap on your next audit.
That gap becomes regulatory exposure, and risk of exploitation by hackers or bad actors.
Model governance is the mature category. These platforms inventory every model and use case, run bias and risk assessments, map controls to regulatory frameworks, and generate audit-ready evidence.
Such tools include Credo AI, Holistic AI, IBM watsonx.governance, OneTrust, and Monitaur.
Agent governance is the newer layer. It's where operational risk has shifted and companies are most blind.
The questions AI agent governance helps you answer are different:
- Identity: Whose identity is this agent acting under, and does that inherit from your identity provider (Okta, Entra ID, JumpCloud) that already holds your employee accounts and group memberships?
- Permissions: Not just “can the agent reach Jira” but “what can it do *inside* Jira (read a ticket vs. delete a project)?”
- Runtime enforcement: Is something sitting inline between the agent and the tool at connection time, or are you reading traces after the fact?
- Shadow AI: Shadow AI is any AI tool, agent, or connection an employee runs without IT approval or security review. Can you see the rogue MCP servers and personal API keys on employee laptops right now?
A model governance platform tells you a model is registered and assessed. An agent governance platform stops an agent from deleting a Salesforce record it was never permissioned to touch.
Both matter and you likely need both, but the wrong pick for your primary gap shows up on the next audit.
Layer 1 (model/policy governance) proves a model is fair and documented. It includes an AI registry, bias and risk assessment, and regulatory-framework mapping.
Layer 2 (Agent governance) proves an agent acted within its scoped rights in approved environments. It includes agent identity, app-aware permissions, an MCP gateway, and shadow AI discovery.
Most platforms in this guide live cleanly in Layer 1. One, Willow, lives in Layer 2.
Why the shift matters for compliance buyers
Every major platform shift in enterprise IT created an identity gap.
On-prem applications got their identity and access layer in Active Directory.
The move to SaaS created a new gap. Hundreds of cloud apps, no central control. Okta (and Entra ID, JumpCloud) filled it with single sign-on, SCIM provisioning (the standard that automatically creates and removes a user's access as they join, move, or leave), and a single audit trail tied to a real employee.
But AI agents have nothing, and enterprises are exposed for exactly that reason.
Agents are the third wave, and right now most enterprises govern them with nothing.
AI agents also happen to be moving faster than anything before. The tooling is multiplying week over week, with a fast-moving open-source community behind it.
Closing this gap requires a speed of action and implementation that IT orgs have previously not had to keep up with, outside of perhaps cybersecurity.
Active Directory and Okta deprovision a leaving employee automatically. Agents running on personal API keys have no equivalent trigger because nothing ties their actions back to a named person.
Willow is the identity and access layer for AI agents, the Agentic Access Platform in this guide. Each agent inherits a real employee's identity from your existing IdP, gets permissions scoped to what it can do inside each tool, and is deprovisioned automatically when that employee leaves. What Okta became for SaaS, Willow is built to be for agents.
The 6 best AI governance platforms for enterprise compliance in 2026
The best platform depends entirely on whether your primary risk lives in models or in agents.
This list profiles each tool for what it governs, then names the buyer it fits.
Five are model/policy governance leaders. One, Willow, is the Agentic Access Platform.
The table below maps every platform to what it governs and which compliance frameworks it documents support for.

Credo AI: the analyst-recognized model and agent governance leader

Credo AI is the strongest pure-play AI governance platform for enterprises that need to discover, assess, and document every AI system across the org.
It centralizes an AI Registry covering models, applications, agents, and shadow AI.
Rather than give you a static point-in-time snapshot, continuous risk assessment runs across a broad set of risk dimensions.
The platform ships with pre-built policy packs for the EU AI Act, NIST AI RMF, ISO 42001, and SOC 2 with audit-ready evidence generation.
Credo AI was named a Leader in the Forrester Wave: AI Governance Solutions (Q3 2025), with the highest possible score (5/5) in 12 criteria, and landed at #6 in Applied AI on Fast Company's World's Most Innovative Companies of 2026.
For Agent heavy workloads, it comes with:
- An Agent Registry with agent cards and dependency graphs.
- A GAIA governance assistant that automates intake and control mapping.
- A public MCP server in preview that exposes the platform to customer-built agents.
However, Credo AI's runtime governance evaluates agent traces and applies policy after behavior is observed. It does not currently evaluate at connection time.
It provides no inline MCP gateway, IdP-inherited agent identity, or per-action permissions inside tools. Enforcement integration with CI/CD pipelines and API gateways is on the public roadmap but not yet shipped as of this writing.
Credo AI is the right pick if you need to govern, document, and prove compliance across a whole AI estate.
Holistic AI: strongest model risk testing with real-time agent oversight

Holistic AI is the best fit for enterprises whose top concern is bias, safety, and adversarial risk in their AI systems.
It runs 40+ specialized tests spanning:
- Bias
- Fairness
- Toxicity
- Hallucination
- Prompt injection
- Jailbreak resistance
Risk scores get mapped to the EU AI Act, NIST AI RMF, ISO 42001, and NYC Local Law 144, covering each model’s regulatory obligations.
AI discovery scans 20+ cloud and SaaS integrations to surface shadow AI and classify it by risk and owner.
On agents, Holistic AI is one of the closest model-governance vendors to runtime control. Its Guardian Agents architecture pairs Sentinel Agents (observe and evaluate every agent action against policy in real time) with Operative Agents (intervene and remediate when thresholds are crossed).
A 2026 update added tool-calling, access control, and cost control for agentic systems in production.
Two gaps matter for enterprise buyers.
Guardian Agents observe and intervene, but there's no inline MCP gateway handling auth at connection time, and the platform cannot inherit agent identity from an employee IdP.
Holistic AI is a strong choice when model risk and bias testing is the core job, and runtime agent oversight is a strong plus.
IBM watsonx.governance: broadest regulatory coverage for large regulated enterprises

IBM watsonx.governance is the platform to beat when multi-jurisdictional regulatory coverage is the deciding factor for your organization.
It supports 200+ regulatory frameworks with automated applicability, evidence collection, and audit-ready reporting. That is the deepest framework coverage among the dedicated AI governance platforms in this comparison.
It's FedRAMP-authorized on AWS GovCloud, a Forrester Wave Leader for AI Governance, and a Gartner AI Governance Platforms Leader (June 2026).
Governance Graph maps the entire AI ecosystem (assets, policies, risks, and regulatory requirements).
Continuous drift and bias monitoring are also included.
But there are three caveats.
First, the agent monitoring is recent (GA December 2025) and less mature than the model governance core.
Second, on-prem deployments require Cloud Pak for Data VPC licensing, which adds cost and complexity.
Third, there's no agent identity layer, no inline MCP gateway, and no endpoint-level shadow AI discovery.
IBM watsonx.governance fits large, regulated, IBM-ecosystem enterprises that need maximum framework breadth.
OneTrust AI Governance: unified GRC with documented MCP policy enforcement

OneTrust is the best fit for enterprises that want AI governance living inside one platform alongside privacy, data governance, and third-party risk.
It extends OneTrust's mature governance, risk, and compliance ecosystem.
MCP policy enforcement with audit logs come built in, alongside agent registration with defined purpose and enforced allowed actions.
It also announced AI agents (Privacy Agent, Third-Party Risk Agent) in September 2025 to automate governance work.
OneTrust's strength is workflow and documentation, not inline runtime defense.
Its runtime controls are policy-triggered via AI Guardrail Enforcement rather than a protocol-level inline gateway.
Its strength is compliance workflows and documentation.
OneTrust is not built with the purpose of intercepting every agent-to-tool connection before it occurs.
The platform also carries a steep learning curve for leaner teams that are unlikely to staff a dedicated GRC function.
It also lacks IdP-integrated agent identity and an inline MCP gateway.
OneTrust is the right choice for enterprises already standardized on OneTrust GRC who want AI in the same data model.
Monitaur: full-lifecycle model governance for regulated industries

Monitaur is built for highly regulated enterprises. Primarily insurance, but also financial services and healthcare.
Its platform runs:
- Define (policy templates and risk methodology)
- Manage (model and use-case inventory with a Common Controls Library)
- Automate, where "FlightSim" pre-deployment simulation grades models before release
- Record which runs continuous production validation for drift and bias
The trade-off, though, is breadth.
Monitaur lacks a dedicated agent registry, dependency graph, or runtime agent governance.
However, it is a strong choice for insurance and financial-services teams that need rigorous, auditable model risk management.
Willow: the agent governance and identity specialist

Willow is the platform to choose when the action an agent takes is the thing keeping you up at night.
Willow governs AI agents at the identity, action, and runtime layers. Three things the model-governance platforms above don't do natively.
Every agent inherits a real employee's identity through the existing identity providers (Okta, Entra ID, JumpCloud).
That includes SCIM provisioning, SSO, and auto-deprovisioning on offboarding. That makes offboarding clean: agents lose access the moment the employee does.
Willow’s greatest strength, though, is in app-aware permissions. With it you can define not only which tools an agent can reach, but also what it can do inside each one.
Define whether an agent can read vs. write vs. delete, on which data, and under what conditions.
Three more top features make Willow a strong candidate for your AI agent governance stack.
(1) An Inline MCP Gateway sits between every agent and tool, enforcing auth at the connection layer and permissions at the action layer, not reading traces after the fact.
(2) Shadow AI discovery at the endpoint uses sensors and Willow for Chrome (a browser extension) to surface rogue MCP servers, personal API keys, and unapproved agents on employee machines.
(3) Audits tied to a real employee mean every agent action is logged, timestamped, and immutable in the Logbook. It comes with pre-built SOC 2, GDPR, HIPAA, and ISO 27001 exports.
Willow sits at the identity and access layer for AI agents.
That means teams can run AI agents at production scale while security has the policy controls, the audit log, and access revocation capability.
Wix's Head of AI Core, Asaf Yonay attributes their success to Willow: "We are six to ten months ahead of most companies in AI adoption. More code to production, fewer incidents."
Across the Wix deployment, Willow governs approximately 5,000 weekly active users using ~600 governed tools. Together, they generate 300,000+ governed tool calls per week (Willow X Wix).
Willow is the right pick when employees are already running agents and you need identity, app-aware permissions, and an audit trail tied to named people.
It's also the most pricing-transparent option of those in this list. It is free for up to 5 users, $15/seat for Startup, custom for Enterprise. All SOC 2 Type II certified.
Depending on your requirements, you can choose from a SaaS, self-hosted, or on-prem/air-gapped deployment, with full feature parity across all three.
Model governance vs agent governance: a side-by-side comparison
The clearest way to choose which platform is right for you is to line up the two governance types on the capabilities that decide an audit.
The table below contrasts the model/policy governance platforms (Credo AI, Holistic AI, IBM, OneTrust, Monitaur as a group) against the agent governance/agentic access platform (Willow) approach.

Read the framework row carefully. Model governance platforms lead with EU AI Act, NIST AI RMF, and ISO 42001. Agent governance leads with SOC 2, GDPR, and ISO 27001.
That difference reflects which audit each was built to pass.
If your compliance pressure is the EU AI Act, start with the model-governance leaders. If it's SOC 2 and access control over what agents touch, start with agent governance.
How to choose an AI governance platform for your enterprise
Choose based on where your risk lives, which frameworks you're audited against, and whether you need documentation or inline enforcement.
The platforms in this guide are good at different jobs, and the wrong fit shows up as a gap on your next audit or a regulatory breach your auditors can’t trace to a named actor.
Work through the below criteria before you shortlist.
Where does your risk live: models or agents?
If it's biased or undocumented models, weigh toward Credo AI, Holistic AI, IBM, OneTrust, or Monitaur.
If it's employees running agents that reach into production systems, weight toward agent governance and Willow.
Many enterprises need both layers.
If your primary compliance pressure is the EU AI Act, ISO 42001, or model bias risk, start with a model-governance platform from this list.
If employees are already running agents against production systems like Jira, Salesforce, or GitHub, start with agent governance. If both are true, plan for two layers: model governance answers the "what did we build" audit; agent governance answers the "what did it do" audit.
Which frameworks are you audited against?
EU AI Act, NIST AI RMF, and ISO 42001 point to model governance.
SOC 2, GDPR, HIPAA, and ISO 27001 access control point to agent governance.
For maximum breadth, IBM's 200+ frameworks are hard to beat.
Do you need documentation or enforcement?
Most model-governance platforms document and assess. They evaluate traces or generate evidence.
But if you need something to block an unpermitted action inline, you need a runtime gateway. You need agent governance.
What's your shadow AI exposure?
Cloud/API scanning (Credo AI, Holistic AI, OneTrust) finds AI in your cloud accounts.
Endpoint sensors (Willow) find personal API keys and rogue MCP servers on laptops.
For most enterprises, it's both.
If your shadow AI risk lives in cloud accounts and SaaS integrations, start with Credo AI, Holistic AI, or OneTrust. They scan cloud and SaaS integrations and classify AI tools by risk and owner.
If it lives on endpoints, including personal API keys, local MCP servers, and unapproved tools running on employee laptops, start with Willow. Its endpoint sensors surface every tool in use before it reaches a production connection.
There is no single best AI governance platform
Pick the model-governance leader that matches your frameworks and ecosystem.
Add the agent-governance layer if employees are already running agents against your systems.
Most enterprises in 2026 will end up running one of each.
For architecture decisions on identity, app-aware permissions, and runtime enforcement, the Willow blog covers the agent-governance layer in depth.
Most enterprises in 2026 will end up running one model-governance platform and one agent-governance layer.
Further Reading and Sources
- Credo AI: https://credo.ai/product
- Holistic AI: https://holisticai.com
- IBM watsonx.governance: https://ibm.com/products/watsonx-governance
- OneTrust: https://onetrust.com/solutions/ai-governance
- Monitaur: https://monitaur.ai/platform
- Willow: https://withwillow.ai/platform
- Wix case study (Willow): https://withwillow.ai/blog/wix-case-study
Frequently asked questions
Enterprise compliance and security teams evaluating AI governance platforms in 2026 consistently reach the same four questions about model governance, agent governance, framework coverage, and cost.

The Willow July Digest: The Fastest Way to Put AI to Work Is to Govern It
Most of what shipped in Willow this July has the same shape. Somewhere, an admin was stuck between two bad options: block a capability outright, or hand it over and hope. Every feature below closes that gap a little further, so the answer stops being "no" or "trust me" and starts being a policy you actually set.
A few of these started as requests customers sent us directly. Here's what changed.
Roll out Claude Code from one console, not one machine at a time
Claude Code Policy Console is now generally available. Before this, locking down Claude Code across an org meant touching settings machine by machine, or waiting for the console to leave beta. Now it's one screen: MDM export, deny-list tiers, model settings, hooks, and a scenario builder, with per-OS install instructions built in. If your rollout was paused waiting for GA, it's ready now.

Delegate the busywork, not the risk
Custom org roles replace all-or-nothing admin access. Until now, giving someone admin capability meant giving them everything, whether they needed to manage skills and connectors or not. Custom roles scope access to exactly what a person's job requires, so a teammate can run day-to-day admin work while security and policy settings stay with whoever should hold them. It's enforced across every API endpoint and admin page, not just the parts of the UI someone happens to click through.

Stop choosing between block and allow
Guard rules can now pause a risky tool call and loop in a human before it runs, instead of forcing a binary decision to block it outright or let it through. A new warn-and-approve action holds the call until someone signs off, with a live notification the moment it fires. This is the same false choice we've been writing about all month: block it and you lose the capability, allow it and you're exposed. A pause is a third option, and now it's a real one.

Guard your org on Chrome, without asking every employee to install anything
The Claude Guard Chrome extension now installs via MDM. IT can push it to every managed machine in one action instead of asking each employee to install it themselves, which in practice meant partial coverage and no way to guarantee everyone was protected. Fleet-wide rollout is now the default path, not the aspirational one.

Your SIEM already watches this. Now it can watch Willow too
CrowdStrike is a supported log destination as of this month. Willow's logs flow into the SIEM your security team already has open, with a delivery-audit view and a one-click test send to confirm it's actually landing. For teams where "does it show up in our SIEM" is a hard requirement before sign-off, this closes that gap directly.

Also shipped this month
Audit logs now redact sensitive data by default, with full detail available on demand for anyone who needs it. Analytics graduated to general availability. Group-level AI Champion roles let teams delegate governance responsibility to a named owner instead of routing everything through central IT. Slack alerts now fire when shadow AI activity is discovered. And a set of reliability fixes landed across the gateway and dashboard that you should notice mostly by not noticing them.

New to watch and read
Two things worth your time if you haven't seen them yet:
- Shalev, Willow's CTO, wrote up how Claude Tag actually works, and why identity is the hard part (7 min read).
- And we put out a new video walking through shadow AI, skills, and plugins, and how to get all three under control in about two minutes (1 min watch).
The pattern, if you're tracking it
None of this month's releases are about adding a new AI capability. They're about making the capabilities you already have safe enough to turn all the way on. That's the bet behind Willow: the fastest way to put AI to work isn't to loosen the controls, it's to build controls precise enough that loosening them stops being the only way to move fast.
Questions about anything above? Reach out to us and the team would love to discuss!

Willow Launches with $7M to Build the Future of Enterprise AI Agent Governance
After a year running quietly inside Wix at the scale of thousands of employees, Willow emerges from stealth as the Agentic Access Platform for the enterprise. Hetz Ventures leads the round.
Herzliya, Israel · June 4, 2026 — Today we're announcing that Willow has raised $7 million in seed funding, led by Hetz Ventures, to build the access layer enterprises need to safely adopt AI agents at scale.
Willow is the AI Basecamp for the enterprise: a unified Agentic Access Platform where every AI agent gets a real identity, scoped access to exactly the tools its task requires, runtime guardrails, and a full audit trail tied to a human. The platform is already running in production at Wix, powering ~5,000 weekly active users across engineering, product, design, HR, finance, and legal. Deployments are now expanding across cyber security, real estate, fintech, and adtech.
This funding accelerates Willow's go-to-market and product development at exactly the moment enterprises are confronting the question they've been avoiding: who is actually using AI inside the company, with what permissions, and how would we know if something went wrong?
The problem: AI agents are running inside your organization. You probably can't see them.
AI adoption inside enterprises didn't follow the SaaS playbook. It didn't come in through procurement. It came in bottom-up.
A developer installs an MCP server on a Tuesday. Finance starts piping reports into an unmonitored tool. Sales runs an unapproved skill that touches the CRM. Someone in marketing builds a vibe-coded app and wires it straight into the company data platform, and suddenly the entire lead base is one GET request away from anyone who finds the endpoint. No ticket. No inventory. No review.
By the time security asks "what do we actually have?", the honest answer is: we don't know.
The numbers back up what every CISO is already feeling:
- 79% of enterprises are deploying AI agents. (PwC, 2025)
- 73% are running multi-agent systems. (HFS / Cognizant)
- 65% have already had an agent-related incident in the last 12 months. (Cloud Security Alliance, 2026)
Most existing AI gateways only secure what enterprises already know about. The real problem is everything they don't: agents on personal API keys, unsanctioned skills with standing access, data leaving through paths no one logs.
The category that emerged in response, AI security as an after-the-fact dashboard, is failing in two directions at once. It tells security what already happened. It tells employees only what they can't do. Neither closes the gap. Both leave enterprises one prompt away from a serious incident.
Why traditional IAM, PAM, and DLP can't fix this
The default reaction has been to bolt agents onto existing identity infrastructure. It doesn't work, and the reason is structural, not configuration.
Identity and Access Management (IAM) was built for humans and predictable service accounts. Stable identities, known sessions, access to apps and files. Agents break every one of those assumptions. They are non-human, autonomous, short-lived, and multi-tool. One agent might touch Jira, Snowflake, and GitHub in a single task, assemble its capabilities at runtime, and act on behalf of a human while making decisions no one pre-approved.
Privileged Access Management (PAM) vaults credentials for privileged humans and known sessions. Agents are neither.
Machine identity issues certs and keys for predictable, service-to-service traffic. Agents are probabilistic, not deterministic.
Legacy DLP watches the network layer. Agent risk lives at the prompt layer. By the time data shows up in a packet, it has already left through a path no one monitored.
Agentic access is a new category because each existing model solves a narrower problem. Human access assumes a person authenticates once and you trust their judgment. Agents are non-human, multi-tool, probabilistic, and act on behalf of humans while making decisions no human pre-approved. That combination requires governing the action, not just the connection. Not "can this agent reach Snowflake," but "which schemas, under which conditions, doing what."
What Willow does: identity, scope, audit, before an agent touches a system
Willow is the control plane underneath every AI agent in the enterprise. One platform that connects any agent (Claude, Cursor, ChatGPT, Codex, Gemini, n8n, custom agents) to any internal system, with the auth, scope, runtime guardrails, and audit trail enterprises actually require.
The platform does five things, on one control plane:
- Identity at the agent layer. Every agent gets a real identity inherited from your existing IdP (Okta, Entra, Active Directory, JumpCloud). No new identity model to build and maintain.
- Scope per task, not per organization. Tools are generated at runtime, scoped to exactly what the agent's task requires. Not blanket OAuth grants. Not standing access. Least privilege, enforced at the point of tool generation, before the agent acts.
- Runtime guardrails. PII redaction, prompt-injection protection, app-aware permissions, and approval workflows that fire before risky actions complete, not after.
- Shadow AI detection. A browser extension and an endpoint agent (pushed through your MDM) surface unsanctioned MCPs, skills, and agents the moment they appear, not after an incident.
- Audit trail tied to a real human. Every action streamed to your SIEM in real time. Full attribution, every time, no exceptions.
The platform also includes a marketplace with over 1,000 ready-to-use connectors, more than 100 skills, and more than 100 plugins, plus the ability to wrap any internal API as an MCP. Deploy as SaaS, dedicated cloud, or self-hosted, including fully air-gapped.
The outcome is the line we use internally: Willow turns "we can't approve that" into "it's already governed."
Proof: what production looks like at Wix
We didn't write this from a whiteboard. Willow has been running in production at Wix for a year, and the numbers from that deployment are the foundation of everything we just said.
- ~5,000 weekly active users, across engineering, product, design, HR, finance, and legal. More than the entire Wix engineering organization.
- 600+ governed tools, all behind Okta SSO with full audit and shadow-AI protection.
- 300,000+ governed tool calls every week, with zero hit to security posture.
What surprised us most wasn't the scale. It was the breadth. The moments that stuck were the ones we didn't anticipate. An office manager who used to walk hundreds of meeting rooms once a month to release the unused ones now runs a single prompt through Claude, governed by Willow, and frees every empty room in minutes. A developer who spent hours on manual data migrations now does it in one prompt through Cursor, scoped to the right systems and audited end-to-end. Hours back, every week, for people who will never write an MCP file.
"Thousands of Wix employees are using AI agents every day, and at our scale, visibility and control over those agents are absolutely critical. To accelerate AI adoption safely, we need guidelines, governance, and full visibility across the company. Willow provides exactly that."
Avishai Abrahami, Co‑Founder and CEO, Wix
Innovid (NYSE: CTV) uses Willow to govern developer machines specifically around exposure to MCP servers and external skills, getting control and reducing AI risk without telling their engineers to stop. Riskified (NYSE: RSKD) is deploying Willow in production. More are coming.
Why Hetz Ventures led the round
"The gateway between AI agents and an enterprise's internal systems is rapidly becoming one of the most overlooked blind spots in enterprise security. What convinced us to lead this round was watching Willow solve the problem inside Wix first, at the scale of thousands of employees, before bringing it to market. Eyal, Shalev, and Idan have built something rare: a governance layer that enterprises actually deploy, rather than another framework that sits on a shelf. They're the right team to define this category."
Guy Fighel, Partner, Hetz Ventures
The thesis: every infrastructure era has had its access layer
This is the part we keep coming back to.
On-prem had Active Directory. SaaS had Okta. Agents need theirs now. That is the access layer Willow is building, and it is being built right now whether enterprises choose it deliberately or assemble it by accident.
Leaders who treat agent governance as a feature they'll bolt on later, or as a tool that belongs only to the security team, will wake up with seven vendors, seven dashboards, no unified identity for their agents, and no neutral way to answer what those agents actually did across a multi-vendor fleet. They will rebuild it as one platform anyway, under far worse conditions, after an incident.
Willow is built for the other path. Choose the access layer on purpose. Govern every agent, in every tool, on behalf of every human, from one control plane.
What's next
The $7M seed accelerates three things:
- Hiring across engineering, product, and GTM. The platform team is growing, and we're investing heavily in the parts of the product that make enterprise AI actually work in production.
- Deeper platform investment. More guardrails, more shadow-AI coverage, more depth on the integrations that make Willow fit into how enterprise teams already operate.
- Expanding deployments. More enterprise customers, more verticals, more of the world's largest organizations adopting governed AI at scale.
If any of this resonates, the easiest next step is to book a 20-minute demo or explore the platform.
Join us
We're hiring across engineering, GTM, product, and design. If you want to build the access layer for the agentic era, see open roles.
About Willow (formerly Webrix)
Founded by former Wix engineers Eyal Ben Ezra (CEO), Shalev Shalit (CTO), and Idan Chetrit (VP Platform), Willow is the Agentic Access Platform for enterprise AI. The company enables organizations to securely connect AI agents to internal systems with runtime permissions, centralized controls, auditability, and full attribution of agent activity. Willow is headquartered in Herzliya, Israel.
