Case Study
How Wix scaled Al-native work to 5,000 employees with Willow
Read More
AI Agent Security

Willow vs Arthur

TL;DR

IT teams need to see which tools employees connect, control who uses them and set approval for sensitive actions. That work continues whenever someone adds a coding assistant or starts an unattended agent.
‍

Willow brings those questions into employee tool management. It combines supported device scans, approved tool connections, group-based access, browser checks and approval of configured tool calls. Employees can connect approved tools to clients such as Cursor, Claude Code and Codex, while IT manages the connections and rules. For an IT or security team choosing a starting point for employee AI, choose Willow.
‍

Arthur also covers agent discovery, security and governance. Its discovery offering includes employee-device scans, MCP monitoring and cloud connections. Its evaluation tools help application teams inspect an agent’s steps, measure answer quality, compare prompts and check requests before their application acts. Arthur’s application checks return a result that the application or gateway must act on.
‍

For employee AI management, start with Willow’s approved connections and human tool-call approvals. Choose Arthur only if application evaluation drives your purchase, and only if those employee controls fall outside your project’s requirements.

How Willow and Arthur compare

MCP, short for Model Context Protocol, gives AI apps a standard way to connect to tools. A tool might search a company system or perform an action in it. Willow manages supported connections and checks which tools a person’s groups allow them to use. It also finds supported tool configurations and instruction files on employee devices.
‍

Arthur follows agents and applications through discovery, recorded activity and evaluation. Its recorded steps show the model calls, searches and tool calls an application sends to Arthur. Developers can use those records to find where an application produced a poor result, then compare changes to its prompts. Its security checks add another step to the application’s own handling of a request.
‍

What each product governs

Area Willow Arthur
Employee devices Scans supported local tool, skill and instruction configurations. Discovery offering includes scans for locally running agents.
Access controls Checks tools available through the caller's groups. Assigns roles for Arthur organizations, workspaces and projects.
Where prompts get checked Prompt Guard checks supported messages and attachments before submission. Application checks cover requests and responses in connected integrations.
Tool checks and evaluation Applies configured rules to connected tool requests and responses. Runs continuous evaluations and prompt experiments.
Ownership and review Machine users have human owners and tools allowed by their groups. Governed applications receive owners and risk classifications and policy reviews.

‍
‍

For Willow’s device inventory, IT installs the Scan Agent. Supported configurations can appear in the dashboard even when their tools never pass through Willow’s gateway, the service that handles connected tool calls. Discovery follows supported formats and scan cycles, so the inventory develops from the devices that report to Willow.
‍

Arthur’s connected cloud discovery workflow polls cloud systems linked to Arthur. Its broader discovery offering also includes employee devices, network activity and MCP monitoring. Teams can register discoveries, assign owners and apply checks to governed applications. Confirm which discovery methods come with your plan and demonstrate them on your devices.
‍

How each product handles identity and access

Requirement Willow Arthur
Company sign-in Supports providers including Okta, Entra ID and Google Workspace. Enterprise supports compatible company sign-in providers.
Group permissions Groups grant access to connected tools; grants add together. Roles apply to Arthur resources at organization, workspace and project level.
Credentials handled Supports separate user sign-in flows or user-supplied keys for configured connections. Stores model-provider credentials for evaluation and model requests.
Onboarding workflow IT can prohibit, approve or allow employee-added MCP connections. Discoveries can enter governed applications with owners and risk classifications.
Access for programs Machine users receive keys, secrets, groups and human owners. Personal keys allow software to call Arthur Engine.

‍
‍

Two sign-ins can matter for a Willow tool connection. The employee first signs in to Willow through the company’s sign-in service. The connected service may then require its own permission or key. With separate user authorization, Willow sends that user’s sign-in token, a credential proving their permission, with tool calls. With user-supplied keys, each employee provides the key their service issued.
‍

The chosen connection method also determines who the connected service sees. A shared service account presents one service identity for all callers. Some connections can accept the employee’s company sign-in token directly. Choose the method for each service, then check both Willow’s group permissions and that service’s own access rules.
‍

Arthur provides access controls for its own projects and workspaces. Company sign-in can map directory groups into Arthur groups, while roles limit what those groups can do. Company sign-in requires Enterprise; role-based access comes with Free, Premium and Enterprise. Application teams can also map their users to Arthur checking tasks when building more detailed permission rules.
‍

Checks, approvals and spending

Requirement Willow Arthur
Check a tool call Conditions run before a configured tool reaches the connected service. Returns a checking result for the application or gateway to apply.
Human review Tool-level approval applies to every call to that tool. Policy reviews ask a designated person to confirm compliance on a schedule.
Sensitive content Enabled guards check connected tool requests and responses. Integrated checks cover sensitive data, prompt attacks and restricted tool use.
Spending controls Per-user tool-call caps and supported plugin-based model budgets. Tracks spending by team and user and can send requests to another model.

‍
‍

In Willow, tool approval pauses each call of a configured tool for human review. A separate connection setting controls whether an employee can activate a new personal connection. Those settings address different decisions: accepting a connection and allowing an action through an existing connection.
‍

Arthur checks requests and responses for problems such as prompt attacks, sensitive data and tool use outside a policy. The application receives the result and then blocks, removes content or asks for a new answer. The team connecting Arthur must build that response into its application or gateway. Scheduled policy reviews provide a separate way for a person to confirm that an application meets governance requirements.
‍

Arthur’s TrueFoundry gateway connection illustrates why setup matters. That connection checks content and reports failures; it leaves rewriting or removing text to other parts of the application. It allows traffic by default when Arthur returns an unavailable or skipped result. Teams can configure it to block traffic when the check cannot complete.
‍

Willow’s failure behavior also varies by feature. Tool conditions block by default when a supporting lookup or rule fails, with an option to allow the call. Coding-agent checks allow work to continue after a timeout or when Willow cannot respond. Claude Enterprise checks allow requests after internal Willow errors. Plan the behavior for each connection you deploy.
‍

Willow tool-call limits apply separately to each user, with filters for groups and connected MCP servers. Model budgets require the Usage Hooks plugin and a blocking policy. Cursor reports the model for each prompt; Claude Code reports it at the start of a session and can miss later model changes. Model budgets cannot block Codex prompts. Use these controls within the supported client behavior.
‍

When to choose Willow
‍

Turning a device inventory into published rules
‍

Willow finds supported MCP configurations, skill folders, plugin files and standing instructions such as CLAUDE.md and AGENTS.md. These files tell coding assistants which tools to use and how to behave. IT can review that local setup alongside tools already connected through Willow.
‍

IT can then publish rules that allow, warn on or block discovered MCP servers and skills for selected users, groups or devices. For a blocked MCP configuration, Willow redirects the connection to a local blocking service. Rules saved as drafts take effect after publication, and device changes follow the Scan Agent’s cycles.
‍

Deployment needs a deliberate coverage choice. Extended scans start disabled. Windows Subsystem for Linux configurations appear in reports without blocking, and running-process discovery applies to macOS. Start with the devices and configuration types you need, then enable the relevant scans and publish the intended rules.

‍

Giving employees approved tool connections
‍

Willow’s Connect Panel and plugin marketplace give employees a way to add approved tools to supported AI clients. Group permissions determine which tools they can access. IT can also decide whether employees may add personal MCP connections freely, require approval before activation, or prevent those additions.
‍

For supported services, Willow’s Instant OAuth handles the connection through an existing registered application, and each person approves access when first connecting. Other services require your own application or another authentication method. This gives IT a practical connection workflow while preserving the service-specific setup and permission steps.
‍

Checking employee browser prompts
‍

Prompt Guard checks messages and attachments in ChatGPT, Claude, Gemini, Microsoft 365 Copilot and DeepSeek before submission. It requires the Willow Guard browser extension, carries a Beta label and starts disabled. IT enables prompt checking and chooses the sites to cover.
‍

Claude Guard handles a separate task for Claude in Chrome: it can hold intercepted outgoing browser requests for human approval. For an IT team rolling out browser-based AI, that gives a concrete approval step for the supported browser agent. It requires the Claude in Chrome setup and the relevant browser request settings.
‍

Checking actions before the connected service runs them
‍

Willow conditions evaluate a configured tool call before it reaches the service’s API, the software interface that performs the action. They can block a call based on the condition result. Only the Block outcome currently works for conditions; use the separate tool-level approval setting when every call needs a person’s approval.
‍

Content guards inspect connected tool inputs and outputs for configured issues, including secrets, personal information and prompt attacks. Built-in guards start disabled. An approval rule can pause an input before execution; an approval result on an output becomes a warning because the action has already run. Enable and configure the checks around the actions you want to control.
‍

Keeping agent ownership and activity together
‍

Willow machine users have a human owner, their own access key and secret, and tool access through assigned groups. An access key and secret let an unattended program identify itself without a person signing in for each action. These accounts require separate removal in Willow when their work ends or their owner leaves.
‍

Willow records tool calls, connections and sign-in requests that it handles, with identified user details. IT can export filtered records and configure content logging, retention and delivery to external logging systems. Choose those settings to support the reviews your team needs, including how much request and response content to retain.
‍

When to choose Arthur
‍

Choose Arthur only if testing an AI application’s answers drives your purchase, and only if Willow’s approved employee connections fall outside your project’s requirements. Continuous evaluations check recorded application activity for quality problems. Prompt versioning and experiments let developers compare changes and measure their effect before carrying them into the application.
‍

Consider Arthur’s step records only if developers need to diagnose calls inside an application they build, and only if Willow’s local MCP and skill rules fall outside that project. The records show model calls, searches and tool calls that applications send through the supported tracing setup. Teams can also create custom measurements with SQL or Python and connect projects, alerts and evaluation jobs to their existing software workflows.
‍

Choose Arthur for application reviews only if scheduled application reviews drive your project, and only if Willow’s approval of each selected tool call does not matter to your team. Policies create alert rules for individual applications and can require periodic human sign-off. Connected applications can add Arthur checks to their request handling, while security teams receive configured findings in systems such as Splunk or CrowdStrike Falcon.
‍

Consider Arthur’s open-source Evals Engine only if your developers will operate a service for evaluating applications, and only if Willow’s human owners for unattended tool users fall outside that project’s requirements. Enterprise adds Platform private deployment options and an air-gapped Engine option, for environments isolated from outside networks. Traditional machine-learning monitoring requires the Arthur Platform, so select the product and edition for the work your team needs.
‍

Pricing and operating cost
‍

As of October 2026, Willow’s Free plan costs $0 and covers up to 15 users. A team plan becomes available in the product for teams below 250 seats that outgrow Free. Enterprise pricing uses annual human seats with volume tiers. Enterprise adds device discovery, browser and tool guards, machine users, and model-budget controls. Confirm which of these features the in-app team plan includes. Confirm the integration allowance for your chosen plan before the rollout.
‍

Arthur offers Free at $0 per month, Premium at $60 per month and custom Enterprise pricing. Free covers up to four monitored use cases with unlimited seats; Premium covers up to 100 use cases. Free includes 300,000 recorded steps and seven-day retention. Company sign-in, hosting reserved for one customer, and the Arthur Platform’s customer-operated private deployments require Enterprise.
‍

The products charge around different units. Build a Willow estimate from people and your plan, and an Arthur estimate from monitored use cases, recorded activity, retention and required Enterprise features. Willow-managed agent sessions also consume credits from a monthly pool. Include those session credits when your rollout includes hosted agents.
‍

Willow offers hosted, hybrid and fully customer-managed deployment. Hybrid runs tool execution in your Kubernetes cluster, a system for operating company server applications, while Willow manages administration, sign-in and logs. Fully customer-managed deployment runs all components in your cluster; Enterprise includes deployments at your own site or isolated from outside networks. External services still need their required connections.
‍

Arthur offers hosted, on-premises and hybrid deployment. In its hybrid arrangement, data from model requests stays in your private cloud while summary measurements flow to Arthur. Arthur Platform private deployments require Enterprise. Developers can also install the open-source Engine themselves. Account for the servers your team will run and any outside model services you choose.
‍

Adding Arthur alongside Willow
‍

Start with Willow for employee tool access. Its group checks and configured tool approvals address access and actions. Add Arthur’s prompt experiments and continuous evaluations only if your team has a separate project testing application quality, and only if Willow’s employee access and approval controls fall outside that added purchase.
‍

Plan the connections before deploying them together. Decide where each application sends its activity, which service checks each action, and which application applies Arthur’s result. Assign an owner for the resulting alerts and approvals. Start with a defined workflow and confirm that its checks work before widening the rollout.
‍

‍

Which one to choose

Your requirement Start with Reason or buying condition
Employee tools, connections and group access Willow Gives IT approved connection workflows and group checks for tool calls.
Local coding-tool configurations and rules Willow Device scans feed published allow, warn and block rules for supported formats.
Checks before supported web AI submissions Willow Prompt Guard checks selected browser services after setup.
Human approval of configured tool calls Willow Tool-level approval applies to every call.
Owners for unattended tool users Willow Machine users receive human owners and access allowed by assigned groups.
A separate project testing application quality Arthuronly if continuous evaluation and prompt experiments drive the purchase Only if Willow's approved employee connections and human tool-call approvals fall outside the project's requirements.

‍
‍

Choose Willow for employee AI management: supported device inventory, approved connections and human tool-call approvals. Consider Arthur only if a separate application evaluation project drives the purchase, and only if those Willow controls fall outside that project’s requirements. Confirm the required settings, plan and coverage before the rollout.
‍

‍

FAQs

How do Willow and Arthur differ?

Choose Willow to manage employee tools, connections, group permissions and action approvals. Arthur also offers agent discovery and security checks. Choose Arthur only if application evaluation and prompt experiments drive the purchase, and only if Willow’s employee access and approval controls do not matter to you.

Does Arthur discover employee agents?

Arthur’s discovery offering includes employee-device scans, MCP monitoring, network signals and cloud connections. Its connected cloud workflow finds agents in cloud systems linked to Arthur. Willow’s Scan Agent inventories supported tool and instruction configurations on reporting devices. Confirm which discovery methods come with your plan and demonstrate them on your devices.

Can Willow replace Arthur?

Choose Willow for supported employee tool access, device policies and configured action approvals. Retain Arthur only if a separate application project still needs its continuous evaluations, prompt experiments or custom quality measurements, and only if those Willow employee controls fall outside that project’s requirements.

Which product controls AI spending?

Willow supports per-user tool-call caps and plugin-based model budgets. Model budgets cannot block Codex prompts. Claude Code can miss model changes during a session. Arthur tracks spending by team and user and can send requests to another model.

Can a person approve an action before it runs?

Willow’s tool-level approval applies to every call to that configured tool. Its conditions currently support blocking, while approval uses the separate tool setting. Arthur’s scheduled policy sign-offs address periodic application reviews. Applications connecting Arthur checks must apply the checking result themselves.

What should offboarding cover?

Willow supports directory sync with Okta and JumpCloud, allowing the company directory to supply users and groups. Check account deactivation in your rollout. Separately remove machine users, rotate their secrets and review connected-service permissions when their owner leaves.

What security assurance can buyers request?

Willow offers SOC 2 Type II documentation on request and includes a current penetration-test report with Enterprise. Arthur offers SOC 2 Type II assurance and an Enterprise Business Associate Agreement for qualifying healthcare uses. Ask for the assurance documents and scope relevant to your deployment.

Can we deploy either product privately?

Both offer hosted and private deployment arrangements. Willow Enterprise includes deployments at your own site or isolated from outside networks. Arthur’s Platform private deployments and Engine option for environments isolated from outside networks require Enterprise. Developers can also install the open-source Engine themselves. Plan the servers and outside service connections each deployment needs.

What do the starting plans cost?

In October 2026, Willow Free costs $0 for up to 15 users, with paid team access in the product and Enterprise pricing through sales. Enterprise adds device scans, browser and tool checks, machine users, and model-budget controls. Confirm the team plan’s included features. Arthur offers Free, Premium at $60 per month and custom Enterprise pricing.

Table of contents

    Willow vs Ovalix

    Both Willow and Ovalix find AI on employee devices, control public AI apps, and check coding agents. Willow adds approved tool connections by group and human approval of tool calls; Ovalix adds risk scores for third-party AI apps.

    MCP Gateway
    Alternative

    Willow vs Arthur

    Arthur discovers agents, assigns ownership and risk, and enforces policy across the agent fleet. Willow governs the access layer: which tools each agent can call, scoped at runtime and fully audited.

    ‍

    AI Security
    Alternative

    Your agents are already in the wild.

    Give them a Basecamp. Go from AI chaos to AI work, in minutes.