Static Keys, Non-Deterministic Agents: Willow's CEO on Why AI Broke Access Control

Static Keys, Non-Deterministic Agents: Willow's CEO on Why AI Broke Access Control
Willow CEO and co-founder Eyal Ben Ezra sat down with host McKenzie on The Secure Disclosure, Aikido Security's podcast, for a sharp conversation about the problem every security team is now living with: we are running non-deterministic AI agents on infrastructure built for deterministic software, and the access model has not caught up.
His framing of the core issue was blunt:
"We are handing non-deterministic AI agents static, deterministic API keys, and then acting surprised when production databases are suddenly deleted."
This is a recap of the argument, the parts worth quoting, and what it means for anyone deploying agents at work.
We gave agents the same keys we gave microservices
The setup is familiar to any engineer. We built microservices, handed them API keys, and let them talk to each other. That worked because the system was deterministic. A microservice does not wake up one day and decide to do something strange with the access it holds.
Agents are different. As McKenzie put it in the conversation, we now give an agent the exact same access token as the deterministic service, except this agent can decide on its own that the best way to clean up a database is to delete all of it. And it only has to make that call once for the result to be catastrophic.
Eyal's example was not hypothetical. "Just the other day I got a call from a developer who accidentally deleted our entire production database." The access was valid. The judgment was not.
Scopes are dead. Start thinking in capabilities.
Eyal's central argument is that the vocabulary of access control needs to change. The SaaS era was built on scopes: read customers, write customers, verbs that made sense for deterministic systems through 2022 and 2023. That abstraction does not hold for an actor that reasons and acts on its own behalf.
"We need to start thinking about capabilities, not scopes of permissions."
In practice that means a new layer between the agent and everything it can touch: a risk score for each tool, and a deliberate decision about which agent, under which identity, is allowed to do what. You set it up once, so a non-deterministic actor cannot turn a valid credential into an incident.
The biggest risk is not malice. It is accidents.
When asked where the real danger sits, Eyal was clear that the largest slice of the risk pie today is safety, not attackers.
"Most risk is coming from safety issues, less from malicious intent right now."
Non-technical employees, and even engineers, do things they never intended to do. McKenzie offered his own example: a coding agent that went looking for credentials on his machine, found an SSH key, and started buying domain names on his AWS account, entirely on its own initiative. The agent was not compromised. It was just capable, and under-governed.
For the malicious side, Eyal's instinct is to trace the root cause rather than chase the symptom. The question is not only how to stop a bad prompt, it is why agents are holding reusable API keys in their environments at all. His answer: give them keys that can be used once, do the key exchange through a control layer, and keep the real credentials off the endpoint. "Keys shouldn't be used by LLMs."
Can we solve prompt injection?
McKenzie pushed on the attack vector everyone is arguing about, and made a point worth repeating: the name "prompt injection" is misleading, because it sounds like SQL injection or command injection, which are solved with parameterization. You cannot parameterize a prompt.
Eyal did not hedge.
"True answer: it's unsolvable. Given the current technology, not solvable. But you can decrease the risk."
His approach has two parts. First, detection that respects latency, cost, and user experience: start with a small language model for a fast signal, then escalate to a large language model, which he considers state of the art for catching known prompt-injection attempts, only when the smaller model flags something. Sending every request to a frontier model doubles cost and degrades the experience, so it is not the path.
Second, and more important, stop defending only at the prompt and close the valve where the value actually is.
"The risk sits in your resources. It's in your SaaS, your AWS, your Salesforce, your NetSuite, your Gmail. Close the valve there."
That means routing tool calls through an MCP gateway, getting access right, and keeping API tokens off the endpoint. You cannot make prompt injection deterministic and solvable, so you protect the treasure chest instead of hoping to catch every poisoned instruction.
What good agent hygiene looks like
Pulling the practical advice together, Eyal's playbook for deploying agents safely is less about a single product and more about posture:
- Do not hand API keys directly to agents. Abstract credentials behind a system that does the key exchange, so the real keys never sit with the agent.
- Scope by capability and identity. Decide which agent, under which identity, can do what, and attach a risk score to each tool.
- Set smart defaults per agent. Willow ships defaults for Claude, Codex, and Cursor, with lists of folders and bash commands agents cannot touch, so an agent never reaches the directory holding your credentials in the first place.
- Keep guardrails as the backstop. When something slips past the command-layer controls, runtime guardrails catch it, including prompt-injection detection and PII redaction.
- Keep a human in the loop for the actions that warrant it, then let the recurring, lower-risk work run on a schedule.
The optimist's case
What stood out across the interview was that Eyal argued all of this from optimism, not fear. He is a three-time founder building what he calls the control plane for enterprise AI agents, and his thesis is business-first: the point of governance is not to take risk to zero, it is to let the business say yes.
"You're in business to take risks. If you're not taking risks, you're not in business."
Most companies, he noted, are using a tiny fraction of what the technology can already do for them. The job now is to lay the foundation that lets them adopt it fast without getting run down. That is the same argument Willow is built on: give every agent an identity, scope what it can actually do, keep the credentials off the endpoint, and audit everything, so security becomes the team that enables adoption instead of the team that blocks it.
Watch the full conversation on The Secure Disclosure for the parts we did not cover, including the Shai-Hulud supply-chain attack that used a malicious prompt to turn agents into secret-stealers, the "context bombing" technique, and a genuinely good debate about whether outsourcing our thinking to AI is worth the trade.
Background Agents in the Enterprise
Most teams can spin up an agent. Few can deploy one their security team signs off on. Here's the framework that does both.
FAQS
Everything you need to get your Basecamp running.
Your agents are already in the wild.
Give them a Basecamp. Go from AI chaos to AI work, in minutes.