Agentic Coding Agents Have a Permissions Problem Nobody Is Selling Around

Factory just raised $200M at a $5B valuation and OpenClaw just patched 23 confirmed vulnerabilities. Both stories point to the same gap: agentic coding agents can't reliably carry permission limits across multi-step work.

Cover art for Agentic Coding Agents Have a Permissions Problem Nobody Is Selling Around

Factory co-founders announced on September 15 that the company raised $200 million at a $5 billion valuation

more than three times what investors assigned it just five months ago . Six days later, OpenClaw published the results of a broad security audit it ran with Trail of Bits through OpenAI's Patch the Planet initiative . Both stories got covered as funding news and security news. They are actually the same story dressed differently: agentic coding agents are moving into real engineering work faster than the permission model underneath them can keep up.

What the OpenClaw audit actually found

The most important output from the Trail of Bits engagement is not the vulnerability count. Trail of Bits submitted 27 private repository advisories and three standalone hardening pull requests; of those advisories, 24 described severity-rated vulnerabilities.

The 24 broke down to zero critical, two high, 16 medium, and six low. All patched, shipped in stable releases. Fine.

What matters is the category of problem that kept appearing. The review exposed a permission risk broader than a missing approval prompt: permissions could disappear, attach to the wrong identity or target, or survive revocation during an active run.

Permissions could be lost when a request handed work to another component - the project's example is simple: a filename generator has no reason to receive tools merely because the request that started it had them. Follow-on work must carry the original limits forward, or avoid receiving access it does not need.

That failure pattern is not a bug OpenClaw wrote. It is a structural property of agentic work. When an agent runs a multi-step task - read the ticket, clone the repo, edit files, run tests, open a PR - each step is a handoff. If the permission envelope does not survive those handoffs cleanly, you either end up with a tool getting access it should not have, or you end up with a stalled task because a check somewhere dropped what it needed.

The advisories were private, and the recap does not provide CVE identifiers or public exploit details. The repair status is OpenClaw's own account of the engagement, not an independent retest published by Trail of Bits. Read it as directional, not final. But the direction is instructive: the first serious audit of a widely deployed open-source agent runtime found that the permission model is the soft tissue, not the model itself.

What Factory's valuation says about where the enterprise bet is going

Factory builds an enterprise platform where autonomous AI agents called Droids handle the full software development lifecycle, not locked to any single language model or deployment environment - customers can run it in the cloud, on-premises, or in fully air-gapped networks.

Named enterprise customers include Nvidia, Blackstone, Royal Bank of Canada, Palo Alto Networks, Adobe, and T-Mobile.

The investment thesis depends on governed, measurable software delivery rather than raw code-generation volume. That framing is deliberate. A coding agent that writes correct code but cannot tell you what it touched, in which context, and under whose authority, is a hard sell to a security team at a bank. Factory's pitch - and the reason Blackstone, Khosla, and Sequoia all re-upped - is that the governance wrapper is worth more than the model.

The company is also building a measurement layer around agent activity. Factory's Agent Effectiveness documentation describes a private-preview capability that links sessions and spending with projects, issues, pull requests and artifacts, while warning that attribution and output views cover only connected systems.

That limitation matters. The product can provide a framework for measuring value, but its documentation is not evidence that customers have already achieved organization-wide productivity gains.

That caveat is worth sitting with. A $5B valuation priced on future governance tooling is a bet that the market is real, not that it has already paid off. Teams evaluating Factory (or Cognition, or any Droids-style platform) should press on that attribution layer hard in a pilot, not after signing.

$5BFactory's September valuationup from $1.5B in April - a 3.3× jump in five months
23confirmed vulnerabilities in OpenClawfound and patched by Trail of Bits, September 2026
2High-severity findings in the auditzero Critical - but the failure class, not the count, is the story
$48BCognition's valuation, September 9the week before Factory closed its round

The multi-model stack that outperforms single-vendor setups

Both stories converge on a practical question for engineering teams: how do you run agentic coding at meaningful scale without either locking into one vendor or losing visibility into what the agents did?

The finding that matters most from teams actually using these tools in 2026: a mixed stack beats any single vendor's stack. A strong orchestrating model combined with specialist models splitting implementation and review across provider lines, with cheaper models absorbing the volume, produces stronger engineering output than the best all-Anthropic or all-OpenAI configuration you can build.

The permission problem compounds in a multi-model setup. If your orchestrator delegates a subtask to a worker model with a narrower tool allowlist, and that worker calls another component, the chain of custody for permissions needs to be explicit at every link. OpenClaw's multi-agent architecture allows you to spin up specialized agents - one for research, another for customer support, a third for data entry - each with isolated memory and distinct tool access. The isolation is the point. The question is whether it holds across the full task graph.

In a LangChain pilot of 20+ debugging workflows, coordinated agent execution produced a 93% reduction in time-to-root-cause compared to historical baselines, with over 200 engineering hours saved across 512 sessions in a single month. Development workflows showed a 65% reduction in execution time, with the biggest gains coming from compressing downstream testing - not code generation.

The testing compression is the non-obvious part. The productivity story teams keep telling is "the agent writes code faster." The data says the time savings are downstream, in test cycles, not in the writing. That changes which agents you prioritize and which integrations matter.

Delegating a bug-fix ticket to a coding agent
Without Beagle
engineer reads the ticket, opens the repo, writes a fix, manually runs tests, opens a PR, tags a reviewer - end to end, most of that is context-switching tax
With Beagle
Beagle surfaces the ticket in Slack, the agent clones the relevant branch in an isolated worktree, runs tests, and posts the PR link for a human to approve before anything merges

Where teams get stuck in practice

Agentic coding is a more autonomous, goal-oriented process, analogous to delegating a complete task to a capable junior engineer who works independently but under supervision. That analogy is useful precisely because it names the supervision requirement. You do not hand a junior engineer unconstrained write access to production config and tell them to figure it out. The same constraint applies here.

Three specific places where teams are hitting walls:

  • Scope creep during autonomous runs. The agent interprets "fix the bug" broadly and touches files it was not supposed to. Without explicit tool-allowlisting at the task level, scope is a prompt boundary - and prompts are soft.
  • Attribution after the fact. When something breaks after an agentic session, the question is not just what changed but why the agent made that specific decision. Most harnesses surface the diff; fewer surface the reasoning chain in a queryable form.
  • Revocation lag. Failures in the OpenClaw audit included permission loss across follow-on tasks, alias mismatches, changing targets, and access persisting after settings were disabled. That last one - access that survives a revocation - is the one that should concern a security lead.

A teammate like Beagle handles the human approval side of this: the agent drafts, a person approves before anything posts or merges, and the reason gets logged. That does not solve harness-level permission bugs, but it keeps a human in the loop at the consequential moments.

Beagle in action#engineering, 10:42am
The ask
'Droids opened a PR on the auth service - can someone check what it actually touched before we merge?'
Beagle drafts
pulls the PR diff, traces the files changed against the service boundary documented in the team's runbook, drafts a summary with a flag on one unexpected config change
You approve
engineer reviews the flag, asks the agent for the reasoning trace, approves the rest - merge happens with the anomaly noted
Do this in your workspace →

Agentic coding agent permissions: common questions

What is the biggest security risk when running a coding agent on real code?

Permission propagation across multi-step tasks is the documented risk that matters most. An agent that starts a task with scoped access can pass unscoped access to a sub-component during handoff. The Trail of Bits audit of OpenClaw in September 2026 found this pattern repeatedly - permissions lost, misrouted, or persisting after revocation. The fix is explicit allowlisting at each component boundary, not just at the task entry point.

How do enterprise coding agent platforms like Factory differ from open-source harnesses?

Factory's platform handles the full software development lifecycle, deployable in cloud, on-premises, or air-gapped networks, without being locked to any single language model. The enterprise proposition is governance and attribution on top of the model - connecting agent sessions to issues, PRs, and spend - rather than the model quality itself, which any harness can access.

Does running a mixed-model coding agent stack actually improve output?

Yes, based on current practitioner data. Teams getting the most out of these tools put their strongest reasoning model in the orchestrator seat and hand the actual work to whichever model is best at that specific job, regardless of who trained it. The tradeoff is complexity: each model boundary is also a permission boundary that needs explicit handling.

Should teams wait for agentic coding tools to mature before adopting them?

No, but scope tightly. Start with tasks that have a clear review gate - test writing, PR descriptions, bug fixes in isolated modules - before delegating anything that touches shared config or auth. The productivity gains in testing cycles are real and low-risk to capture now. Broader SDLC delegation is where governance tooling needs to catch up first.

What should you look for in a pilot before committing to an agentic coding platform?

Press on three things: how the platform scopes tool access per agent and per task, what the attribution layer actually covers (and what it warns it cannot cover), and whether permissions survive a mid-task revocation. Ask to see the audit log from a real session, not a demo. Public information about even well-funded platforms often includes no annual recurring revenue, retention, contract-value, gross-margin or cash-burn figures

  • treat valuation as hype signal, not quality signal.
Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle