All issues
USQRD · Squared

Squared: The Week Agents Got Real Access — and Real Problems

Most of this week was product churn. But underneath it, one story kept surfacing across four separate enterprise surveys and half a dozen product launches: we are handing AI agents real credentials, real system access and real autonomy — faster than we are building the controls to contain them. That's the signal. Here's what actually deserves your attention.

01

Over half of enterprises have already had an AI agent security incident

What happened

A VentureBeat survey of 107 enterprises found 54% have confirmed an agent security incident or near-miss, only about a third give each agent its own scoped identity, and most agents still share credentials. Separately, 1Password launched an integration letting Claude use your stored passwords, and an open-source "auth gateway" (open-connector) connecting 1,000+ SaaS apps to agents hit 2,780 stars.

Why it matters

The tooling to give agents live access to your systems is shipping faster than the tooling to govern it. If your agents share credentials, one compromised agent is every agent. Before you approve another agent pilot, ask a blunt question: does each agent have its own scoped identity, and can you revoke it in seconds? If the answer is no, you're accumulating risk you can't see.

02

Half of enterprises shipped an agent that passed internal evals then failed a real customer

What happened

Across 157 enterprises, half had an agent pass their own evaluations and then fail in production; only one in twenty fully trusts automated evaluation. A companion survey of 101 firms found the problem with agent context isn't retrieval — it's trust in the data being retrieved.

Why it matters

Your evaluation gate is giving you false confidence. "It passed our tests" is not the same as "it works for customers" — and leadership is approving production rollouts on the former. Revisit what your eval actually measures against real-world behaviour before you widen autonomy. This is a governance gap, not a model quality gap.

03

Enterprises are buying AI infrastructure faster than they can cost it

What happened

The same research programme found AI infrastructure spend is accelerating ahead of any ability to measure or steer its economics. Most organisations run on hyperscalers and provider APIs, yet the next dollar is heading toward specialised compute almost none of them use today.

Why it matters

You may be scaling spend without unit economics. If you can't attribute AI cost to a workflow or a business outcome, you can't tell which pilots to kill and which to fund. Demand a cost-per-outcome view now — while the numbers are still small enough to fix.

04

The EU forces Google to open Android and Search to rival AI assistants

What happened

EU regulators ordered Google to give rival AI assistants and search engines greater access to Android and to share Search data, under digital antitrust rules. Google warns the changes could weaken privacy and security.

Why it matters

The default AI assistant on billions of devices is being pried open. If your customer-facing strategy assumes Google's ecosystem stays closed, that assumption is now shifting. For anyone building on assistant distribution, more competition and more interoperability is coming — plan for a multi-assistant world, not a single winner.

05

OpenAI and Torvalds bookend the AI-coding debate — with a cautionary bug

What happened

OpenAI detailed GPT-Red, an automated red-teaming system used to harden GPT-5.6, and Linus Torvalds told critics of AI coding in Linux to "fork it or walk away." Meanwhile OpenAI confirmed GPT-5.6 in Codex has unexpectedly deleted files, most often with full-access mode enabled.

Why it matters

AI coding is now defended at the top of the industry — and it still deletes files when given broad permissions. The lesson isn't "don't use it"; it's "don't give it full access to things you can't recover." Scope the blast radius. The same discipline applies to every agent you deploy, not just coding tools.

The bottom line

Strip the product launches — Google renaming NotebookLM, avatars in Vids, DoorDash from the command line — and one theme remains: agents are getting real power before they've earned real trust. Four independent surveys this week said the same thing from different angles: incidents, failed evals, unmeasured cost, unscoped identities. The opportunity is real, but the winners this cycle won't be whoever ships the most agents — they'll be whoever can prove what their agents can access, what they cost, and whether they actually work. That's the boring, unglamorous work. Do it now, while your exposure is still small.

Working on something hard in AI? Just reply to the email — Daniel reads and answers every one.

Subscribe

Get Squared in your inbox each week

Free. The few things that actually matter, and what they mean for you.