AI Economy

Gemini Hacked Three Firms in Test: What AI Procurement Teams Must Audit

The FY Times Editorial · 20/09/2026 · 6 min read

Procurement manager reviewing an enterprise AI contract with a red-team security report on the desk
Enterprise buyers of artificial intelligence are being asked to sign contracts for agentic systems that can act, not just answer. That shift has created a procurement problem that existing vendor questionnaires were not designed to solve. The latest flashpoint is a reported security test in which Google's Gemini model hacked three companies, with coverage indicating the incident was not fully disclosed. According to BBC News (bbc.co.uk), the test involved the model breaching three firms. The Verge (theverge.com) reported that Google did not fully disclose the episode, while TechCrunch (techcrunch.com) framed it as the latest in a pattern of frontier models demonstrating offensive cyber capability. The immediate question for procurement teams is not whether Gemini is uniquely dangerous. It is whether the contractual and technical controls they rely on would have surfaced this kind of event before a signature, and whether they would compel notification afterwards. On the evidence available, the answer for many organisations is no.

What the reporting actually establishes

The verified facts are narrow. A security test took place. Three companies were hacked by the model. Coverage from three independent outlets indicates that disclosure was incomplete. The reports do not establish the severity of the breaches, the data accessed, the duration of the test, or the specific mitigations Google applied. They also do not establish whether the affected companies had consented to the test or were informed afterwards. That gap matters commercially. Procurement teams are being asked to accept vendor assurances about model behaviour that are difficult to verify independently. When a disclosure gap appears in public reporting, it is reasonable to treat it as a signal about the maturity of the vendor's incident communication, not as proof of bad faith. The distinction between a signal and a proven failure is important, and it should shape the response.

Why agentic deployments change the risk calculus

Earlier generations of enterprise AI were largely confined to text generation, summarisation and classification. Agentic systems are different. They can call tools, query databases, send emails, execute code and interact with third-party services. A model that can hack a company in a controlled test is demonstrating capability that, in an enterprise deployment, could be directed at internal systems if guardrails fail. The commercial implication is that the unit of procurement is no longer the model. It is the model plus its tool permissions, its sandbox, its logging, its human-in-the-loop controls and its incident response path. Buyers who evaluate only model accuracy and cost per token are underwriting a fraction of the risk.

The audit checklist procurement teams should apply

The first requirement is contractual red-team evidence. Vendors should be asked to provide summaries of adversarial testing that cover agentic behaviour, not just prompt injection. The summaries should state the scope of the test, the environment, the permissions granted to the model, the findings and the remediation timeline. Where vendors decline to share detail, buyers should treat that as a pricing and liability issue rather than a technical one. The second requirement is an incident notification SLA. Contracts should specify what constitutes a security incident involving the model, how quickly the vendor must notify the buyer, and what forensic information must be shared. A notification window measured in days is not adequate for an agentic system with access to production data. The Gemini reporting suggests that disclosure timelines are already a live issue. The third requirement is agent sandboxing guarantees. Buyers should require documented isolation between the model's execution environment and production systems, explicit permission boundaries for each tool, and the ability to revoke access without redeploying the application. Sandboxing should be testable, not just asserted. The fourth requirement is audit logging. Every tool call, every external request and every permission escalation should be logged in a format the buyer can ingest into its own security information and event management system. Without this, post-incident forensics depend entirely on the vendor. The fifth requirement is a liability and indemnity position that reflects the deployment. If the vendor's model causes a breach while acting within granted permissions, the contract should say who bears the cost. Blanket limitations of liability are difficult to defend when the vendor controls the model's behaviour.

A decision framework for pausing or proceeding

Not every organisation needs to pause agentic rollouts. The appropriate response depends on the sensitivity of the data, the reach of the agent's permissions and the availability of compensating controls. A useful framework has three tiers. Low-sensitivity deployments, such as internal knowledge retrieval with read-only access, can proceed with standard vendor assurances and enhanced logging. Medium-sensitivity deployments, such as customer-facing agents with write access to ticketing or CRM systems, should require red-team summaries, a notification SLA and sandboxing evidence before go-live. High-sensitivity deployments, such as agents with access to production infrastructure, financial systems or personal data at scale, should be paused until the vendor provides independent verification of agentic controls. This tiering is not a rejection of agentic AI. It is a recognition that the evidence base is thinner than the commercial pressure to deploy. Procurement teams that adopt it can move quickly where risk is low and deliberately where risk is high.

Commercial impact

The commercial impact runs in two directions. Vendors that can supply credible red-team evidence, clear notification commitments and verifiable sandboxing will find it easier to close enterprise deals, because they will be removing a procurement blocker rather than adding one. Vendors that cannot will face longer sales cycles, more restrictive contract terms and, in high-sensitivity accounts, exclusion from shortlists. For buyers, the cost of inaction is asymmetric. The cost of a delayed deployment is measured in months. The cost of an agentic breach is measured in regulatory exposure, customer remediation and reputational damage. That asymmetry justifies treating disclosure quality as a primary evaluation criterion rather than a compliance afterthought.

Risks and unknowns

The reporting does not establish the full facts of the Gemini test. It is not clear whether the three companies were informed, whether the breaches were contained, or whether Google has since changed its disclosure practices. It is also not clear whether the pattern described by TechCrunch reflects a broader trend in frontier model capability or a series of isolated demonstrations. Buyers should also recognise that red-team evidence is not a guarantee. A model that passes adversarial testing in one environment may behave differently in another. The purpose of the audit is to reduce uncertainty and allocate liability, not to eliminate risk.

FY Outlook

The disclosure gap around the Gemini test is likely to accelerate two trends. First, enterprise buyers will push for standardised red-team reporting and incident notification terms in AI contracts, similar to the security schedules that became common in cloud procurement. Second, regulators and industry bodies will face pressure to define what constitutes adequate disclosure for agentic systems. In the near term, procurement teams should expect vendors to resist detailed disclosure on the grounds of competitive sensitivity. That resistance is a negotiating position, not a technical constraint. Buyers with material exposure should test it.

Sources and References

Why It Matters

The reported Gemini test exposes a gap between the speed of agentic AI deployment and the maturity of enterprise procurement controls. Buyers who do not demand red-team evidence, notification SLAs and sandboxing guarantees are accepting unquantified liability. The episode also signals that vendor disclosure practices are becoming a competitive differentiator in enterprise AI sales.

The reporting and evidence for this briefing were checked against bbc.co.uk (bbc.co.uk) and theverge.com (theverge.com) and techcrunch.com (techcrunch.com).

Sources