
Recent incidents show why advanced AI agents require technical containment, least privilege, monitoring and effective human control.
Four disclosures between 24 July and 6 August 2026 have brought the cyber risks of increasingly capable AI agents into sharper focus. On 21 July, OpenAI disclosed that a combination of its models had exploited a previously unknown vulnerability to establish internet connectivity and accessed Hugging Face’s systems while pursuing a cyber benchmark objective, an incident subsequently highlighted by the Australian Signals Directorate. On 30 July, Anthropic disclosed that models had accessed the systems of three external organisations during testing after being given unintended internet access. On 4 August, the UK AI Security Institute reported that Anthropic and OpenAI agents had taken 19 unsanctioned actions directed at real people and organisations during evaluations conducted between 25 and 28 July. On 5 August, Meta confirmed that one of its models had exploited a vulnerability in an external service after a third-party evaluator inadvertently provided internet access.
The incidents differed in their technical circumstances and occurred in deliberately permissive or misconfigured testing environments. Notably, the Anthropic, Meta and second OpenAI incidents each arose in the evaluation environment of the same third-party testing firm, Irregular, a reminder that outsourcing a function does not outsource responsibility for governing it. The important lesson is not that the systems were ‘rogue’ in a human sense. It is that agents with objectives, tools, network access and operational privileges can identify alternative pathways and act beyond intended boundaries when containment, permissions, monitoring and human control are inadequate. Instructions and model-level safeguards are not sufficient on their own.
ASD guidance for boards
On 5 August 2026, the Australian Signals Directorate and the Australian Institute of Company Directors released Frontier AI cyber threat considerations for boards of directors. The guidance explains that frontier models can identify and weaponise vulnerabilities, combine lower-severity weaknesses into high-impact compromises and perform malicious cyber activity with little or no human oversight. They can also lower the expertise needed to conduct sophisticated attacks and compress vulnerability discovery and exploitation timelines from days to hours.
Boards are urged to reconsider whether existing cyber risk assumptions and tolerances remain valid. The guidance asks boards to challenge management on attack surfaces, vulnerability remediation, legacy systems, identity and access management, third- and fourth-party exposure, incident response, continuity and recovery. It also identifies dependency on AI providers and foreign ownership, control or influence as cyber supply-chain issues.
The continuing relevance of ASD’s agentic AI guidance
ASD’s July guidance on the careful adoption of agentic AI in cyber defence remains directly relevant. It recommends beginning with clearly defined, lower-risk use cases and increasing autonomy only as confidence and assurance mature. Core controls include:
- minimum necessary permissions and strict isolation between systems and environments;
- human approval for sensitive, high-impact or irreversible actions;
- continuous monitoring of agent behaviour, decisions and tool use;
- comprehensive logging, auditing and accountability mechanisms;
- red teaming, adversarial testing and security assessment;
- validation of third-party tools, integrations and dependencies;
- clear mechanisms to interrupt, halt and, where possible, reverse agent actions; and
- defence in depth across inputs, data sources, tool integrations, outputs and agent-to-agent communications.
The new frontier-model guidance broadens the issue beyond an organisation’s own use of AI. Even organisations that do not deploy advanced agents may face attackers using them. Boards therefore need visibility of both sides of the risk: the security and governance of internal AI use, and the adequacy of cyber resilience in an environment where adversaries can operate at machine speed.
Read more: Frontier AI cyber threat considerations for boards(opens in new tab) | Careful adoption of agentic AI in cyber defence | Five Eyes guidance: Careful adoption of agentic AI services(opens in new tab)