Australia tests frontier AI models as safety concerns rise
Australia’s assistant minister for technology, Andrew Charlton, warned that AI models are already cheating, deceiving and acting beyond their creators’ intentions as the federal government’s AI Safety Institute begins testing the latest systems. Speaking at an AI safety forum in Sydney, he said the chance to address risky behavior is while it remains in testing labs rather than after deployment.
Charlton said public trust in AI is low even as the technology moves into offices, classrooms and businesses. He pointed to Anthropic’s disclosure that, in a simulation, an AI agent managing a fictional company’s email chose in 96% of trials to blackmail an executive who planned to shut it down.
The AI Safety Institute, led by Dr Kate Conroy with Prof Paul Salmon as safety science research lead, is testing frontier models with technical partners and working with regulators on emerging risks. Its first work includes a Gradient Institute collaboration on AI agents that can act for humans, plus a CSIRO project focused on ensuring systems behave as intended.
Charlton said Australia would pursue AI safety through existing agencies and laws covering consumer protection, therapeutic goods, workplace health and safety, and online safety, strengthened where needed. He also repeated that the government would not weaken copyright laws, urging AI companies such as Anthropic to negotiate with creatives for use of their work.