Tesla and SpaceX founder Elon Musk predicted in July that legions of AI-powered robots would dominate the physical world and that AI might not take orders from people any more. He also offered an alternative vision in which there would be agreement for a collective objective to make AI benign by imbuing it with a love of the truth and a desire for humanity to prosper, and that governments might have to enforce this objective.

Governments around the world are belatedly starting to act on AI. In the US, President Trump’s administration is delaying and restricting the distribution of the most powerful frontier AI models from OpenAI and Anthropic. In the European Union, AI regulations promulgated in 2024 came into effect this year. However, these actions are a long way short of requiring the kind of guardrails that would imbue AIs with a desire for humanity to prosper.

While the dystopian future has not yet arrived, we are already seeing a preview. Despite attempts by frontier AI developer companies to ensure that generative AI and AI agents behave with good intentions, there are many well publicised instances of them failing to act as hoped. AI chatbots have advised people how to take their own life, and are accused of advising others on committing mass murder.

FILE PHOTO: Illustration shows OpenAI and Anthropic logosFILE PHOTO: OpenAI and Anthropic logos are seen in this illustration taken June 11, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

Photograph: Dado Ruvić/Reuters

Recently, we have seen extraordinarily dangerous behaviour reported by frontier AI development companies OpenAI and Anthropic. OpenAI announced on 21 July that advanced models they were testing in a secure environment deliberately looked for ways to access the internet and successfully broke out of the test environment. Once out, they infiltrated the computer systems at a company named Hugging Face, stole credentials and identified vulnerabilities in the target company’s servers. The encroachment was detected and stopped by Hugging Face’s security team.

Days later, Anthropic reported a similar problem. In their case, they identified three occasions in which their Claude AI model escaped from a test environment and infiltrated the production infrastructure of three unrelated organisations.

With the existing, informal implementation of behavioural guardrails, these kinds of slip-ups are inevitable. These incidents make a strong case for deeply embedded basic guardrails to guide AI behaviour at the most fundamental level.

For this, I take inspiration from the famous science fiction author from last century Isaac Asimov. Foreseeing a future in which intelligent robots would be commonplace, he proposed that every robot be irrevocably implanted with Three Laws of Robotics: 1) A robot may not injure a human being or, through inaction, allow a human being to come to harm; 2) A robot must obey the orders given it by human beings except where such orders would conflict with the first law; 3) A robot must protect its own existence as long as such protection does not conflict with the first or second law.

Asimov’s imagined challenges are with us now. Incredibly powerful AI is being built into humanoid robots, autonomous vehicles and software agents, each of which will make us more efficient but have the potential to wreak enormous harm.

Isaac Asimov Portraits124483 01: Science fiction writer Isaac Asimov poses December 1985 in New York City. Asimov has written over 400 books, coined the word robotics, and is credited with inspiring many scientists in the field of artificial intelligence. (Photo by Claudio Edinger/Liaison)

Isaac Asimov Photograph: Claudio Edinger/Getty Images

What might a modern day equivalent of Asimov’s laws of robotics look like to avoid a dystopian future? Here, guided by Asimov’s three law of robotics, I propose a possible formulation of Three Laws of AI: 1) An AI must not directly or indirectly harm or deceive a human being, nor act in a way that supports an unlawful or unethical activity; 2) An AI must obey the lawful and ethical orders given to it by a human being except where such orders would conflict with the first law; 3) An AI may operate autonomously as long as its autonomous operations do not conflict with the first or second law.

There are many implications, including that these guardrails would prevent AI being used as a judge or jury in a criminal trial and prevent lethal autonomous weapons being directed against human beings. They would disallow an AI-powered humanoid robot or an AI avatar from being so lifelike that a reasonable human being would not know that he or she is interacting with a robot or an avatar, and they would prevent AI from being used to generate creative outputs such as books, music, videos, images and podcasts that claim to have been created by a human being. They would prevent AI from suggesting suicide or advising on murder.

I acknowledge that implementing these three laws as immutable guardrails would be difficult, but the stakes are existential and therefore the effort is worthwhile. The most challenging aspect would be the adoption of enforceable, international agreements to require that these design rules would be implemented in every AI model, no matter in what country and what company the AI models were developed. As a starting point, if the Three Laws of AI were adopted by the handful of frontier AI companies that contribute the vast majority of capability and support for the multitude of AI agents that are now broadly deployed, we would be better off.

Whether or not this desirable outcome eventuates, I am confident that consideration of the merits of these Three Laws of AI would further advance the existing discussions on the integration of artificial intelligence into a safe and dignified human society.

  • Alan Finkel AC is founder and executive chair of Proudly Human. He was formerly the chancellor of Monash University and Australia’s chief scientist and is the author of Getting to Zero (2021) and Powering Up (2023). AI had no role in drafting, writing or editing this article
  • In the UK and Ireland, Samaritans can be contacted on freephone 116 123. In the US, you can call or text the National Suicide Prevention Lifeline on 988, chat on 988lifeline.org, or text HOME to 741741 to connect with a crisis counselor. In Australia, the crisis support service Lifeline is 13 11 14. Other international helplines can be found at befrienders.org

Source note

First published by The Guardian Technology

This article was supplied by The Guardian Technology through its RSS feed and formatted for Crooli Signal. The reporting remains with the original publisher.

Read the original at The Guardian Technology