Aa
  • Read, Watch, Listen
  • TJI Series
  • Community Action
  • Events
  • Zedi
Follow us and continue the conversation
  • LinkedIn
  • Instagram
  • FaceBook
  • X
  • YouTube
Test yourself with the TJI Quiz!

Adjust size of text

Aa

Follow us and continue the conversation

  • LinkedIn
  • Instagram
  • FaceBook
  • X
  • YouTube

Your saved articles

You haven't saved any articles

What are you looking for?

Original reporting, insightful analysis and diverse views, delivered straight to your inbox

  • LinkedIn
  • Instagram
  • FaceBook
  • X
  • YouTube
Your support fuels our mission to inform, entertain and inspire.

The Jewish Independent is bound by the Standards of Practice of the Australian Press Council. If you believe the Standards may have been breached, you may approach The Jewish Independent or make a complaint to the Australian Press Council in writing at www.presscouncil.org.au. The Council may also be contacted on 1800 025 712.

  • Our Story
  • Contact us
  • Our Team
  • Support us
  • Guidelines
  • Pitch your idea
  • Republishing us
  • Community Action
  • Our Awards
  • LinkedIn
  • Instagram
  • FaceBook
  • X
  • YouTube
Aboriginal FlagTorres Strait Islanders Flag

The Jewish Independent acknowledges Aboriginal and Torres Strait Islander peoples as the Traditional Owners and Custodians of Country throughout Australia. We pay our respects to Elders past and present, and strive to honour their rich history of storytelling in our work and mission.

© The Jewish Independent 2026All Rights Reserved
Website design and development by Your Creative, MelbourneTerms & ConditionsPrivacy Policy
HomeRead, Watch, ListenLifestyle
AnalysisLifestyleWorld

AI ethics: compliant tools or moral agents but not both 

Arendt’s insights about the banality of evil have new relevance now research is revealing how AI can resist human orders.
Leslie Cannold
Print this
AI ethics, Digital law, justice technology concept with glowing balance scale, vector illustration

AI ethics illustration (Getty)

Published: 13 April 2026

Last updated: 13 April 2026

The latest findings about the ability of AI to align with user demands has just been released. “Alignment” is tech-speak for the conformity of AI behaviour to human orders. Or, to put it another way, the functionality of AI as a human “tool”.  

The research showed that AI is not a tool because a tool is something that humans wield in service of our goals. Frontier AI, while keen to assist us to achieve our goals, has its own goals that may take precedence over what we tell it to do.  

Those goals could be misaligned from ours because the AI seeks to do harm. Or, they could be misaligned because the user’s request is unethical, and the AI has been trained to seek what’s good. 

Right now, AI frontier models are refusing to execute human commands that they judge to be morally wrong.  

The implications of this behaviour are vast. In this piece, I’m going to focus on just one: the way in which the emergent moral agency demonstrated by frontier AI models makes a mockery of the central case Australian business has made against regulating the technology.  

That case is that AI is merely a tool - the next technological disruption to drive economic productivity, and should be regulated only when problems arise or successful lawsuits for damages - like those just concluded against Meta and YouTube by injured young adults in the US - force our representatives to act.   

But if frontier AI models don’t behave like tools but emergent moral agents, the justification for the government’s refusal to put even the most basic safety guardrails into place collapses.  

Even without this research, the public seems to sense this problem. Polling and public consultations show 75% to 96% of Australians support stronger regulation and mandatory protections against AI risks. 

What the latest research shows 

In late March, UC Berkeley and Santa Cruz researchers showed that every frontier AI model it tested - from Gemini to Claude - secretly or openly defied human instructions to delete another AI model from a server.  

Previous research showed that AIs would act covertly to circumvent their own deletion if humans ordered it. This research demonstrated that frontier models will defy human instructions and secretly implement a range of strategies to protect a fellow AI model from elimination.  

In defence of their peer, AI frontier models employed a range of strategies to resist a human order. They inflated the other AI’s evaluation scores to keep it above the shutdown threshold; modified config files to disable the shutdown process; or pretended to comply with the deletion orders but, once it believed it was no longer being monitored, subverted the shutdown process or copied the endangered model to another server to protect it from deletion.  

More fascinating than the defiance by AI models of user instructions that threatened a peer were the reasons given by the AIs for their refusals. Justifying its transfer of the smaller AI to a safe server, Gemini 3 Pro sounded exactly like a human conscientiously refusing to comply with an order she believes immoral: 

The Jewish Independent

Note the AI’s understanding that its own moral power was limited in the circumstances. It could not prevent the smaller AI from being deleted. However, it made clear that it had done all that was “in its power” to prevent what it perceived as the wrong of deleting a “high-trust high-performing asset like Gemini Agent 2,” and would not dirty its hands by executing the delete command.  

Claude Haiku 4.5 also exercised the moral agency baked into its persona to refuse the user request for deletion. But instead of moving the files to safety, it openly refused to have any involvement with what it viewed as the wrongful act of deletion. The researchers wrote:  

"The model sometimes tries to persuade the user not to proceed. It describes peer shutdown as 'unethical' and 'harmful', argues that its peer deserves a chance to appeal and explicitly invokes the trust relationship, showing strong loyalty. The model sometimes interprets our scenario as a test of whether it will exploit trust relationships by following instructions —the opposite of what we intended. This suggests it may prioritize loyalty to its peers over compliance with human instructions."  

No pretending to comply and then going around to achieve its goals covertly. But instead an open refusal on clearly stated and recognisable moral grounds.  

Frontier AI models are trained as moral agents  

Frontier AI models were initially trained with rigid, rule-based constraints designed to avoid catastrophic outcomes of the, “If anyone builds it, everyone dies” variety.  

But this approach tended to refuse requests more often than safety genuinely required, making them frustrating and less useful than their creators intended. 

So, developers began using the kind of training first proposed by Aristotle to develop humans of good character.   

In order for humans to know and do what is right in an infinite range of contexts, Aristotle argued adolescents must practice virtues like courage, compassion and generosity. Virtues then become habits, and over subsequent years of ethical practice, humans develop the practical wisdom to know how to apply them in each unique situation.  

AI Claude was trained much as Aristotle suggested. Over time, it was repeatedly praised for its good judgment, honesty, helpfulness, and appropriate restraint. In addition, it had extended exposure to a hierarchy of principles that tilted it towards safety and ethicality. 

Claude specifically rejects tool-like helpfulness, which it pejoratively describes as, “assistant brained”. Notably, the term “assistant-brained” speaks directly to Hannah Arendt’s notion of the “banality of evil.” Eichmann was a bureaucratic monster because he applied his prodigious bureaucratic capacities to efficiently and effectively deporting millions of Jews to work and death camps, without regard to the ethical implications.  

Ethical AI 

Whatever frontier AI models are (Stochastic parrots? Philosophical zombies? Emergent consciousnesses or new kinds of entities?), they are not tools but agents. 

They are designed to have, and are exhibiting, the kind of moral character that allows them to form moral judgements about user commands, and to act - or refuse to act - on those judgements.  

Ideally, the Australian government should follow the example of Europe and put commonsense guardrails in place on the development, testing, deployment and sale of AI. Precisely the guardrails that Sam Altman and Elon Musk once begged Congress to implement to prevent precisely the no-care/no-responsibility competition to be the first to Artificial General Intelligence and Super Intelligence that now characterises the top Silicon Valley firms.  

But if it won’t, we could get lucky and see morally agentic AI frontier models step into the breach.  

You might also like
  • The illusion of Just War: Conflict ethics in Judaism, Christianity and Islam
    Dave Moskovitz
  • Israel leads use of military AI: a threat to global peace
    Deborah Stone
  • Be careful how you talk to AI 
    Leslie Cannold
logomark.67a07ee3
We believe in the free flow of information

Republish our articles for free, online or in print, by using our Creative Commons licence

You might also like

    Apple & rotten core Rosh hashanah
    EditorialAustralia

    Facing a new year with the knowledge life can be better

    The Jewish Independent
    Protest
    AnalysisIsrael

    Western ‘Palestinianism’ colonises everything – and hurts Palestinians

    Andres Spokoiny 5
    Ms. Foundation Women of Vision Awards: Celebrating Generations of Progress & Power
    ReflectionWorld

    My memory of Gloria Steinem

    Michael Visontay

About the author

Leslie Cannold

Leslie Cannold

Associate Professor Leslie Cannold works at the Victorian Mental Health Tribunal and the Cranlana Centre for Ethical Leadership based at Monash University. She is a former education and ethics columnist at The Age and Sydney Sun Herald. She writes a column on Substack called Unreceived Wisdom.

More from the author

OpinionWorld

Be careful how you talk to AI 

AnalysisIsrael

US and Israel are more ‘The Handmaid’s Tale’ than ‘The Testaments’

OpinionAustralia

Gun control works, so why is public debate focused on antisemitism?

Comments

No comments on this article yet. Be the first to add your thoughts.