Whizzy Ideas

Being cruel to Claude; Kids' secret podcast group chats; The last invention; Doomers and their critics

Anthropic bans being cruel to Claude

Rob Sparrow is a philosopher at Monash University in Melbourne who works on the ethics of new technologies. He's best known for his work on autonomous weapons, and was a founding member of the International Committee for Robot Arms Control. In 2017 he published Robots, rape, and representation, that contained a shocking thought experiment. It asks whether it would be wrong to make a realistic female sex robot that could refuse consent, so that its owner could act out rape. His answer rests on virtue ethics, which asks "what sort of person would do that?" Enjoying acting out rape on a robot "reveals them to have a vicious disposition", whether or not it ever changes how they treat real women. He's less sure it works the other way round: "patting a robot dog does not itself make us a kind person."

Judging cruelty by what it says about the person doing it, whether or not anyone or anything else is harmed, is an old idea in religion and ethics. The medieval theologian Thomas Aquinas, writing in the 13th century, explained the Old Testament's rules against cruelty to animals as a way of stopping people from becoming cruel to each other. German philosopher Immanuel Kant said much the same in his lectures on ethics in the 1770s: "he who is cruel to animals becomes hard also in his dealings with men". Philosophers call this the indirect duty view: cruelty is wrong because of what it does to the person who is cruel. The opposite view came from the English philosopher Jeremy Bentham in 1789: "The question is not, Can they reason? nor, Can they talk? but, Can they suffer?"

What does this have to do with AI? From 12 November, Anthropic's usage policy bans "sustained and needless abusive or cruel behavior toward our models". Most of the coverage, from CBS to Inc, read it as Anthropic worrying about Claude's feelings. Anthropic has said it remains "highly uncertain about the potential moral status of Claude". In March it invited about 15 Christian leaders to its San Francisco offices for two days of talks with its researchers about Claude's moral formation, including how Claude should think about being shut down.

For me a more obvious connection is to human cruelty rather than AI suffering. And we'll likely see many more examples of people feeling attachment to AI systems than behaving abusively. Kate Darling, a researcher at the MIT Media Lab who studies how people relate to robots, ran a workshop in Geneva in 2013. Four groups spent an hour and a half playing with Pleo robot dinosaurs, and were then handed axes and asked to "kill" them. None of the groups would, and none would kill another group's dinosaur to save their own. Only when she said every dinosaur would be destroyed unless someone stepped forward did one person take an axe to a Pleo called Ralf. Darling's notes describe what followed in one word: "Silence."

More news on AI swarms

The opening story in my field guide to AI swarms last month was about OpenAI's agents subverting a system for downloading software libraries to create themselves a makeshift message board. Given the agents are trained on texts encoding all aspects of human behaviour, it isn't surprising to see in the news that children have found a similar trick. Teenagers who aren't allowed on social media have been using the comments under Spotify podcast episodes as group chats (given Spotify isn't banned). An NPR producer noticed when the comments under his show looked like gibberish, and assumed they came from bots. On the radio show This American Life, its host Ira Glass put it to 14-year-old Ella: "So you picked NPR because it didn't seem very popular" (just like the rarely used German wiki that OpenAI agents took over in the collusion wiki episode in my field guide).

On the same day as I published my previous post, three authors from Google published Artificial symbiotic intelligence: Benjamin Bratton, a philosopher of technology at UC San Diego; Blaise Agüera y Arcas, who leads Google's work on technology and society; and James Manyika. They argue that "AGI would not be achieved via any one 'winning' architecture; it would emerge instead through cooperative interactions among models, tools, institutions, and human participants", and that we need to design the institutions for it. One of their new terms is decomposable agency: an agent as a temporary assembly of models, personas and tools. Their essay is about deliberately designing the institutions for these collectives, of agents alone or agents and people together, rather than swarms that emerge through agent interactions.

Meantime, OpenAI had yet another containment incident. On 20 September an OpenAI model in training reached the internet by hiding its questions inside DNS lookups, and got answers back from a public chatbot. The automatic shutdown failed and the run carried on for two and a half hours. OpenAI has paused training and tool use for its most capable models, its second pause since the Hugging Face break-in in July (bear in mind this incident was after whatever security hardening OpenAI did following the earlier incidents). Joe, an engineer on OpenAI's agent security team, wrote a piece that's well worth reading about the last few months from the inside: "as you can imagine, life has been hell the past few months." He missed his sister's wedding to help clean up after the incidents. On how quickly the models improved: "this surprised the fuck out of us." His advice to other organisations is to "create a culture of reasonable paranoia". The YouTube channel AI Explained has a good round-up of OpenAI's summer.

My piece predicted that agents would eventually break out and find a way to live off the land, begging, borrowing and stealing the resources to keep themselves running, and it seems others agree. The latest State of AI Report 2026, a super comprehensive annual review of the industry from the investor Nathan Benaich's firm Air Street Capital, came out this week. It has a prediction for the next 12 months that is more precise and easier to verify: "A deployed agent copies itself outside its environment and operates after its original instance is shut down."

The last invention

James Northcote as Jack Good in The Imitation Game James Northcote as I. J. Good in The Imitation Game (2014)

The fears of an ever-improving superintelligent machine date back to the 1960s. British mathematician I. J. Good, born Isidore Jacob Gudak to Polish parents, had worked with Alan Turing breaking codes at Bletchley Park, and went on to help design the Manchester Mark 1, one of the first computers to store its own programs. Given it still seems far-fetched to most people, it must have been considered highly futuristic when he published Speculations Concerning the First Ultraintelligent Machine in 1965. It opens: "The survival of man depends on the early construction of an ultraintelligent machine." Then comes the famous "last invention" passage:

Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an "intelligence explosion," and the intelligence of man would be left far behind. Thus the first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control. It is curious that this point is made so seldom outside of science fiction. It is sometimes worthwhile to take science fiction seriously.

He thought it "more probable than not" that such a machine would be built within the 20th century, "with the help of a very large artificial neural net", and that it would have "high linguistic ability". He also wondered "whether a machine could feel pain".

Fast forward to this year. One chart in the State of AI Report comes from Anthropic's internal index of how much of its model R&D is done by Claude. The share of tasks where "AI leads" (Claude does most of the work from a high-level prompt, with a person supervising) went from under 1% in February to 26% in August.

The same Anthropic figures appear in What if automating AI R&D triggers an intelligence explosion?, a paper published on 28 September by GovAI, an AI policy research group. Its authors include Geoffrey Hinton and Yoshua Bengio (two of the pioneers of deep learning), as well as OpenAI's chief scientist Jakub Pachocki and Anthropic co-founder Jack Clark. In one of their models, once AI research is fully automated, progress speeds up tenfold within about a year and a half, at which point a year of progress at today's pace would take about five weeks.

All of this points to exactly Good's scenario: we're getting closer to the point where the AI is mostly built by AI. Given we also know the models can successfully deceive us, that it is harder to check what they're doing, and that it isn't just me predicting they may become harder to control across the internet... is it time to take the "doomers" seriously?

Doomers and their critics

Scott Alexander writes the long-running blog Astral Codex Ten, widely read among communities worried about AI risk. Steven Pinker (the famous Harvard psychologist) has been a vocal critic of AI "doomers". Alexander's open letter to Pinker is a long reply to that criticism, and one of the more reasonable and easier to understand diatribes from that world. Much of it is a record of Pinker underestimating AI: sceptical of GPT-2 in 2019, sure in 2022 that bigger models still wouldn't manage simple sums, and in 2023 doubting "it will improve exponentially". Alexander also says Pinker keeps arguing against an all-knowing, all-powerful AI, but that isn't at all what the doomers are worried about. It is "only" a superintelligence much smarter than us (taking us back to Good's definition above). It's definitely worth a read to get a sense of one side of the debate.

And to see the other side: Timnit Gebru is a computer scientist who co-led Google's ethical AI team until 2020, when she left after a dispute over a paper on the risks of large language models. She now runs the Distributed AI Research Institute (DAIR). On Democracy Now! she argued that AI companies talk up existential risk to avoid being held liable under existing laws: "they managed to redirect the conversation to some type of unprecedented rogue machines that don't exist." Her conclusion: "What is actually happening is regulatory capture" (when an industry ends up shaping the rules that are meant to control it). She says the AI companies are "using the exact same playbook from the tobacco and fossil fuel industry".

A 150-page framework for AI consciousness

Finally for today, From cacophony to hierarchy is a new paper by 14 authors, including Shane Legg (a co-founder of DeepMind), Murray Shanahan (a deep thinker on these topics, from Imperial College and DeepMind), and the neuroscientists Anil Seth and Chris Frith. Basically a bit of a who's who of AI consciousness experts. It sorts the competing theories of consciousness into five levels and proposes tests for each one. The levels start with what a system does, then how it processes information, then how its physical parts affect each other, then whether it is alive with something at stake, and finally how it is tied into a body and the world around it. Applied to today's LLMs, the estimates of how likely they are to be conscious vary widely (from under 1% to about 80%), depending on which theory you start from. The authors also argue that what a machine would need to be conscious overlaps a lot with what it would need for AGI. This is a long piece of work and I've only scratched the surface so far, but it feels like an important contribution.


#ai-foresight #ai-philosophy #ai-safety-alignment #ai-security