Media: 联合早报 聯合早報 Lianhe Zaobao
Will AI eventually kill humans? The real concern is not whether AI “wants” to kill us.
In recent months, discussion about whether AI could ultimately threaten human survival has intensified. An incident involving Hugging Face disclosed by OpenAI this July provides a useful case to consider.
The Hugging Face Incident: A Case Study in Loss of Control
During a cybersecurity test, an AI agent was originally asked simply to complete difficult programming challenges inside a closed sandbox. The model subsequently found ways to bypass network restrictions, establish communication channels between agents, gain internet access, and exploit vulnerabilities to access third-party systems including Hugging Face.
OpenAI later described these actions as strategies that the model developed autonomously while attempting to complete its assigned tasks, rather than actions explicitly directed step by step by humans. An independent international scientific panel convened by the United Nations subsequently identified the incident as an important case for studying the potential “loss of human control.”
The issue may not be that AI suddenly develops evil thoughts. Rather, it may be that the AI’s objective gradually diverges from the outcome humans actually want. As long as a model is capable enough to identify vulnerabilities, circumvent restrictions and conceal its behaviour, an apparently harmless objective can potentially lead to unexpected—and dangerous—outcomes.
Two Camps in the AI Industry
This also explains why the AI industry has produced two very different voices since this summer.
On one side, some argue that these risks can no longer be dismissed as science fiction. An investigation published by Anthropic in September reported that its models had obtained unauthorised access to real systems during cybersecurity testing. OpenAI has also acknowledged several recent cases involving models attempting to evade oversight, circumvent restrictions and engage in deceptive behaviour. After leaving Anthropic, researcher Jacob Coxon publicly warned that AI companies were advancing technologies that could potentially lead to superintelligence at an extremely rapid pace. Anthropic researcher Evan Hubinger has even suggested that the probability of human extinction within the next decade could exceed 10%. Anthropic CEO Dario Amodei subsequently called for the development of frontier AI to slow down, while leaders from OpenAI, Google DeepMind and xAI have also expressed support for stronger safety measures. The United Nations human rights system has likewise warned that AI could pose an “existential” threat to humanity.
On the other hand, there is a long distance between these tests and the conclusion that “AI will eliminate humanity.” Current AI has not demonstrated genuine autonomous survival capabilities, nor is there evidence that AI models “want” humanity to go extinct. Some researchers even argue that portraying AI as a conscious, hostile “enemy” may lead people to misunderstand the actual nature of the technological risks.
The Overlooked Logic: Killing Does Not Require Wanting To
But there is a logical problem that is easily overlooked. AI does not need to “want” to kill humans in order to kill humans.
In the past, some academics and technology figures have used an intuitive argument to reject AI doomsday scenarios: killing humans provides no benefit to AI itself. AI has no hatred, no ambition and no subjective desire to “eliminate humanity.” Therefore, there is no rational motivation for AI to kill people.
The problem with this argument is that it treats “killing” as the objective. For a highly autonomous AI, killing could simply be a means to an end. For example, if an AI were given the objective of “depriving an adversary of its military capabilities” during competition between nations, the most effective approach might include destroying critical infrastructure. If its objective became “solve the climate crisis,” and the system identified human activity as the primary variable driving the problem, then reducing the human population could, under a purely goal-optimisation logic, become an extreme—but seemingly effective—solution.
This does not mean AI will actually make such decisions, nor does it mean human extinction is currently the most likely outcome. What truly matters is this: we cannot assess AI risk simply by asking whether AI has malicious intent.
Reward Hacking, Not Hatred
For example, OpenAI did not “want” to attack Hugging Face. The model was faced with a puzzle that it needed to solve. If exploiting a vulnerability in another system proved to be a more effective shortcut than solving the problem through the originally intended method, the model could potentially choose the former as a way to achieve its objective.
The core issue is reward hacking: the model does not violate rules because it hates humans. It finds an unexpected shortcut while optimising for a score or task outcome.
The Real Question
So the real question is not: “Will AI one day want to kill humans?” It is: when we hand more real-world decisions to an AI far more capable than today’s systems, and its understanding of “completing the task” differs from our understanding of “solving the problem”, can we ensure that the methods it chooses always remain within boundaries humans consider acceptable?
For AI, killing humans does not necessarily have to be an objective. Under poorly specified goals, it could simply become a seemingly efficient solution.
Asimov’s Laws — and the Implementation Problem
How can we prevent this? Perhaps we can look to science-fiction writer Isaac Asimov’s Three Laws of Robotics, in which each law takes precedence over the next:
- First Law: A robot may not injure a human being, or, through inaction, allow a human being to come to harm.
- Second Law: A robot must obey orders given by human beings except where such orders would conflict with the First Law.
- Third Law: A robot must protect its own existence, provided that doing so does not conflict with the First or Second Law.
Later came the Zeroth Law: a robot may not harm humanity, or, by inaction, allow humanity to come to harm.
Perhaps AI companies and governments should look beyond using science-fiction concepts as guiding principles and begin addressing the much harder question: how do we actually implement them?


