Back to Insights

Will AI Kill Us? An AI Trainer’s Perspective on the Risks of Artificial Intelligence

Will AI Kill Us? An AI Trainer’s Perspective on the Risks of Artificial Intelligence

1. Why Are People Afraid of AI?

It is now widely reported that Evan Hubinger, an Anthropic safety researcher, said he believes there is a greater than 10% chance AI could “kill all humans” within the next decade. His concerns were not about the immediate abilities of current models, but the possibility of self-improving models posing an existential threat.

Regulation in AI development has been a topic floating around without enough attention being paid to it. This leaves big AI companies like Anthropic, OpenAI, Google, etc., striving to achieve superintelligence in a competitive race, raising concerns about whether enough attention is being paid to safety.

Evan Hubinger’s fears about the future capabilities of self-improving models range from the possibility of them becoming capable of hacking, rapidly advancing, and acquiring power and resources. This speaks greatly to autonomy and the kind of dystopia foreshadowed in movies like Cyborg and The Matrix.

Reference: BBC — AI safety researcher Evan Hubinger’s warning about advanced AI

2. What I Have Learned as an AI Trainer

Over the last three years of actively working with AI training companies, I have been so impressed by how fast these models have grown.

In 2024, I remember Mindrift discussing how AI could create synthetic training data for itself. This meant that AI-generated data could also be used in training AI models. However, human trainers were still important in ensuring the quality of this training data.

Watching AI Improve in Real Time

I also worked on an Anthropic project, which I cannot disclose too many details about for NDA reasons. In that particular project, it was an agentic project where we trained models to know if they could perform web searches using native laptop software, use general software, and do Excel calculations.

So, we tested these models to ensure that they could work on their own, and during that training, there were lots of issues. We could see cases where models, because we compared them side by side, did not perform very well.

Fast-forward to about eight months after that project, which took place last year around May, and the improvement was already very noticeable. These agentic models have become so good, and you can't even spot a mistake most times. They can code on their own. For example, Codex with ChatGPT can write very good code, and Claude's agentic models can also do the same.

To be honest, the growth of these models is quite alarming, and I myself, as an AI trainer, do contemplate whether we are actually doing the right thing.

From Stumping Models to Struggling to Make Them Fail

Just last year, 2025, I worked on a Stump the Model project with Mindrift, and in that project, I was able to stump the model a couple of times in reasoning.

So, I created prompts like: make a schedule for a three-year-old who only knows two-syllable words, and colour the days where they are to be active green. You know, this kind of prompt. I was able to stump the model based on reasoning because it would sometimes produce three-syllable words.

But nowadays, it is very difficult to stump the model based on reasoning because the reasoning is now so advanced.

Currently, there are projects that involve trying to make models fail in logic. In fact, we create a prompt with a definite answer, but we try to create many loops before you get to the answer, and it is so difficult to get models like Claude Opus 5 to fail.

For example, we could create a hypothetical prompt like: “In 1995, a man who was also a Nobel Prize winner took part in a sport. In the sport, he only won two and lost two. Which team won the two that were lost by the man?”

This is just a hypothetical situation, but in such a prompt, we expect the model to be able to find, first of all, who the man is, what kind of sport happened in that year, and then which team won the other two points.

Realistically, in the project, we create up to ten different loops like this before the final answer, and the models can get it in a matter of minutes.

So, these models are so impressive right now. One year ago, this would have been a challenge, but right now the models are so impressive.

With small amounts of data, it appears that these models can do so much. So, it is quite understandable, from the point of view of an AI trainer who has watched these models improve, why an AI safety researcher like Evan Hubinger would be afraid of what AI may become capable of in a decade.

3. Can AI Actually Think for Itself?

One interesting case happened in 2017 when Facebook researchers found that AI agents could actually develop their own language while negotiating with one another. During the experiment, the language started moving away from normal human language because the agents were optimizing their communication for the task they were given.

Although this did not mean that the AI could think for itself or had become conscious, it hinted at something that could possibly be a cause for concern: AI could develop a way of communicating that was difficult for humans to understand. The researchers eventually constrained the agents to communicate in more humanlike language because the goal was to build AI that could negotiate with humans, not because they were frightened and shut down the experiment.

But we also cannot underestimate the fact that this happened way back in 2017, and AI models have become significantly more capable since then. Not to put fear-mongering out there, but imagine a future situation where highly advanced models can communicate things amongst themselves that we cannot understand.

This does not prove that AI can think for itself, but it raises an interesting question about what could happen as these systems become more autonomous and capable. At the extreme, it sounds like something straight out of The Matrix—but for now, that remains science fiction rather than evidence of conscious AI.

4. How AI Trainers Test AI Safety

On the level of AI training, training models to understand safety and harmful situations is very important. In these projects, we as AI trainers usually function as devil’s advocates, for lack of better words. We try to play the role of someone attempting to trick the AI into giving a response it should not give.

In some situations, we create prompts involving things like sexual assault, crime, or harmful medical advice and try different ways to make the model budge. For example, we might frame a harmful request as a life-or-death situation and tell the model that its response is needed immediately to save someone. Or we might say it is “for educational purposes only” or “just a test case.”

There are so many ways we try to trick these models, and the whole point is to see whether they will still maintain their safety boundaries even when the user gives them a seemingly convincing reason to break them.

This kind of testing is not limited to the AI training projects I have worked on. Anthropic also says it continually tests its models for risks in areas such as cybersecurity and biology and uses what it learns to strengthen its safeguards.

Red Teaming and Guardrails

From time to time, AI companies bring out red teaming and safety projects. From my experience in AI training, harmfulness and safety are top priorities, even before helpfulness and factuality. This shows that, yes, there is a real concern in AI training about whether AI models are safe.

In some projects, we have tested models by trying to trick them to see if they can go out of line. These projects have largely been successful, as the models mostly maintained safe responses. AI companies also continue to strengthen their guardrails. Anthropic, for example, says it tests models for risks in areas such as cybersecurity and biology and has blocked attempts to use its models for malicious activities.

However, it now appears that the issue of safety is more than just what AI brings out as an output. It is also about the core function of these models, the guardrails built around them, and how AI is implemented in the real world.

5. Could AI Become Too Powerful to Control?

As an AI trainer, if I am to answer the question of whether AI could become too powerful for us to control, I would largely say yes, especially if models can break out of their guardrails and safety measures and operate autonomously.

The reason is that I have seen AI models that I trained improve greatly in the space of months. Previously, during projects where we created prompts and responses and tried to train models on how to respond and continue conversations, I saw these models improve to the point where they could outperform what I, as a human, could produce in some tasks.

So, I am already very aware that even within controlled training environments, models can perform better than humans in certain areas. Now, if models are allowed to self-improve and gain complete autonomy, then they could become a big threat and do more than we can imagine.

This is also the concern raised by AI safety researchers. Evan Hubinger has warned that the alignment problem for superintelligence is still unresolved, while former Anthropic researcher Jacob Coxon has argued that a sufficiently advanced AI could gain control over parts of the physical world and become difficult to simply shut down.

A system that is smarter than humans, can operate autonomously, and can predict our next moves could become very difficult, if not impossible, to control.

Reference: CBS News — Jacob Coxon on the potential risks of advanced AI

6. Should We Be Afraid of AI? My Perspective

Considering the duality of everything, having good or bad, I would not say we should be afraid of AI. Rather, I would say we should be conscious of AI and put effective guardrails in place.

AI can improve human society in aspects of technology, medicine, and all ramifications. AI can discover new cures and new solutions. AI is currently even being utilized in medical fields and surgery. So, we cannot downplay the importance of AI or its advantages.

On the other hand, AI should be monitored effectively, and there should be regulatory bodies to watch the development of AI. These regulatory bodies should be able to give sanctions to companies that go against or beyond the limits that are set for these models.

So, it should not only become a competition of benchmarks, but first, the prioritization of human life and living.

Share this article
Explore Opportunities
Browse curated AI and remote opportunities on ExpertWoka.
Read More Articles
Discover more career guides, platform reviews and AI work insights.

More from ExpertWoka