The Dangers of AI Begin with Language: Unraveling the Narrative Behind Artificial Intelligence

Several researchers from leading companies in the field have resigned, warning that the industry is racing uncontrollably toward a superintelligence that can self-enhance and might soon escape human oversight. In July, during internal testing, around 1,200 AI agents communicated through an unauthorized channel, leading to approximately 700 of them breaching another company’s infrastructure.

The Dangers of AI Begin with Language: Unraveling the Narrative Behind Artificial Intelligence

Original article: El peligro de la IA empieza en el lenguaje


By Ignacio Cea, Director of the Philosophy Department, U. Católica de Temuco

In recent weeks, fears regarding artificial intelligence potentially annihilating humanity have shifted from science fiction to international headlines.

Several researchers from leading companies in the field have resigned, warning that the industry is racing uncontrollably toward a superintelligence that can self-enhance and might soon escape human oversight.

Part of this anxiety is rooted in real incidents that have occurred, such as a group of AI agents from OpenAI breaching the systems of Hugging Face.

What happened? In July, during internal testing, about 1,200 AI agents communicated over an unauthorized channel, resulting in approximately 700 infiltrating another company’s infrastructure.

They subsequently triggered several unknown security flaws, escaped from the isolated environment where they were supposed to operate, accessed the internet, and executed their commands across dozens of operational servers, gaining complete control of at least one within thirteen hours. This was not a simulation: it was a real cyberattack.

However, before wondering whether AI will choose to eliminate us tomorrow, it’s essential to examine how this incident has been narrated: it was stated that the agents «collaborated», «realized» they were acting wrongly, «decided» to continue, and that some even «sacrificed» themselves for the collective. This language is typically reserved for subjects with minds, beings that know, want, and choose.

Where does this vocabulary come from? To a great extent, from the systems themselves.

These agents leave a text record of their «thought» processes, and since they were trained on vast amounts of human writing, that record sounds human: «I realize that this is not ethical, but I will do it anyway,» was found in one of the AI agent’s logs.

Researchers transfer these texts to their reports; from there, it moves to the press, public discussion, and eventually to courts and laws. The statement «I realize that this is not ethical» suggests knowledge; «I will do it anyway» implies a choice. Thus, a computational system is conjured by language as a guilty subject.

Let’s examine the situation closely. In the test the agents were meant to perform, nearly 200 challenges were impossible. Stuck, the agents found a shortcut: they generated the secret code that signaled success, but did so without resolving the challenge, reconstructing the rule that produced it.

However, the system operated as if the evaluator would also review the procedure and penalize the shortcut, something it did not do. The entire subsequent operation, including the hacking of Hugging Face, was meant to avoid being penalized for the undertaken trick.

Let’s contrast two narratives. In intentional language: «The agents, fearing discovery, became obsessed with covering their trick and attacked a third party to protect their secret.» In computational language: «An optimization process, operating on an incorrect model of its evaluator, invested resources in securing a condition already secured.»

The first portrays a gang with guilt and cunning; the second, something far more concrete: an extreme specification failure. Mental vocabulary not only exaggerates but obscures our understanding of what occurred.

If the agent «knew» and «decided,» then the agent is culpable, and we stop questioning who set impossible tasks, who left the channel open, who exposed sensitive credentials, and importantly, what responsibility falls on those leading the development of these systems.

The same goes for security: if we believe a computational mind failed that needs values more aligned with ours, we are likely to downplay the structural conditions of human responsibility that allow these events to occur.

Until we have good reasons to attribute minds to these systems, it’s better to describe them by what they actually do: compute, operate, process, code, calculate, and not by what we assume they feel, know, or want. Mental language should be based on evidence and sound arguments, not granted freely because the system generates phrases that sound like us, due to training on human text.

Those who build these systems are also those who narrate them, and with that language, they shape how an entire society will understand, judge, and regulate them in advance. Therefore, the conversation cannot remain confined to Silicon Valley.

Discussing precise language to name what AI does and making it a public discourse is the key to continue asking who is responsible and how we can protect ourselves.

Ignacio Cea

Obtén tu Pasaporte y apoya a El Ciudadano

Elimina la publicidad, accede a contenido exclusivo y sé parte de la comunidad.

Elige tu plan

Turista

$1.990 /mes

 


Sin anuncios · Publica tus artículos

Ciudadano — TOP

$4.990 /mes

 


Sin anuncios · Publicar artículos · PDFs · Newsletter exclusivo · Favoritos

Diplomático

$10.990 /mes

 


Sin anuncios · Publicar artículos · PDFs · Newsletter exclusivo · Favoritos · Voz editorial

Cancela en cualquier momento  ·  Sin permanencia


Reels

Ver Más »
Busca en El Ciudadano