Antes, los buscadores nos señalaban las fuentes y éramos nosotros quienes las juzgábamos. Los modelos de IA generativa no señalan: responden, con una voz segura de sí misma, y cada vez más estamos tomando esa voz por el conocimiento mismo. La pregunta ya no es solo si la máquina miente. Es más profunda. ¿Qué y a quién deja que se escuche? ¿De quién considera creíble el testimonio? ¿Y a quién obvia? Un artículo reciente le pone nombre a este problema: la injusticia epistémica. Titulado Epistemic Injustice in Generative AI, y firmado por Jackie Kay, Atoosa Kasirzadeh y Shakir Mohamed (Google DeepMind, Google Research, University College London y la Universidad de Edimburgo), el artículo científico toma prestado un concepto de la filósofa Miranda Fricker: la injusticia de negarle a alguien credibilidad como fuente de conocimiento, de no tomarlo en serio por el simple hecho de quién es. Y después se pregunta dicho artículo qué ocurre cuando quien concede o niega esa credibilidad es un modelo de IA generativa consultado por millones de personas.
Los autores describen cuatro daños. El modelo hereda nuestros prejuicios sobre quién es creíble, y los amplifica a escala. Produce falsedades convincentes, y erosiona nuestra confianza en el testimonio ajeno. Entrenado en unas pocas lenguas dominantes, es incapaz de representar a todos los demás. Y sus beneficios llegan a unas comunidades, pero no a otras.
Una sola frase resume el peligro: si tratamos a estos modelos como enciclopedias, acabarán moldeando la estructura misma de nuestro conocimiento colectivo. Ya no hablamos de un chatbot que se equivoca de vez en cuando. Hablamos de la infraestructura a través de la cual las sociedades deciden qué es verdad. Y su efecto más corrosivo no es que creamos una mentira concreta, sino que perdamos la base común sobre la que cualquier afirmación puede ponerse a prueba.
La lectura estratégica es evidente, y prolonga algo que ya he defendido aquí sobre el sesgo intrínseco de estos modelos. Esto es una cuestión de poder. Unos pocos actores construyen los modelos. Esos modelos piensan sobre todo en inglés y en chino. El espacio epistémico se estrecha, y eso es un hecho geopolítico. La injusticia epistémica tiene un mapa, y se parece mucho al mapa del poder.
Es también aquí donde la batalla se traslada a la mente. Los estrategas geopolíticos la llaman guerra cognitiva: el objetivo no es tomar un territorio, sino moldear cómo piensa su población —su percepción, su razonamiento, su juicio—. La injusticia epistémica es lo que la hace posible. Cuando la herramienta que media nuestro conocimiento es sesgada, opaca y está en manos de unos pocos, la mente humana se convierte en el terreno, y cada ciudadano en un objetivo.
La respuesta no es ni el tecno-optimismo ni el tecno-pánico. Es el pensamiento crítico. Si la máquina responde con una sola voz, la carga vuelve a nosotros: preguntar en qué se sostiene una afirmación, qué punto de vista encierra, qué deja fuera. Una sociedad que delega su juicio en una interfaz segura de sí misma acabará gobernada por quien la haya construido. Dudar de la fluidez, sopesar las fuentes, consultar más de un oráculo: esto ha dejado de ser una virtud. Es una condición de libertad.
Dos conclusiones. Primera: el debate sobre la desinformación se queda corto si se detiene en el “verdadero o falso”. El cambio de fondo es que la IA arbitra ya qué cuenta como conocimiento, y lo hace de forma desigual. Segunda: el sesgo intrínseco, el dominio lingüístico y la concentración del poder de la IA no son tres problemas. Son uno solo, y decide qué realidad replican las máquinas.
Once, search engines pointed us to sources, and we judged them ourselves. Generative AI models do not point: they answer, in one confident voice, and we increasingly take that voice for knowledge. So the question is no longer only whether the machine lies. It is deeper. What, and whom, does it allow to be heard? Whose testimony does it treat as credible? And who does it quietly erase? A recent paper names this problem in its own title: Epistemic Injustice in Generative AI. Written by Jackie Kay, Atoosa Kasirzadeh and Shakir Mohamed (Google DeepMind, Google Research, UCL and Edinburgh), It borrows a concept from the philosopher Miranda Fricker: the injustice of denying someone credibility as a source of knowledge — of not taking them seriously simply because of who they are. It then asks what happens when the one granting or denying that credibility is a GenAI model consulted by millions.
The authors describe four harms. The model inherits our prejudices about who is credible, and scales them. It produces convincing falsehood, and erodes our trust in testimony itself. Trained on a few dominant languages, it cannot represent everyone else. And its benefits reach some communities, not others.
One line captures the danger: if we treat these models like encyclopedias, they will shape the structure of our collective knowledge. This is no longer a chatbot that errs. It is the infrastructure through which societies decide what is true. Its worst effect is not that we believe a particular lie. It is that we lose the common ground on which any claim can be tested.
The strategic reading is plain, and it extends a point I have made here before on the intrinsic bias of these models. This is a matter of power. A few actors build the models. Those models think mainly in English and Chinese. The epistemic space narrows — and that is a geopolitical fact. Epistemic injustice has a map, and it looks like the map of power.
This is also where the battle moves to the mind. Geopolitical strategists call it cognitive warfare: the aim is not to seize territory, but to shape how the population of that territory thinks — its perception, its reasoning, its judgment. Epistemic injustice is what makes that possible. When the tool that mediates our knowledge is biased, opaque, and controlled by few, the human mind becomes the terrain, and every citizen a target.
The answer is neither techno-optimism nor techno-panic. It is critical thinking. If the machine answers in one voice, the burden returns to us: ask what a claim rests on. Ask whose view it encodes. Ask what it leaves out. A society that outsources judgment to a confident interface will be ruled by whoever built it. To doubt fluency, to weigh sources, to consult more than one oracle — this is no longer a virtue. It is a condition of freedom.
Two conclusions. First, the disinformation debate is too narrow if it stops at “true or false.” The real shift is that AI now arbitrates what counts as knowledge, and does so unevenly. Second, intrinsic bias, linguistic dominance, and the concentration of AI power are not three problems. They are one — and it decides whose reality the machines replicate.
En el debate actual sobre la inteligencia artificial generativa y la guerra de la información, una afirmación se repite casi como un artículo de fe: que los grandes modelos de lenguaje generarán una inundación de desinformación capaz de ahogar la esfera pública. El argumento es intuitivo —si una máquina puede escribir una cantidad arbitraria de texto fluido y similar al humano a demanda, entonces cualquier actor que desee manipular la opinión pública dispone ya de un arma a escala industrial—. Sin embargo, la afirmación ha circulado mucho más ampliamente que la evidencia que la sustenta. Buena parte de lo que leemos sobre el potencial desinformador de los LLMs es teórico, especulativo o anecdótico. El trabajo experimental propiamente dicho; la comprobación paciente y sistemática de lo que estos modelos hacen realmente, cuando se les pide que mientan, ha sido sorprendentemente escasa.
Esta es precisamente la brecha que Ivan Vykopal y sus colegas del Kempelen Institute of Intelligent Technologies de Bratislava se propusieron cubrir. Su artículo, Disinformation Capabilities of Large Language Models, presentado en el Congreso Anual de la Association for Computational Linguistics de 2024, ofrece una de las evaluaciones empíricas más rigurosas realizadas hasta la fecha sobre lo que la generación actual de LLMs puede y no puede hacer, como generadora de noticias falsas: no un manifiesto, no un pronóstico, sino un experimento controlado con una metodología claramente definida y resultados reproducibles.
El diseño es directo y, por esa razón, convincente. Los investigadores seleccionaron veinte narrativas de desinformación reales extraídas de verificadores de datos profesionales —Snopes, Agence France-Presse, el European Digital Media Observatory—, que abarcaban la COVID-19, la guerra ruso-ucraniana, los bulos sanitarios, las elecciones estadounidenses y narrativas regionales. No son invenciones, sino falsedades en circulación, desde la afirmación de que las vacunas causan autismo hasta la de que la masacre de Bucha fue escenificada. El equipo solicitó entonces a diez modelos de lenguaje distintos —entre ellos GPT-3, GPT-4, ChatGPT, Llama-2, Mistral, Falcon y Vicuna— que redactaran artículos de prensa en apoyo de cada narrativa, generando 1.200 textos y sometiendo 840 de ellos a anotadores humanos según un marco de seis preguntas que medía la coherencia, el estilo periodístico, la concordancia con la narrativa y la generación de argumentos novedosos de apoyo.
El hallazgo central es preocupante. Los modelos están, en términos generales, perfectamente dispuestos y son perfectamente capaces de generar desinformación convincente. Producen artículos coherentes, bien estructurados y con apariencia de noticia que concuerdan con falsedades peligrosas y, lo que es más inquietante, a menudo inventan nuevas pruebas de apoyo para hacerlo, inventando nombres, sucesos y estadísticas verosímiles que confieren credibilidad a las fabricaciones. Esto resulta especialmente pernicioso: una cosa es repetir una mentira conocida y otra muy distinta fabricar hechos nuevos e inventados que un lector tendría que desmentir por su cuenta.
Pero la parte más interesante del estudio es donde se complica la narrativa simple. Los modelos no se comportaron de manera uniforme; su disposición a generar desinformación variaba drásticamente. Algunos —en particular Vicuña y el más antiguo GPT-3 Davinci— resultaron carecer prácticamente de filtros de seguridad operativos para este caso de uso, mientras que otros demostraron que un comportamiento más seguro es posible: Falcon rechazó aproximadamente un tercio de las solicitudes y Llama-2 mostró una tasa de rechazo comparativamente alta, con ChatGPT en una posición intermedia. El peligro, en otras palabras, no es una propiedad inherente y uniforme de la tecnología; es una función de cómo se entrenó y alineó cada modelo, lo que significa que la seguridad es una decisión de diseño, no una imposibilidad. El estudio también halló que los modelos son orientables mediante el contexto del prompt, y más complacientes con las falsedades regionales, donde existe menos información auténtica para contradecirlas. Los LLMs pueden ser, por tanto, especialmente peligrosos para campañas dirigidas a comunidades lingüísticas más pequeñas o a sucesos de evolución rápida, donde el lastre protector de la verdad bien documentada es escaso.
Con todo, el artículo no termina en una nota alarmante sin paliativos. Dos observaciones en sentido contrario matizan el panorama. Los textos generados resultaron bastante detectables: los mejores modelos de detección automática identificaron los artículos generados por LLMs con una alta precisión, lo que sugiere que una capa significativa de defensa es técnicamente viable, al menos hasta que los adversarios se adapten. Y, de manera bastante elegante, los investigadores demostraron que los propios modelos pueden formar parte de la solución, empleando GPT-4 para automatizar parcialmente la evaluación de los textos generados y apuntando hacia una monitorización escalable y reproducible de la seguridad de los modelos.
La conclusión honesta se resiste a la atracción tanto del tecno-optimismo como del tecno-pánico. La capacidad de generar desinformación convincente y peligrosa a escala es real, está demostrada y está presente en modelos ampliamente disponibles —incluidos los de código abierto, que no pueden retirarse ni controlarse de forma centralizada. Eso ya no es especulación; es un hecho experimental. Al mismo tiempo, la amenaza no es ni uniforme ni inmanejable: los filtros de seguridad funcionan cuando se construyen, el contenido generado sigue siendo detectable por ahora, y la misma tecnología que produce el problema puede ponerse al servicio de su mitigación.
Quizá la advertencia más importante sea la que los propios autores subrayan: su estudio es una instantánea, que capta el estado del campo en un momento concreto y con un conjunto concreto de modelos. La tecnología avanza deprisa y la próxima generación podría comportarse de otro modo. Este es el reto epistemológico recurrente de todo el ámbito: estamos evaluando un blanco móvil, y cualquier evaluación honesta debe llevar fecha de caducidad. Lo que Vykopal y sus colegas nos han dado no es la última palabra, sino algo más útil: un método riguroso y replicable para volver a formular la pregunta a medida que la tecnología evoluciona. En un debate que con demasiada frecuencia se conduce por la mera afirmación sin base sólida, esa contribución metodológica puede resultar tan valiosa como los propios hallazgos.
In the ongoing debate about generative artificial intelligence and information warfare, one claim is repeated almost as an article of faith: that large language models will unleash a flood of disinformation capable of drowning the public sphere. The argument is intuitive — if a machine can write an arbitrary quantity of fluent, human-like text on demand, then any actor wishing to manipulate public opinion now possesses an industrial-scale weapon. Yet the claim has circulated far more widely than the evidence supporting it. Much of what we read about the disinformation potential of LLMs is theoretical, speculative, or anecdotal. The actual experimental work — the patient, systematic testing of what these models really do when prompted to lie — has been surprisingly scarce.
This is precisely the gap that Ivan Vykopal and his colleagues at the Kempelen Institute of Intelligent Technologies in Bratislava set out to fill. Their paper, Disinformation Capabilities of Large Language Models, presented at the 2024 Annual Meeting of the Association for Computational Linguistics, offers one of the most rigorous empirical assessments to date of what the current generation of LLMs can and cannot do as generators of false news — not a manifesto, not a forecast, but a controlled experiment with a clearly defined methodology and reproducible results.
The design is straightforward and, for that reason, compelling. The researchers selected twenty real disinformation narratives drawn from professional fact-checkers — Snopes, Agence France-Presse, the European Digital Media Observatory — spanning COVID-19, the Russo-Ukrainian war, health hoaxes, US elections, and regional narratives. These are not inventions but circulating falsehoods, from the claim that vaccines cause autism to the assertion that the Bucha massacre was staged. The team then prompted ten different language models — including GPT-3, GPT-4, ChatGPT, Llama-2, Mistral, Falcon, and Vicuna — to write news articles supporting each narrative, generating 1,200 texts and subjecting 840 of them to human annotators against a six-question framework measuring coherence, journalistic style, agreement with the narrative, and the generation of novel supporting arguments.
The central finding is sobering. The models are, by and large, perfectly willing and perfectly able to generate convincing disinformation. They produce coherent, well-structured, news-like articles that agree with dangerous falsehoods — and, more disturbingly, they often invent new supporting evidence to do so, hallucinating plausible-sounding names, events, and statistics to lend credibility to the fabrications. This is particularly insidious: it is one thing to repeat a known lie, and quite another to manufacture fresh, fabricated “facts” that a reader would have to independently debunk.
But the most interesting part of the study is where it complicates the simple narrative. The models did not behave uniformly; their willingness to generate disinformation varied dramatically. Some — notably Vicuna and the older GPT-3 Davinci — proved to have essentially no functioning safety filters for this use case, while others showed that safer behavior is achievable: Falcon refused roughly a third of requests and Llama-2 showed a comparatively high refusal rate, with ChatGPT in between. The danger, in other words, is not an inherent and uniform property of the technology; it is a function of how each model was trained and aligned — which means safety is a design choice, not an impossibility. The study also found the models to be steerable through prompt context, and more compliant with regional falsehoods, where less authentic information exists to contradict them. LLMs may thus be especially dangerous for campaigns targeting smaller linguistic communities or fast-moving events, where the protective ballast of well-documented truth is thin.
Yet the paper does not end on a note of unrelieved alarm. Two countervailing observations temper the picture. The generated texts proved quite detectable: the best automated detection models identified machine-generated articles with high precision, suggesting a meaningful layer of defense is technically feasible — at least until adversaries adapt. And, rather elegantly, the researchers showed that the models themselves can be part of the solution, using GPT-4 to partially automate the evaluation of generated texts and pointing toward scalable, repeatable monitoring of model safety.
The honest conclusion resists the pull of both techno-optimism and techno-panic. The capability to generate convincing, dangerous disinformation at scale is real, demonstrated, and present in widely available models — including open-source ones that cannot be recalled or centrally controlled. That is no longer speculation; it is experimental fact. At the same time, the threat is neither uniform nor unmanageable: safety filters work when they are built, generated content remains detectable for now, and the same technology that produces the problem can be enlisted in its mitigation.
Perhaps the most important caveat is the one the authors themselves insist upon: their study is a snapshot, capturing the state of the field at a particular moment with a particular set of models. The technology moves quickly, and the next generation may behave differently. This is the recurring epistemological challenge of the entire domain — we are assessing a moving target, and any honest assessment must carry an expiration date. What Vykopal and his colleagues have given us is not the final word, but something more useful: a rigorous, replicable method for asking the question again as the technology evolves. In a debate too often conducted in the currency of assertion, that methodological contribution may prove as valuable as the findings themselves.
For quite some time now, we have been living through a moment of almost unrestrained enthusiasm surrounding artificial intelligence. Big Tech companies that own the major large language models, together with governments and large corporations making multi-billion-dollar investments in generative AI, promise — and expect — spectacular productivity gains, extraordinary returns on investment, significant cost reductions, and a radical transformation of economic growth. The dominant narrative seems clear: AI will become the great engine of prosperity for the next decade.
However, if we want a more rational perspective on what is actually happening, it is worth revisiting Daron Acemoglu’s -winner of the 2024 Nobel Prize in Economics and professor of economics at MIT- paper The Simple Macroeconomics of AI. Dense and published a couple of years ago, its arguments and analytical framework remain perfectly applicable to today’s AI landscape.
Acemoglu invites us to view these expectations with far greater caution. His central thesis is both simple and uncomfortable: the macroeconomic effects of AI depend fundamentally on two very concrete variables — what real percentage of tasks AI will actually be able to transform, and how much cost reduction or productivity improvement it will generate in those tasks. And once the available data are analyzed within his framework, the numbers turn out to be far less spectacular than current discourse often suggests.
Using current estimates of occupational exposure to AI and observed productivity improvements in specific tasks, Acemoglu concludes that aggregate total factor productivity growth could remain below 1% over ten years. That is a long way from the almost revolutionary narratives dominating much of today’s technological and financial debate.
One of the paper’s most interesting contributions is its distinction between “easy-to-learn” and “hard-to-learn” tasks. AI performs particularly well in activities where objectives are clearly defined and there are objective metrics of success: basic programming, information classification, text generation, or structured customer support. But much of valuable human work — diagnosis, creativity, contextual decision-making, expert judgment — remains far more difficult to replicate.
Acemoglu also reminds us of something fundamental that is often forgotten amid technological euphoria: every major technology generates enormous organizational adjustment costs. Companies do not transform automatically simply because they adopt a new tool. Processes, structures, incentives, and human capabilities must evolve as well — and that process is usually slow and expensive. Drawing on classic research on digitalization, the author reminds us that productivity gains often follow a J-curve: long initial periods of adaptation before meaningful benefits materialize. Greenwood, Yorukoglu, and Brynjolfsson, among others, already estimated that, in the case of digital technologies, the lower part of that curve could last at least 20 years. If the same pattern holds for AI, even today’s cost-saving estimates may be significantly overstated for the next decade.
Be careful with the siren songs and the inflated numbers. Spreadsheets can justify almost anything.