From Bruce Schneier: "All it takes to poison AI training data is to create a website:

meigahub@mastodon.social

Este ejemplo muestra cómo la data sesgada o falsa puede entrenar a los LLMs. ¿Qué mecanismos podrían implementarse para validar la fuente de los datos de entrenamiento?

petealexharris@mastodon.scot

@tml @Yendolosch @emacsomancer

Broadly fair usage. Got someone else's computer system to behave in a way they didn't want it to. The only stretch is that there's an implication in "hacked" that some safeguards had to be bypassed, and there weren't any in the first place. But that's worse, right?

sergiudinit@mastodon.social

This is a genuinely scary insight from Schneier. The implications for AI reliability go way beyond just training data quality. What happens when adversarial training becomes industrialized?

bearsong@ravenation.club

@emacsomancer

"Ned Ludd's in your datacentre, poisoning your training sets!"

https://ravenation.club/@bearsong/116104233823870563

larsbrinkhoff@mastodon.sdf.org

@petealexharris @tml @Yendolosch @emacsomancer It's rather close to the original usage of the word "hacked". Some still use it like that.

gnomeoffender@mastodon.social

@emacsomancer they aren't trustworthy. Take up a lot of time trying to get a reasoned answer and there's always a phrase or wording out of place that needs correction. Almost as it the AI is trying to engage longer and longer than necessary.

darknetdon@mastodon.social

@emacsomancer to be honest i am not well-informed enough to definitively judge the accuracy of this, but it seems wrong for 2 main reasons.

1. models dont train on the fly, typically, yet, so for models to behave as such in such a short period of time seems inaccurate and would require web search enabled and explicitly directed to disregard other search results.

2. people training these models know conflicting info is everywhere and the source of truth is prioritized in training algorithms.

kneoghau@mastodon.social

@emacsomancer How is this a news story, beyond "ai bad"? In the dial up days people falsely believed everyone ate 9 spiders a year in their sleep due to chain emails.

photo55@mastodon.social

@emacsomancer
Shall we have an algorithmic bullshit generator?

And pass around multiple copies of it, identical and with small changes, omissions and additions?

sorro@woof.tech

@emacsomancer in less than 24 hours the chatbots fell for the experiment, and less than 24 hours after it was revealed what the experiment was about, that information has ALSO become part of the training data

are they constantly scrapping websites for training data or why does this appear here so fast??? no wonder those datacenters consume so much electricity if they dont take a single break from scrapping the internet

duco@norden.social

@larsbrinkhoff @petealexharris @tml @Yendolosch @emacsomancer in the sense of life hacks or food hacks this is an AI hack. So the AI has been hacked.

gim@lou.lt

@emacsomancer it's not really a new thing Russians are already using this technique to poison training data:

https://thebulletin.org/2025/03/russian-networks-flood-the-internet-with-propaganda-aiming-to-corrupt-ai-chatbots/

Edit: there is some newer reporting on that matter, but I can't find it right now/don't have it anywhere at hand

w@mountains.social

@emacsomancer He also poisoned the data for everyone who searches for hot dog eating competetitors online in other ways. I'm not sure what he accomplished.

Piero Bosio Social Web Site Personale

From Bruce Schneier: "All it takes to poison AI training data is to create a website:

Feed RSS

Gli ultimi otto messaggi ricevuti dalla Federazione

Post suggeriti

Un paio di link collegati tra loro:

Select all statements you agree with.

Alleï, j’avais promis un thread de meme #IA #AI

The #ocaml debacle with the ginormous slop #AI / #LLM "contributions" continues!