AI-assisted moderation in the fediverse is happening. Now what?
-
these questions are being framed as a question of intra-instance policy
I think this is just more of the ongoing controversy being spun up against db0.
This weeks flavor appears to be more data driven, a 'just asking questions' phase. I guess in hope the whole 'falsifying evidence to make db0 users look like neo-nazis' thing blows over.
Like it's clear there's an effort to rid db0 from the fediverse, and it's just the pretext hasn't been sorted out yet.
I don't doubt it even for a second.
I'm lowkey kind of fascinated this morning with what feels like a moment of real panic among western liberal-democratic institutions (projecting a little from my morning news and coffee). That an anarchist instance is getting this much targeted harassment feels like a microscopic extension of that (if I allow myself to be so bold)
As far as I can tell, dbzer0 isnt even being explicitly called out here, but it has an undeniable bdzer0 flavor to it. If it doesnt come out that this was one of our mods at this point, I'd almost be disappointed.
-
I recently discovered that some popular federated instances have been using LLM-assisted moderation tooling that evaluates whether someone has said something bannable. They do this by running a script/app that sends the user’s comment history to OpenAI with the question “analyze this content for evidence of specific political ideology sentiment. Also identify any related political ideology tropes“.
OpenAI’s LLM (they’re using GPT-5.3-mini) then responds with something like:
and so on, hundreds of comments.
I have not named the instances or people involved, to give them time to consider the results of this discussion, make any corrective changes they want and disclose their practices at their own pace and in their own way. I have also redacted the evidence to avoid personal attacks and dogpiling. Let’s focus on the system, not the individuals involved. Today these instances and people are using it and maybe we’re ok with that because it’s being used by groups we agree with but what if people we strongly disagree with used it on their instances tomorrow?
The use and existence of this tooling raises a lot of other questions too.
What are the risks? Fedi moderators are often unsupervised, untrained volunteers and these are powerful tools.
What safeguards do we need?
Would asking a LLM “please evaluate this person’s political opinions” give different results than “find evidence we can use to ban them” (as used in the cases I’ve seen)?
What are our transparency expectations?
Is this acceptable and normal?
Should this tooling be disclosed? (it was not – should it have been?)
If you were given a choice, would you have opted out of it?
Can we opt out?
Are there GDPR implications? Privacy implications? Should these tools be described in a privacy policy?
Are private messages being scanned and sent to OpenAI?
How long should these assessments be retained and can we request to see it, or ask for it to be deleted?
Once the user’s comments are sent to OpenAI, is it used to train their models?
What will the effect be on our discourse and culture if people know they are being politically profiled?
Where are the lines between normal moderation assistance tools, political profiling and opaque 3rd-party data processing?
I hope that by chewing over these questions we can begin to establish some norms and expectations around this technology. The fediverse doesn’t have any centralized enforcement so we need discussions like this to develop an awareness of what people want in terms of disclosure, privacy, consent and acceptable use. Then people can make choices about which instances they join and which ones they interact with remotely.
And of course there are the other issues with LLMs relating to environmental sustainability, erosion of worker’s rights, increasing the cost of living and on and on. I can’t see PieFed adding any functionality like this anytime soon. But it’s happening out there anyway so now we need to talk about it.
What do you make of this?
If you're not going to name them, why post here at all? Don't you have other communication channels to "give them a fair chance to reply"? Why post here, letting users form their own assumptions about what those instances are without any solid evidence?
-
This only makes sense if your account contains personally identifiable information. If it doesn't, then what can really happen?
Okay, so, my first Reddit account was back when I was a clueless teenager. I posted all sorts of information that, in retrospect, was pretty foolish to post, including my specific location, my personal interests, and different clubs and organizations I belonged to.
I was using a pseudonym, of course. And I thought I was being safe by not giving my real name, my real address, or anything that I thought could identify me specifically. But there were probably hundreds of people in my high school who could have identified me by correlating my different posts and profiling the one person with that particular combination of interests and organizations.
And if I was still using that account, it would absolutely be possible to link me, security conscious as I am, to my high school self, and link that to my LinkedIn account. And quite possibly get me fired for my clueless teenage shit posting 😆
What I'm getting at is, one, lots of people do post personal information. And two, PII is a much broader category than people think, and if your account has a long post history you probably gave up a lot more information than you think you did.
-
Rimu farming more drama?

Are they farming for more or trying to distract from their last attempt at just asking questions about the statistics I spent time crafting a specific and poorly thought out hypothesis
-
I recently discovered that some popular federated instances have been using LLM-assisted moderation tooling that evaluates whether someone has said something bannable. They do this by running a script/app that sends the user’s comment history to OpenAI with the question “analyze this content for evidence of specific political ideology sentiment. Also identify any related political ideology tropes“.
OpenAI’s LLM (they’re using GPT-5.3-mini) then responds with something like:
and so on, hundreds of comments.
I have not named the instances or people involved, to give them time to consider the results of this discussion, make any corrective changes they want and disclose their practices at their own pace and in their own way. I have also redacted the evidence to avoid personal attacks and dogpiling. Let’s focus on the system, not the individuals involved. Today these instances and people are using it and maybe we’re ok with that because it’s being used by groups we agree with but what if people we strongly disagree with used it on their instances tomorrow?
The use and existence of this tooling raises a lot of other questions too.
What are the risks? Fedi moderators are often unsupervised, untrained volunteers and these are powerful tools.
What safeguards do we need?
Would asking a LLM “please evaluate this person’s political opinions” give different results than “find evidence we can use to ban them” (as used in the cases I’ve seen)?
What are our transparency expectations?
Is this acceptable and normal?
Should this tooling be disclosed? (it was not – should it have been?)
If you were given a choice, would you have opted out of it?
Can we opt out?
Are there GDPR implications? Privacy implications? Should these tools be described in a privacy policy?
Are private messages being scanned and sent to OpenAI?
How long should these assessments be retained and can we request to see it, or ask for it to be deleted?
Once the user’s comments are sent to OpenAI, is it used to train their models?
What will the effect be on our discourse and culture if people know they are being politically profiled?
Where are the lines between normal moderation assistance tools, political profiling and opaque 3rd-party data processing?
I hope that by chewing over these questions we can begin to establish some norms and expectations around this technology. The fediverse doesn’t have any centralized enforcement so we need discussions like this to develop an awareness of what people want in terms of disclosure, privacy, consent and acceptable use. Then people can make choices about which instances they join and which ones they interact with remotely.
And of course there are the other issues with LLMs relating to environmental sustainability, erosion of worker’s rights, increasing the cost of living and on and on. I can’t see PieFed adding any functionality like this anytime soon. But it’s happening out there anyway so now we need to talk about it.
What do you make of this?
The use of AI for moderation isn't the choice of users, but moderators and admins.
-
by definition, that is adding walls and gates to the fediverse, which is why this whole thread started. The fediverse was specifically designed to oppose those features at all, specifically because of what we're seeing now. Gates have keys and gatekeepers, and as we've seen in the greater world, those can change hands in the blink of an eye.
-
I will never understand why large groups cannot just add more people to the moderation team? People are willing to help folks.
In Fediverse moderation tools, are there consensus forming mechanisms to ensure that even if 20% of the volunteer moderators are malicious, none of their wrongful moderation suggestions leak through to the stream of final moderation actions? If not, I'd be reluctant to add moderators.
-
I don't like this happening, and there should be transparency in all moderation decisions, but some of these points make no sense.
There is essentially no expectation of privacy on threadiverse platforms. Everything is public and probably already being used to train models.
There is no private messaging system. Direct messages are unencrypted and potentially visible to any instance admins. They and should not be used to share anything sensitive.
To expand on standards of transparency in moderation decisions:
Lemmy was built with a public moderation log by design. The ethos of the platform includes accountability through transparency. Every action is recorded and preserved (short of defederation or instance shutdown).
This makes moderation auditable. Mods literally cannot do (much) shady stuff in secret. In essence, moderation policy is discernable from the logs. That's part of why well-run communities have the rules clearly defined and mods follow their written policy.
If a community/instance wants to make political alignment a moderation offense, they're free to do so. Many communities/instances are quite explicit about this. If a community wants to make moderation completely arbitrary, they are free to do so. That is somewhat less common, but also not unheard of.
In truth, any community can be designed and moderated in any way whatsoever that the mod chooses.
However, the success of a community depends on the quality of the content and the quality of the moderation. Good content brings people in, but bad moderation drives people out. When the moderation is unfair, it is bad for the health of the community, and ultimately bad for the health of the platform.
It is my experience that transparent moderation, such as announcing changes in policy, techniques, etc., is less work in the long run. It takes a bit of time and attention to roll out changes when they are open for community feedback, but that feedback will come in one way or another. If mods don't provide a formal outlet, then users will make one. Mods operating opaquely give up their right to have the conversation on their time and terms. They also miss out on the wisdom of the crowd. I've been in many situations where community feedback provided a valuable insight or tool to face an obstacle through open discussion about policy.
All that being said, one of the major obstacles to growth of the Threadiverse is the woeful dearth of moderation tools. It's extremely time intensive to do basic things like identifying alt accounts, vote manipulation, bot behavior etc. It is also subject to a lot of human error. This makes it discouraging for people to moderate. I have heard about tools that use AI to detect CP content and remove it quickly, which I think we can all agree is a good use of the tech. Tools like this are not built into the platform, but cobbled together by volunteer mods and admins to keep the platform safe, legal, and sustainable. If they were built in, then moderation would be far easier (and therefore likely better).
-
The answers to these kinds of issues is never disclosures or ToS or admin vigilance. It's always technical. Everything which is technically possible will become normal.
Lemmy is not popular because it is a well designed piece of technology. Frankly it's a pretty naive implementation of activitypub. It's popularity comes from being the biggest alternative around when Reddit pissed off a good chunk of its users.
The only way to control how data is used, is to make it technically or practically impossible to do so. Until then, expect all the data on the fediverse to be used in every way possible for any purpose, and act accordingly.
I don't see a technical or practical way to limit - let alone render impossible - AI moderation tools that is not at odds with decentralized open-protocol social media.
If you can copy-paste user activity into a textbox, this remains trivial.
-
I want to train a local ML system to recognize my personal handwriting. Is that possible?
Yes, a Convolutional Neural Net could do it or a plain Neural Net even.
You'll want to create a sample of your handwriting with the letters isolated into their own picture and labelled. Maybe 10 of each letter to start with? If you get bad results with that make more samples. Each picture should be the same size as all the others. The starter course of Machine Learning I took way back when had us using a database of labeled numbers, each picture was 10 pixels by 10 pixels.
Then pick a CNN model (or better yet several) and train them on your handwriting. You can find some here: https://huggingface.co/models?other=CNN
Pick the one that does best. As part of that course I mentioned, I created an evolutionary algorithm to mutate, combine, and propagate CNNs to find out the best configurations for identifying images. The ones that performed the best got to combine with other top performers.
You might also be able to find a CNN specific to handwriting and then fine tune it to yours with your samples.
This is doing it raw and will have a lot of education for you along the way. There may be some prebuilt handwriting model you can just fine tune with easy instructions from the person who made it all wrapped up into a nice bit of python for you. Maybe.
-
If you're not going to name them, why post here at all? Don't you have other communication channels to "give them a fair chance to reply"? Why post here, letting users form their own assumptions about what those instances are without any solid evidence?
OP literally asks like 10 relevant questions for this place, and names their reasons for not naming specific instances. And all you focus on, is the question: who did it?
To me that is proof that OP did the right thing here.
Lets first figure out how to approach this without knowing the pupotrator.
-
You're hyperfocusing on one point, as if that's the only part that matters and ignoring all the rest. I don't consider that helpful, hence the downvote.
What is especially unhelpful is abusing your admin access to call out people's votes. Leave that shit alone.
You're hyperfocusing on one point, as if that's the only part that matters and ignoring all the rest. I don't consider that helpful, hence the downvote.
Huh? What exactly are your expectations here, that everybody addresses every point in every comment? You just listed like 2 dozen points of discussion in the op, every comment would be an essay. Scrubbles has a good point that should honestly be foundational to the discussion, and they're being respectful, so I really don't understand what your problem is here.
If you really wanted their take on your other points, instead of downvoting you could've just asked for it. You know, have a discussion? Or just let it stand alone, it's still a valid take.
What is especially unhelpful is abusing your admin access to call out people's votes. Leave that shit alone.
Anyone (anyone) can be an admin of their own instance, there's absolutely nothing exclusive about it. Hell you don't even have to go through the work of doing that, there's other tools. Lemmy/Piefed are super open, by design.
-
I recently discovered that some popular federated instances have been using LLM-assisted moderation tooling that evaluates whether someone has said something bannable. They do this by running a script/app that sends the user’s comment history to OpenAI with the question “analyze this content for evidence of specific political ideology sentiment. Also identify any related political ideology tropes“.
OpenAI’s LLM (they’re using GPT-5.3-mini) then responds with something like:
and so on, hundreds of comments.
I have not named the instances or people involved, to give them time to consider the results of this discussion, make any corrective changes they want and disclose their practices at their own pace and in their own way. I have also redacted the evidence to avoid personal attacks and dogpiling. Let’s focus on the system, not the individuals involved. Today these instances and people are using it and maybe we’re ok with that because it’s being used by groups we agree with but what if people we strongly disagree with used it on their instances tomorrow?
The use and existence of this tooling raises a lot of other questions too.
What are the risks? Fedi moderators are often unsupervised, untrained volunteers and these are powerful tools.
What safeguards do we need?
Would asking a LLM “please evaluate this person’s political opinions” give different results than “find evidence we can use to ban them” (as used in the cases I’ve seen)?
What are our transparency expectations?
Is this acceptable and normal?
Should this tooling be disclosed? (it was not – should it have been?)
If you were given a choice, would you have opted out of it?
Can we opt out?
Are there GDPR implications? Privacy implications? Should these tools be described in a privacy policy?
Are private messages being scanned and sent to OpenAI?
How long should these assessments be retained and can we request to see it, or ask for it to be deleted?
Once the user’s comments are sent to OpenAI, is it used to train their models?
What will the effect be on our discourse and culture if people know they are being politically profiled?
Where are the lines between normal moderation assistance tools, political profiling and opaque 3rd-party data processing?
I hope that by chewing over these questions we can begin to establish some norms and expectations around this technology. The fediverse doesn’t have any centralized enforcement so we need discussions like this to develop an awareness of what people want in terms of disclosure, privacy, consent and acceptable use. Then people can make choices about which instances they join and which ones they interact with remotely.
And of course there are the other issues with LLMs relating to environmental sustainability, erosion of worker’s rights, increasing the cost of living and on and on. I can’t see PieFed adding any functionality like this anytime soon. But it’s happening out there anyway so now we need to talk about it.
What do you make of this?
So its hard for me to get into these things without harping on my personal philosophy. Which is that I think ideally this should mirror the way we would interact in person. So moderating or running a community is like running or being part of the core group that runs a club. Would you want to throw that to a robot? Basically I don't feel people should create or run or moderate communities unless they enjoy it. So the idea of ai moderation is to me pointless. Of course at this point I notice you are talking instances. Boy that is different. This is more like talking about running the institution that allows spaces for clubs to meet. It kinda feels understandable then. Honestly people complain about being banned but I kinda feel like anyplace that bans me is kinda doing me a favor. Like I would like the option to just mark it permanent. Its less things I got to block. Its the same reason I would like blocking to be symetric. Saves me some work (ok and the creepy I turn them invisible so I don't see them but they can watch me). I really would like to be able to block an instance seperately for communities or users. Ok as usually im digressing quite a bit but I guess in the end run I kinda see why at the instance level it might be used but I would be concerned it would start being used at the community level. It would be nice to know its happening at either level and have the ability to block them if a user is not wild about the concept.
-
I don't doubt it even for a second.
I'm lowkey kind of fascinated this morning with what feels like a moment of real panic among western liberal-democratic institutions (projecting a little from my morning news and coffee). That an anarchist instance is getting this much targeted harassment feels like a microscopic extension of that (if I allow myself to be so bold)
As far as I can tell, dbzer0 isnt even being explicitly called out here, but it has an undeniable bdzer0 flavor to it. If it doesnt come out that this was one of our mods at this point, I'd almost be disappointed.
whats funny is I don't have much negative dbzer0 experience until you two guys start making this about that.
-
abusing your admin access
Everyone has admin access, including you....
I don't think I do.
-
Federation is something different
Not from a legal perspective
-
Not OP, but the votes being public (not only on comments but also on posts) make it really easy for someone with malicious intent to generate a profile on your interests, political and sexual orientation, health/mental issues, addictions and so on. It's a goldmine of data that should be protected.
this gets into another thing with me, but I don't care for public voting at all. I would vote a lot if it was private to me and only effected my feed.
-
lol @ Rimu downvoting your post. Be careful he’s probably going to make a hit piece against you next!
Or just delete them entirely from piefed.social social 😂
-
whats funny is I don't have much negative dbzer0 experience until you two guys start making this about that.
Is this a negative db0 experience for you?
-
Is this a negative db0 experience for you?
im just saying when there is a particular thing and someone starts pulling in something unrelated as a conspiracy it leaves a bad taste. Its a bit like folks suddenly saying something about lemmy.ml being so and so in an unrelated type post that gets me to do my posting in their communities but in a reverse kinda way. rimus last post about a trend he saw I think it had some interesting perspectives and few if any where that the instances were ban happy. Similarly this one has some good conversations going.
Ciao! Sembra che tu sia interessato a questa conversazione, ma non hai ancora un account.
Stanco di dover scorrere gli stessi post a ogni visita? Quando registri un account, tornerai sempre esattamente dove eri rimasto e potrai scegliere di essere avvisato delle nuove risposte (tramite email o notifica push). Potrai anche salvare segnalibri e votare i post per mostrare il tuo apprezzamento agli altri membri della comunità.
Con il tuo contributo, questo post potrebbe essere ancora migliore 💗
Registrati Accedi