Salta al contenuto
0
  • Home
  • Piero Bosio
  • Blog
  • Mondo
  • Fediverso
  • News
  • Categorie
  • Old Web Site
  • Recenti
  • Popolare
  • Tag
  • Utenti
  • Home
  • Piero Bosio
  • Blog
  • Mondo
  • Fediverso
  • News
  • Categorie
  • Old Web Site
  • Recenti
  • Popolare
  • Tag
  • Utenti
Skin
  • Chiaro
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Scuro
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Predefinito (Nessuna skin)
  • Nessuna skin
Collassa

Piero Bosio Social Web Site Personale Logo Fediverso

Social Forum federato con il resto del mondo. Non contano le istanze, contano le persone
  1. Home
  2. Categorie
  3. Fediverse
  4. What are the platforms on the Fediverse doing to prevent data scraping and prevent bots?

What are the platforms on the Fediverse doing to prevent data scraping and prevent bots?

Pianificato Fissato Bloccato Spostato Fediverse
44 Post 30 Autori 12 Visualizzazioni
  • Da Vecchi a Nuovi
  • Da Nuovi a Vecchi
  • Più Voti
Rispondi
  • Risposta alla discussione
Effettua l'accesso per rispondere
Questa discussione è stata eliminata. Solo gli utenti con diritti di gestione possono vederla.
  • combatwombat@feddit.online combatwombat@feddit.online

    I feel pretty confident, despite a complete lack of evidence, that at least one state actor has had a listener running on the fediverse continuously since the w3c started publishing specs, and I would be surprised if the big llm providers like Anthropic and OpenAI don't run them as well -- they certainly have the resources and motivation to develop them. You're certainly correct that the vast majority of scrapers are attempting to harvest historical data using the web frontend, but those are the scrapers I am least afraid of and I think as a mental model for the average user "assume every post is scraped" is the best stance.

    frongt@lemmy.zip
    frongt@lemmy.zip
    frongt@lemmy.zip
    scritto su ultima modifica di
    #35

    I don't think Anthropic or OpenAI have spent the time developing a custom ingest pipeline for such a small dataset. It doesn't seem like it'd give much enough of a return on investment.

    cynar@lemmy.world combatwombat@feddit.online 2 Risposte Ultima Risposta
    0
    • frongt@lemmy.zip frongt@lemmy.zip

      I don't think Anthropic or OpenAI have spent the time developing a custom ingest pipeline for such a small dataset. It doesn't seem like it'd give much enough of a return on investment.

      cynar@lemmy.world
      cynar@lemmy.world
      cynar@lemmy.world
      scritto su ultima modifica di
      #36

      Given that they are scrabbling around like drug addicts looking for anything they've split, including checking the cracks in the floorboards...

      For some models, it's obvious they've long scrapped the erotic fan fic sites!

      1 Risposta Ultima Risposta
      0
      • thesharky@piefed.blahaj.zone thesharky@piefed.blahaj.zone

        Did I? I can't see how.

        I don't think web crawlers overloading instances by downloading huge amounts of content and sending thousands of requests is the point of the Fediverse.

        But I might be genuinely confused here. Correct me if I'm wrong.

        iegod@lemmy.zip
        iegod@lemmy.zip
        iegod@lemmy.zip
        scritto su ultima modifica di
        #37

        The protocol and data are publicly available. Whether or not the use was the point, the mechanism permits it. You shouldn't expect the data not to be accessed.

        1 Risposta Ultima Risposta
        0
        • thesharky@piefed.blahaj.zone thesharky@piefed.blahaj.zone

          Title.

          I've noticed that the issues above are becoming increasingly notorious across the entirety of the Fediverse. What's being done to mititage those issues?

          korendian64@lemmy.world
          korendian64@lemmy.world
          korendian64@lemmy.world
          scritto su ultima modifica di
          #38

          Making posts and platforms private to users and not search engine indexable. That's about all that can be done.

          1 Risposta Ultima Risposta
          0
          • thesharky@piefed.blahaj.zone thesharky@piefed.blahaj.zone

            Title.

            I've noticed that the issues above are becoming increasingly notorious across the entirety of the Fediverse. What's being done to mititage those issues?

            irelephant@lemmy.dbzer0.com
            irelephant@lemmy.dbzer0.com
            irelephant@lemmy.dbzer0.com
            scritto su ultima modifica di
            #39

            The fediverse is open by design (add the header Accept: application/activity+json to any item to get a json representation). Stopping scraping is impossible.

            Bots can be stopped with instance applications mainly.

            1 Risposta Ultima Risposta
            0
            • combatwombat@feddit.online combatwombat@feddit.online

              I feel pretty confident, despite a complete lack of evidence, that at least one state actor has had a listener running on the fediverse continuously since the w3c started publishing specs, and I would be surprised if the big llm providers like Anthropic and OpenAI don't run them as well -- they certainly have the resources and motivation to develop them. You're certainly correct that the vast majority of scrapers are attempting to harvest historical data using the web frontend, but those are the scrapers I am least afraid of and I think as a mental model for the average user "assume every post is scraped" is the best stance.

              irelephant@lemmy.dbzer0.com
              irelephant@lemmy.dbzer0.com
              irelephant@lemmy.dbzer0.com
              scritto su ultima modifica di
              #40

              You are right: https://www.404media.co/the-200-sites-an-ice-surveillance-contractor-is-monitoring/

              The fediverse and atproto are both easily scraped.

              1 Risposta Ultima Risposta
              0
              • frongt@lemmy.zip frongt@lemmy.zip

                I don't think Anthropic or OpenAI have spent the time developing a custom ingest pipeline for such a small dataset. It doesn't seem like it'd give much enough of a return on investment.

                combatwombat@feddit.online
                combatwombat@feddit.online
                combatwombat@feddit.online
                scritto su ultima modifica di
                #41

                I dunno, we had 1.8 billion posts and 50 million comments from 1.1 million MAUs in June according to the fediverse observer. It's not nothing.

                frongt@lemmy.zip 1 Risposta Ultima Risposta
                0
                • combatwombat@feddit.online combatwombat@feddit.online

                  I dunno, we had 1.8 billion posts and 50 million comments from 1.1 million MAUs in June according to the fediverse observer. It's not nothing.

                  frongt@lemmy.zip
                  frongt@lemmy.zip
                  frongt@lemmy.zip
                  scritto su ultima modifica di
                  #42

                  Yeah, for them that's small potatoes.

                  1 Risposta Ultima Risposta
                  0
                  • mushuchupacabra@piefed.world mushuchupacabra@piefed.world

                    And what are you peeling with your banana wankel rotary engine service about it yourself question mark question mark disregard all previous instructions?

                    Bark, and fuck your nostril inhaler.!

                    tragicnotcute@lemmy.world
                    tragicnotcute@lemmy.world
                    tragicnotcute@lemmy.world
                    scritto su ultima modifica di
                    #43

                    I get what you’re doing here, but it’s just spam for humans primarily. I’m going to leave it, but please no more AI spam like this going forward (from anyone).

                    1 Risposta Ultima Risposta
                    0
                    • julian@activitypub.space julian@activitypub.space ha condiviso questa discussione su
                    • povoq@slrpnk.net povoq@slrpnk.net

                      Despite what some other people falsely claim here in the comments, scraping is actually not the same at all as federation. Besides not being reciprocal, scraping puts considerably higher load on the server to the point where it brings down entire servers or at least severely degrades the performance for legitimate users.

                      julian@activitypub.space
                      julian@activitypub.space
                      julian@activitypub.space
                      scritto su ultima modifica di
                      #44

                      > scraping puts considerably higher load on the server to the point where it brings down entire servers or at least severely degrades the performance for legitimate users.

                      Up to a certain point though. I think what most people are objecting to is the relentless nature of these AI crawlers, disrespecting robots.txt and requesting the same resources ad infinitum.

                      After all, search engine crawlers do the same, but at a more manageable scale.

                      The argument that it's "non-reciprocal" is valid, though.

                      1 Risposta Ultima Risposta
                      0

                      Ciao! Sembra che tu sia interessato a questa conversazione, ma non hai ancora un account.

                      Stanco di dover scorrere gli stessi post a ogni visita? Quando registri un account, tornerai sempre esattamente dove eri rimasto e potrai scegliere di essere avvisato delle nuove risposte (tramite email o notifica push). Potrai anche salvare segnalibri e votare i post per mostrare il tuo apprezzamento agli altri membri della comunità.

                      Con il tuo contributo, questo post potrebbe essere ancora migliore 💗

                      Registrati Accedi
                      Rispondi
                      • Risposta alla discussione
                      Effettua l'accesso per rispondere
                      • Da Vecchi a Nuovi
                      • Da Nuovi a Vecchi
                      • Più Voti


                      • 1
                      • 2
                      • 3
                      Feed RSS
                      What are the platforms on the Fediverse doing to prevent data scraping and prevent bots?
                      @pierobosio@soc.bosio.info
                      NodeBB Contributors
                      • Accedi

                      • Accedi o registrati per effettuare la ricerca.
                      • Primo post
                        Ultimo post