> And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds.
> And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds.
There's a difference between someone running a scraping tool occasionally and bots constantly and rapidly re-scraping the same site over and over again
The author's website is responsible for storing its own data. AI services currently treat the entire web as their storage and cache layer.
This is an important context: the site is more likely to be targeted by scrapers because it is a curated collection of scraped information.
Do the bots care? Seemingly very little intelligence in many of them. Could be as simple as the site has more pages, so more traffic.
Loot first, ask questions later.
A curated collection of people who give away money. Was ever sweeter honey ever found in a pot?
Similar to the dose making the poison - the thing that jumped out at me in this blog was the ratio of scraping to visits. Unless OP is scraping thousands of times a day I don't really think they're in the same class as the bots they are blocking.
OP here: I'm not scraping thousands of times per day! Usually just a few times per year.
But you could just be one of thousands of bots targeting the same sources you’re scraping.
Who scrapes the scrapemen?
Live by the scraper, die by the scraper