Well Reddit would have a lot less scraper traffic if it hadn't shut down the feed API 3 years ago in an attempt to keep its data closed and only distributed to Google (which paid money for it). Self-inflicted. No sympathy.
Well Reddit would have a lot less scraper traffic if it hadn't shut down the feed API 3 years ago in an attempt to keep its data closed and only distributed to Google (which paid money for it). Self-inflicted. No sympathy.
I doubt it. The scraper botnets aren't going to bother with niceties like APIs, they just come in over HTTPS and grab anything they can.
There was a whole mini-industry around scraping specifically Reddit using the API feed because it was so open. AI companies first trained on Pushshift, which is like Common Crawl for Reddit, which is why Reddit shut down Pushshift (with legal threats IIRC) at the same time.
If there were a well-known URL that retrieved all content added to a site after a given date, and a meaningful number of sites implemented it, I bet the biggest scrapers would use it.
Except, of course, that then sites would start lying and sending abridged, poisoned, or nonexistent data.
Nice things cannot be had, no matter how many levels you go down.