Eprints2archives does behave somewhat like a web crawler, so maybe it should pay attention to robots.txt. Here are some resources to start looking at what needs to be done: https://developers.google.com/search/reference/robots_txt https://tools.ietf.org/html/draft-koster-rep-00 https://yoast.com/ultimate-guide-robots-txt/ https://en.wikipedia.org/wiki/Robots_exclusion_standard https://opensource.googleblog.com/2019/07/googles-robotstxt-parser-is-now-open.html https://github.com/google/robotstxt https://www.scrapehero.com/how-to-prevent-getting-blacklisted-while-scraping/ https://medium.com/ub-women-data-scholars/let-the-robot-do-your-work-web-scraping-with-python-9c147fb7690f https://www.promptcloud.com/blog/how-to-read-and-respect-robots-file/ https://opensource.googleblog.com/2019/07/googles-robotstxt-parser-is-now-open.html
Eprints2archives does behave somewhat like a web crawler, so maybe it should pay attention to robots.txt.
Here are some resources to start looking at what needs to be done:
https://developers.google.com/search/reference/robots_txt
https://tools.ietf.org/html/draft-koster-rep-00
https://yoast.com/ultimate-guide-robots-txt/
https://en.wikipedia.org/wiki/Robots_exclusion_standard
https://opensource.googleblog.com/2019/07/googles-robotstxt-parser-is-now-open.html
https://github.com/google/robotstxt
https://www.scrapehero.com/how-to-prevent-getting-blacklisted-while-scraping/
https://medium.com/ub-women-data-scholars/let-the-robot-do-your-work-web-scraping-with-python-9c147fb7690f
https://www.promptcloud.com/blog/how-to-read-and-respect-robots-file/
https://opensource.googleblog.com/2019/07/googles-robotstxt-parser-is-now-open.html