We have a Toolforge tool, Scholia. we currently seeing a lot of activity that looks like an incompent bot activity. We have specified a robots.txt https://scholia.toolforge.org/robots.txt switch specify that bots can only index very few pages and not crawl the site. However, we see in the uwsgi.log file request that I interpret as a bot that does not honor robots.txt. Du to the Toolforge proxy, it is difficult for us to throttle or ban the bot.
We also see a lot of problems with reaching the SPARQL endpoint from Toolforge. We have some server-side WDQS request and they often (or always?) fail at the moment. The bot crawls the pages with server-side WDQS requests, so there might be a connection.
We are currently moving some queries from SPARQL to API, but whether this will help on the more fundamental problem, I do not know.
Do Toolforge/Wikimedia people have suggestions for what we can do? One approach would be to move Scholia behind a login.