So I think it would be useful if we generated something like 1000-2500 stop words using TF-ISF per wiki, then run a cross-TF-IDF on all wikis and take the top 500 to 1000 words common in all wikis.
I would expect these to include interwiki language ISO codes and words like http, com, net etc from urls.