We need a production service for a SPARQL Endpoint for Commons. Estimates on hardware needs are a complete guess as we don't have a baseline for the load on this service and only vague ideas about the data increase over time.
A few data points:
* current estimate is about 5M pages with structured data on Commons
* largest increase so far was +1M / month
* in comparison WDQS has ~11B triples
* current journal size on sdcquery01 (test server) is 2.5G, but that data dump is > 6 month old
* we want at least 3 servers per DC, to provide enough redundancy in case of hardware failure
* current test server runs with 8G of heap
Estimated specs (oversized to try to account for growth over the lifetime of those servers):
* 6x single Xeon (4C/8T), 64G RAM, 500G usable SSD space RAID1 / 10 software, 1G NIC
Minimal specs right now (estimate):
* 6x single Xeon (2C/4T), 32G RAM, 100G usable HDD space, 1G NIC