While troubleshooting Elastic cluster latency in the parent task ( T387176 ), we've noticed a few things that could speed up diagnoses in the future:
- Add API calls to check thread pool rejections, hot threads, etc to our Wikitech docs .
- Add per-host alerts for thread pool rejections
- Update the Elasticsearch node comparison dashboard to use Prometheus instead of Graphite.
Creating this ticket to fulfill the above AC.