We should collect at least basic metrics on each service to prometheus and possibly alert from there. Some services might report by email, i.e. for failed job runs.
Performance metrics:
- Dagster - dagster_log_to_db populates a db table and grafana job status chart
- Metabase
- https://www.metabase.com/docs/latest/installation-and-operation/observability-with-prometheus
- Minio
- Trino (partially done...)
- https://nil1729.github.io/trino-jmx-monitoring
- https://trino.io/docs/current/admin/jmx.html
- Hive Standalone Metastore T405228: Enable prometheus-jmx-exporter for FR hive-standalone-metastore
Alerting:
- Dagster - slack notification, nagios check_procs
- Metabase - nagios check_procs
- Minio - nagios check_minio which scrapes exported prometheus metrics
- Trino - API poll using nagios check_http
- Hive Standalone Metastore