Data Engineering Bug Report or Data Problem Form.
What kind of problem are you reporting?
- Access related problem
- Service related problem
- Data related problem
For a service related problem:
- What is the nature of the issue?
We recently increased the size of the presto cluster from 5 to 15 worker nodes.
However, the result was a reduction in stability of the cluster, with many failed queries and timeouts. We investigated this as an incident in ticket: T325331: Superset: Presto backend: Unable to access some charts
We tried various combinations of new vs old hosts, but it did not seem to be related to which hosts were in use.
- 5 hosts works ok
- 10 hosts demonstrates significant instability
- 15 hosts is even worse
We also tried increasing the amout of JVM heap available to the coordinator process, but it didn't help.
- What are the steps to reproduce the issue?
Enable and start presto-server processes on those servers where it is currently disabled: an-presto10[06-15]
- What happens?
Timeouts and instability with presto in superset and from the command-line.
- What should happen instead?
The cluster should perform faster and more reliably with 15 nodes instead of 5.
For the DE Team to fill out
Which systems does this effect?
- Hive
- Druid
- Superset
- Turnilo
- WikiDumps
- Wikistats
- Airflow
- HDFS
- Goblin
- Scqoop
- Dashiki
- DataHub
- Spark
- Jupyter
- Modern Event Platform
- Event Logging
- Other: Presto
Impact Assessment:
Does this problem qualify as an incident?
- Yes
- No
Does this violate an SLO?
- Yes
- No
| Value Calculator | Rank |
|---|---|
| Will this improve the efficiency of a teams workflow? | 1-3 |
| Does this have an effect of our Core Metrics? | 1-3 |
| Does this align with our strategic goals? | 1-3 |
| Is this a blocker for another team? | 1-3 |