What:
We have observed signs of potential memory leak in the Orchestrator. We need to investigate to try and confirm if a leak indeed exists. If so, identify root causes and next steps for mitigation.
Why:
- memory usage increase over time and has ended frequently in service restarts or OOM kills in Prod logs
- performance struggles with response delays seemingly correlated with uptime in user experience