| Status | Subtype | Assigned | Task | ||
|---|---|---|---|---|---|
| Open | None | T433975 [draft] Q1 FY2026-27 LLM Platform Milestone 2 | |||
| Open | kevinbazira | T433569 qwen36: Support thinking mode on the OpenAI-compatible endpoints | |||
| Resolved | BWojtowicz-WMF | T431089 Set ephemeral-storage requests/limits on ML predictor containers to prevent node disk eviction | |||
| Open | DPogorzelski-WMF | T433973 Update kserve helm chart to include LLMInferenceService | |||
| Resolved | kevinbazira | T433977 Support OpenAI-compatible tool calling in the qwen36 model-server | |||
| Open | BWojtowicz-WMF | T424923 Support GPU Partition specification in deployments | |||
| Open | None | T436928 ml-serve-eqiad: Repartition two MI300X nodes from 24GB to 96GB partitions | |||
| Open | kevinbazira | T434059 Enable Structured Outputs for LLMs | |||
| Open | None | T429237 monitoring: Grafana dashboard for LLM serving on MI300X | |||
| Open | None | T429236 monitoring: View GPU usage per LLM deployment/model |
[draft] Q1 FY2026-27 LLM Platform Milestone 2
[draft] Q1 FY2026-27 LLM Platform Milestone 2
Assigned To
None
Authored By
| isarantopoulos | |
| Aug 4 2026, 1:32 PM |
Project Tags
Referenced Files
None
Subscribers