Page MenuHomePhabricator

Enable FP8 KV Cache for Qwen36-27b
Closed, ResolvedPublic

Event Timeline

Change #1309662 had a related patch set uploaded (by Bartosz Wójtowicz; author: Bartosz Wójtowicz):

[machinelearning/liftwing/inference-services@main] qwen36: Allow configuring vLLM KV cache dtype

https://gerrit.wikimedia.org/r/1309662

Change #1309662 merged by jenkins-bot:

[machinelearning/liftwing/inference-services@main] qwen36: Allow configuring vLLM KV cache dtype

https://gerrit.wikimedia.org/r/1309662

Change #1310091 had a related patch set uploaded (by Bartosz Wójtowicz; author: Bartosz Wójtowicz):

[operations/deployment-charts@master] ml-services: bump llm-qwen36-27b image and set fp8 KV cache dtype

https://gerrit.wikimedia.org/r/1310091

Change #1310091 merged by jenkins-bot:

[operations/deployment-charts@master] ml-services: bump llm-qwen36-27b image and set fp8 KV cache dtype

https://gerrit.wikimedia.org/r/1310091