Page MenuHomePhabricator

qwen36: Support thinking mode on the OpenAI-compatible endpoints
Open, Needs TriagePublic

Description

Summary

Thinking mode currently only works on the KServe :predict endpoint via the reasoning flag. The OpenAI chat completions endpoint always disables it, and when enabled the <think> trace comes back inline in the text instead of a separate field.

Technical notes

apply_chat_template() in the model server hardcodes enable_thinking=False — we could support vLLM's chat_template_kwargs to toggle it per request. The engine already sets reasoning_parser="qwen3" but that only applies in vLLM's own OpenAI serving layer, so we need to split the <think> block ourselves when building the response.

We should also expose vLLM's per-request thinking_token_budget (see docs), promised to the PSI team in T433470#12202348.

Acceptance criteria

  • Thinking mode can be toggled per request on the OpenAI chat completions endpoint
  • The thinking trace is returned as a separate field (reasoning) instead of inline text
  • thinking_token_budget can be set per request

Event Timeline

Change #1324337 had a related patch set uploaded (by Kevin Bazira; author: Kevin Bazira):

[machinelearning/liftwing/inference-services@main] qwen36: support thinking_token_budget on :predict endpoint

https://gerrit.wikimedia.org/r/1324337

Change #1324337 merged by jenkins-bot:

[machinelearning/liftwing/inference-services@main] qwen36: support thinking_token_budget on :predict endpoint

https://gerrit.wikimedia.org/r/1324337

Change #1324680 had a related patch set uploaded (by Kevin Bazira; author: Kevin Bazira):

[machinelearning/liftwing/inference-services@main] qwen36: support per-request thinking on chat completions

https://gerrit.wikimedia.org/r/1324680

Change #1324680 merged by jenkins-bot:

[machinelearning/liftwing/inference-services@main] qwen36: support per-request thinking on chat completions

https://gerrit.wikimedia.org/r/1324680

Change #1324755 had a related patch set uploaded (by Kevin Bazira; author: Kevin Bazira):

[machinelearning/liftwing/inference-services@main] qwen36: return the thinking trace separately from the answer

https://gerrit.wikimedia.org/r/1324755

Change #1324755 merged by jenkins-bot:

[machinelearning/liftwing/inference-services@main] qwen36: return the thinking trace separately from the answer

https://gerrit.wikimedia.org/r/1324755

Change #1325449 had a related patch set uploaded (by Kevin Bazira; author: Kevin Bazira):

[operations/deployment-charts@master] ml-services: deploy qwen36-27b and qwen3-14b isvcs with support for thinking features

https://gerrit.wikimedia.org/r/1325449

Change #1325449 merged by jenkins-bot:

[operations/deployment-charts@master] ml-services: deploy qwen36-27b and qwen3-14b isvcs with support for thinking features

https://gerrit.wikimedia.org/r/1325449

We can now turn on thinking for each request on the chat completions endpoint by sending "chat_template_kwargs":{"enable_thinking":true}. The sampling values follow the mode. Thinking mode uses temperature 1.0, top_p 0.95 and presence_penalty 0.0. The other mode uses 0.7, 0.8 and 1.5.

qwen36-27b chat completions endpoint: request with the enable_thinking flag
$ curl -s https://inference.svc.eqiad.wmnet:30443/openai/v1/chat/completions -X POST \
  -H "Host: llm-qwen36-27b.llm.wikimedia.org" \
  -H "Content-Type: application/json" \
  -d '{"model":"llm-qwen36-27b",
       "messages":[{"role":"user","content":"How many prime numbers are between 100 and 200?"}],
       "max_tokens":2000,
       "chat_template_kwargs":{"enable_thinking":true}}'

{
    "id": "9135ad65badde6fe",
    "object": "chat.completion",
    "created": 1786622923,
    "model": "llm-qwen36-27b",
    "choices": [
        {
            "index": 0,
            "message": {
                "role": "assistant",
                "content": null,
                "refusal": null,
                "annotations": null,
                "audio": null,
                "function_call": null,
                "tool_calls": [],
                "reasoning": "Here's a thinking process:\n\n1.  **Understand the User's Question:** The user asks for the count of prime numbers between 100 and 200. I need to clarify if \"between\" is inclusive or exclusive. Usually, \"between X and Y\" in math contexts can be ambiguous, but typically it means strictly between (100 < p < 200). However, since 100 and 200 are not prime, it doesn't matter for the count. I'll assume strictly between 100 and 200, which is equivalent to primes in the range [101, 199].\n\n2.  **Define Prime Numbers:** A prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself.\n\n3.  **Identify the Range:** 101 to 199 (inclusive).\n\n4.  **List/Find Primes in Range:** I'll systematically check each odd number from 101 to 199 for primality. (Even numbers > 2 are not prime).\n    I can use trial division up to sqrt(199) \u2248 14.1, so I only need to check divisibility by primes up to 13: 2, 3, 5, 7, 11, 13.\n\n    Let's go through the numbers:\n    - 101: Check 2,3,5,7,11,13. Not div by 2,3 (sum=2),5,7 (7*14=98, 7*15=105),11 (11*9=99, 11*10=110),13 (13*7=91, 13*8=104). \u2192 Prime\n    - 103: Not div by 2,3 (sum=4),5,7 (7*14=98, 7*15=105),11 (11*9=99, 11*10=110),13 (13*8=104). \u2192 Prime\n    - 105: Div by 5 \u2192 Not prime\n    - 107: Not div by 2,3 (sum=8),5,7 (7*15=105),11 (11*10=110),13 (13*8=104, 13*9=117). \u2192 Prime\n    - 109: Not div by 2,3 (sum=10),5,7 (7*15=105),11 (11*10=110),13 (13*8=104, 13*9=117). \u2192 Prime\n    - 111: Sum=3 \u2192 Div by 3 \u2192 Not prime\n    - 113: Not div by 2,3 (sum=5),5,7 (7*16=112),11 (11*10=110),13 (13*8=104, 13*9=117). \u2192 Prime\n    - 115: Div by 5 \u2192 Not prime\n    - 117: Sum=9 \u2192 Div by 3 \u2192 Not prime\n    - 119: 7*17=119 \u2192 Not prime\n    - 121: 11*11=121 \u2192 Not prime\n    - 123: Sum=6 \u2192 Div by 3 \u2192 Not prime\n    - 125: Div by 5 \u2192 Not prime\n    - 127: Not div by 2,3 (sum=10),5,7 (7*18=126),11 (11*11=121),13 (13*9=117, 13*10=130). \u2192 Prime\n    - 129: Sum=12 \u2192 Div by 3 \u2192 Not prime\n    - 131: Not div by 2,3 (sum=5),5,7 (7*18=126, 7*19=133),11 (11*12=132),13 (13*10=130). \u2192 Prime\n    - 133: 7*19=133 \u2192 Not prime\n    - 135: Div by 5 \u2192 Not prime\n    - 137: Not div by 2,3 (sum=11),5,7 (7*19=133),11 (11*12=132),13 (13*10=130). \u2192 Prime\n    - 139: Not div by 2,3 (sum=13),5,7 (7*19=133),11 (11*13=143),13 (13*10=130, 13*11=143). \u2192 Prime\n    - 141: Sum=6 \u2192 Div by 3 \u2192 Not prime\n    - 143: 11*13=143 \u2192 Not prime\n    - 145: Div by 5 \u2192 Not prime\n    - 147: Sum=12 \u2192 Div by 3 \u2192 Not prime\n    - 149: Not div by 2,3 (sum=14),5,7 (7*21=147),11 (11*13=143),13 (13*11=143, 13*12=156). \u2192 Prime\n    - 151: Not div by 2,3 (sum=7),5,7 (7*21=147, 7*22=154),11 (11*14=154),13 (13*11=143, 13*12=156). \u2192 Prime\n    - 153: Sum=9 \u2192 Div by 3 \u2192 Not prime\n    - 155: Div by 5 \u2192 Not prime\n    - 157: Not div by 2,3 (sum=13),5,7 (7*22=154),11 (11*14=154),13 (13*12=156). \u2192 Prime\n    - 159: Sum=15 \u2192 Div by 3 \u2192 Not prime\n    - 161: 7*23=161 \u2192 Not prime\n    - 163: Not div by 2,3 (sum=10),5,7 (7*23=161),11 (11*15=165),13 (13*12=156, 13*13=169). \u2192 Prime\n    - 165: Div by 5 \u2192 Not prime\n    - 167: Not div by 2,3 (sum=14),5,7 (7*24=168),11 (11*15=165),13 (13*13=169). \u2192 Prime\n    - 169: 13*13=169 \u2192 Not prime\n    - 171: Sum=9 \u2192 Div by 3 \u2192 Not prime\n    - 173: Not div by 2,3 (sum=11),5,7 (7*25=175),11 (11*16=176),13 (13*13=169). \u2192 Prime\n    - 175: Div by 5 \u2192 Not prime\n    - 177: Sum=15 \u2192 Div by 3 \u2192 Not prime\n    - 179: Not div by 2,3 (sum=17),5,7 (7*25=175),11 (11*16=176),13 (13*13=169, 13*14=182). \u2192 Prime\n    - 181: Not div by 2,3 (sum=10),5,7 (7*26=182),11 (11*17=187),13 (13*14=182). \u2192 Prime\n    - 183: Sum=12 \u2192 Div by 3 \u2192 Not prime\n    - 185: Div by 5 \u2192 Not prime\n   "
            },
            "logprobs": null,
            "finish_reason": "length",
            "stop_reason": null,
            "token_ids": null,
            "routed_experts": null
        }
    ],
    "service_tier": null,
    "system_fingerprint": null,
    "usage": {
        "prompt_tokens": 26,
        "total_tokens": 2026,
        "completion_tokens": 2000,
        "prompt_tokens_details": null
    },
    "prompt_logprobs": null,
    "prompt_token_ids": null,
    "prompt_text": null,
    "kv_transfer_params": null
}
qwen36-27b chat completions endpoint: request without the enable_thinking flag
$ curl -s https://inference.svc.eqiad.wmnet:30443/openai/v1/chat/completions -X POST \
  -H "Host: llm-qwen36-27b.llm.wikimedia.org" \
  -H "Content-Type: application/json" \
  -d '{"model":"llm-qwen36-27b",
       "messages":[{"role":"user","content":"How many prime numbers are between 100 and 200?"}],
       "max_tokens":2000}'

{
    "id": "bd3023d5e608d828",
    "object": "chat.completion",
    "created": 1786623243,
    "model": "llm-qwen36-27b",
    "choices": [
        {
            "index": 0,
            "message": {
                "role": "assistant",
                "content": "To determine how many prime numbers are between 100 and 200, we can list them systematically or count them based on known mathematical data.\n\n### Step 1: Definition and Range\nA prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself. We are looking for primes $p$ such that $100 < p < 200$. Note that 100 and 200 are not prime, so the inclusion or exclusion of the endpoints does not change the count.\n\n### Step 2: Listing the Prime Numbers\nWe can identify these primes by checking odd numbers between 100 and 200 for divisibility by small primes (2, 3, 5, 7, 11, 13).\n\n**Primes in the 100s:**\n1.  101\n2.  103\n3.  107\n4.  109\n\n**Primes in the 110s:**\n5.  113\n6.  127 (Note: 111 is div by 3, 117 is div by 9, 119 is $7 \\times 17$)\n7.  131\n8.  137\n9.  139\n\n**Primes in the 140s:**\n10. 149\n\n**Primes in the 150s:**\n11. 151\n12. 157\n\n**Primes in the 160s:**\n13. 163\n14. 167\n15. 173\n16. 179\n\n**Primes in the 180s:**\n17. 181\n18. 191\n19. 193\n20. 197\n21. 199\n\n### Step 3: Verification and Count\nLet's double-check the count by grouping them:\n*   100\u2013109: 101, 103, 107, 109 (4 primes)\n*   110\u2013119: 113 (1 prime) -> *Wait, let's re-verify 110-119.*\n    *   111 ($3 \\times 37$), 113 (Prime), 115 ($5 \\times 23$), 117 ($9 \\times 13$), 119 ($7 \\times 17$). So only **113**.\n*   120\u2013129: 127 (1 prime) -> *Check 121 ($11^2$), 123 ($3 \\times 41$), 125 ($5^3$), 127 (Prime), 129 ($3 \\times 43$).* So only **127**.\n*   130\u2013139: 131, 137, 139 (3 primes)\n*   140\u2013149: 149 (1 prime) -> *Check 141 ($3 \\times 47$), 143 ($11 \\times 13$), 147 ($3 \\times 49$).* So only **149**.\n*   150\u2013159: 151, 157 (2 primes) -> *Check 153 ($9 \\times 17$), 159 ($3 \\times 53$).*\n*   160\u2013169: 163, 167 (2 primes) -> *Check 161 ($7 \\times 23$), 169 ($13^2$).*\n*   170\u2013179: 173, 179 (2 primes) -> *Check 171 ($9 \\times 19$), 177 ($3 \\times 59$).*\n*   180\u2013189: 181 (1 prime) -> *Check 183 ($3 \\times 61$), 187 ($11 \\times 17$), 189 ($9 \\times 21$).*\n*   190\u2013199: 191, 193, 197, 199 (4 primes)\n\nSumming these up:\n$4 + 1 + 1 + 3 + 1 + 2 + 2 + 2 + 1 + 4 = 21$\n\n### Conclusion\nThere are **21** prime numbers between 100 and 200.\n\nThe list is:\n101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199.",
                "refusal": null,
                "annotations": null,
                "audio": null,
                "function_call": null,
                "tool_calls": [],
                "reasoning": null
            },
            "logprobs": null,
            "finish_reason": "stop",
            "stop_reason": null,
            "token_ids": null,
            "routed_experts": null
        }
    ],
    "service_tier": null,
    "system_fingerprint": null,
    "usage": {
        "prompt_tokens": 28,
        "total_tokens": 1258,
        "completion_tokens": 1230,
        "prompt_tokens_details": null
    },
    "prompt_logprobs": null,
    "prompt_token_ids": null,
    "prompt_text": null,
    "kv_transfer_params": null
}

The two runs make different prompts. The prompt_tokens values are different (26 and 28). This shows that the flag gets to the model. The run with the flag can use all the max_tokens while the model thinks. The budget in the third comment prevents this.

This applies to the chat completions endpoint. The :predict endpoint already has this functionality through its reasoning flag as shown below:

qwen36-27b predict endpoint: request with reasoning flag
$ curl -s https://inference.svc.eqiad.wmnet:30443/v1/models/llm-qwen36-27b:predict -X POST \
  -H "Host: llm-qwen36-27b.llm.wikimedia.org" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"How many prime numbers are between 100 and 200?",
       "reasoning":true,"max_tokens":2000}'

{
    "model_name": "llm-qwen36-27b",
    "response": "Here's a thinking process:\n\n1.  **Understand the User's Question:** The user asks for the count of prime numbers between 100 and 200 (inclusive or exclusive? Usually \"between\" means exclusive, but in math contexts, it's often inclusive of the bounds if they are primes, though 100 and 200 are not primes anyway. I'll assume the range is (100, 200) or [100, 200], it doesn't matter since 100 and 200 are composite).\n\n2.  **Define Prime Numbers:** A prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself.\n\n3.  **Determine the Range:** Numbers from 101 to 199.\n\n4.  **List/Count Primes in Range:** I need to systematically identify primes between 100 and 200.\n   I'll list them out or recall/count them.\n   Let's list all primes up to 200 and subtract those up to 100, or just list the ones in the range.\n   Primes < 100: There are 25 primes.\n   Primes < 200: There are 46 primes.\n   So, primes between 100 and 200 = 46 - 25 = 21.\n   Let me verify this by listing them to be sure.\n\n   Primes between 100 and 200:\n   101, 103, 107, 109,\n   113, 127,\n   131, 137, 139,\n   149,\n   151, 157,\n   163, 167,\n   173, 179,\n   181, 191, 193, 197, 199.\n\n   Let's count them:\n   100s: 101, 103, 107, 109 (4)\n   110s: 113 (1) [111=3*37, 117=9*13, 119=7*17]\n   120s: 127 (1) [121=11^2, 123=3*41, 129=3*43]\n   130s: 131, 137, 139 (3) [133=7*19]\n   140s: 149 (1) [141=3*47, 143=11*13, 147=3*49]\n   150s: 151, 157 (2) [153=9*17, 159=3*53]\n   160s: 163, 167 (2) [161=7*23, 169=13^2]\n   170s: 173, 179 (2) [171=9*19, 177=3*59]\n   180s: 181 (1) [183=3*61, 187=11*17, 189=9*21]\n   190s: 191, 193, 197, 199 (4)\n\n   Total: 4 + 1 + 1 + 3 + 1 + 2 + 2 + 2 + 1 + 4 = 21.\n\n   Matches the known fact: \u03c0(200) = 46, \u03c0(100) = 25, so 46 - 25 = 21.\n\n5.  **Formulate Response:** State clearly that there are 21 prime numbers between 100 and 200. Optionally list them for completeness. Keep it concise.\n\n   Check: \"between 100 and 200\" - does it include 100 and 200? Neither are prime, so it doesn't affect the count. I'll specify \"strictly between\" or just give the number. I'll also list them to be helpful.\n\n   Final answer: 21.\n   I'll double-check the list quickly:\n   101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199.\n   Count: 21. Correct.\n   I'll provide the answer directly and list them. All good.\u2705\n   Output matches response.\u2705\n   Proceed. \n   Structure: Direct answer first, then optional list/verification. Keep it precise.\n   \"There are **21** prime numbers between 100 and 200.\"\n   Add list for clarity.\n   Done. \n   Self-Correction/Verification during thought: \n   Wait, let's verify \u03c0(200) and \u03c0(100) values. \n   Primes up to 100: 2,3,5,7,11,13,17,19,23,29,31,37,41,43,47,53,59,61,67,71,73,79,83,89,97. That's 25. Correct.\n   Primes up to 200: 25 + 21 = 46. Correct.\n   So 21 is definitively correct.\n   I will output the answer. \n   All steps verified. No issues. \n   Output generation. \n   [Done] \n   *Self-Correction/Refinement during generation prep*: \n   I should make sure the phrasing matches standard math conventions. \"Between 100 and 200\" typically excludes endpoints, but since neither is prime, it's fine. I'll just state 21 and list them. \n   Ready. \n   Output matches the final response.\u2705\n   [Output Generation] -> \"There are **21** prime numbers between 100 and 200. Here they are: 101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199.\"\u2705\n   All good. \n   Proceeds. \n   [Final Check of the Prompt]: \"How many prime numbers are between 100 and 200?\" -> Direct answer expected. \n   Matches. \n   Done. \n   [Output Generation] (matches the refined version)\u2705\n   *(Note: I'll keep it concise as requested by the prompt style)*\n   \"There are **21** prime numbers between 100 and 200.\" + list. Perfect. \n   [Done]\n</think>\n\nThere are **21** prime numbers between 100 and 200. \n\nHere they are in order:\n101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199.",
    "prompt_tokens": 26,
    "completion_tokens": 1916
}

The isvc now returns the thinking trace in its own reasoning field. The trace is not in the answer text. This applies to both non-streamed replies and to streamed replies. Each streamed piece goes to delta.reasoning or to delta.content.

The field name is reasoning. It is not reasoning_content, as this task first said. I have updated the acceptance criterion.

qwen36-27b chat completions endpoint: thinking trace in its own field and answer in content
$ curl -s https://inference.svc.eqiad.wmnet:30443/openai/v1/chat/completions -X POST \
  -H "Host: llm-qwen36-27b.llm.wikimedia.org" \
  -H "Content-Type: application/json" \
  -d '{"model":"llm-qwen36-27b",
       "messages":[{"role":"user","content":"How many prime numbers are between 100 and 200?"}],
       "max_tokens":3000,
       "chat_template_kwargs":{"enable_thinking":true},
       "thinking_token_budget":100}'

{
    "id": "9b170fc9bedf5d36",
    "object": "chat.completion",
    "created": 1786625437,
    "model": "llm-qwen36-27b",
    "choices": [
        {
            "index": 0,
            "message": {
                "role": "assistant",
                "content": "\n\nThere are **21** prime numbers between 100 and 200.\n\nHere is the complete list:\n101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199",
                "refusal": null,
                "annotations": null,
                "audio": null,
                "function_call": null,
                "tool_calls": [],
                "reasoning": "Here's a thinking process:\n\n1.  **Understand the User's Question:** The user asks for the count of prime numbers between 100 and 200.\n\n2.  **Define the Range:** \"Between 100 and 200\" typically means exclusive of the endpoints, but since 100 and 200 are not prime anyway, it doesn't matter. I'll consider integers from 101 to 1"
            },
            "logprobs": null,
            "finish_reason": "stop",
            "stop_reason": null,
            "token_ids": null,
            "routed_experts": null
        }
    ],
    "service_tier": null,
    "system_fingerprint": null,
    "usage": {
        "prompt_tokens": 26,
        "total_tokens": 258,
        "completion_tokens": 232,
        "prompt_tokens_details": null
    },
    "prompt_logprobs": null,
    "prompt_token_ids": null,
    "prompt_text": null,
    "kv_transfer_params": null
}

This applies to the chat completions endpoint. On :predict the trace stays inline in the text. That is the usual format of that payload as shown below:

qwen36-27b predict endpoint: thinking trace stays inline
$ curl -s https://inference.svc.eqiad.wmnet:30443/v1/models/llm-qwen36-27b:predict -X POST \
  -H "Host: llm-qwen36-27b.llm.wikimedia.org" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"How many prime numbers are between 100 and 200?",
       "reasoning":true,"max_tokens":3000}'

{
    "model_name": "llm-qwen36-27b",
    "response": "Here's a thinking process:\n\n1.  **Understand the User's Question:** The user is asking for the count of prime numbers between 100 and 200. I need to be careful about the boundaries: \"between 100 and 200\" usually means exclusive (101 to 199), but sometimes inclusive. Since 100 and 200 are not prime, it doesn't matter in this case. I'll count primes from 101 to 199 inclusive.\n\n2.  **Define Prime Numbers:** A prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself.\n\n3.  **Identify the Range:** 101 to 199.\n\n4.  **Method to Find Primes in Range:**\n   - I can list them out systematically or use a known list.\n   - I know there are 25 primes up to 100.\n   - I can recall or calculate the primes between 100 and 200.\n   - Let's list them carefully:\n     101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199.\n   - Wait, let me verify this list. I might have missed some.\n   - Let's go through numbers 100-200 systematically:\n     100: composite\n     101: prime\n     102: even\n     103: prime\n     104: even\n     105: div by 5\n     106: even\n     107: prime\n     108: even\n     109: prime\n     110: even\n     111: div by 3 (1+1+1=3)\n     112: even\n     113: prime\n     114: even\n     115: div by 5\n     116: even\n     117: div by 3 (1+1+7=9)\n     118: even\n     119: 7*17 = 119 (composite)\n     120: even\n     121: 11*11 (composite)\n     122: even\n     123: div by 3\n     124: even\n     125: div by 5\n     126: even\n     127: prime\n     128: even\n     129: div by 3\n     130: even\n     131: prime\n     132: even\n     133: 7*19 = 133 (composite)\n     134: even\n     135: div by 5\n     136: even\n     137: prime\n     138: even\n     139: prime\n     140: even\n     141: div by 3\n     142: even\n     143: 11*13 = 143 (composite)\n     144: even\n     145: div by 5\n     146: even\n     147: div by 3\n     148: even\n     149: prime\n     150: even\n     151: prime\n     152: even\n     153: div by 3\n     154: even\n     155: div by 5\n     156: even\n     157: prime\n     158: even\n     159: div by 3\n     160: even\n     161: 7*23 = 161 (composite)\n     162: even\n     163: prime\n     164: even\n     165: div by 5\n     166: even\n     167: prime\n     168: even\n     169: 13*13 = 169 (composite)\n     170: even\n     171: div by 3\n     172: even\n     173: prime\n     174: even\n     175: div by 5\n     176: even\n     177: div by 3\n     178: even\n     179: prime\n     180: even\n     181: prime\n     182: even\n     183: div by 3\n     184: even\n     185: div by 5\n     186: even\n     187: 11*17 = 187 (composite)\n     188: even\n     189: div by 3\n     190: even\n     191: prime\n     192: even\n     193: prime\n     194: even\n     195: div by 5\n     196: even\n     197: prime\n     198: even\n     199: prime\n     200: composite\n\n   - Let's count the primes I found:\n     101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199.\n     Count: 21.\n\n5.  **Cross-Verify with Known Data:**\n   - I recall that the number of primes up to 100 is 25.\n   - The number of primes up to 200 is 46.\n   - So primes between 100 and 200 (exclusive of 100, inclusive of 200? 200 isn't prime) = 46 - 25 = 21.\n   - This matches my count.\n   - Let's double-check the list to be absolutely sure:\n     1. 101\n     2. 103\n     3. 107\n     4. 109\n     5. 113\n     6. 127\n     7. 131\n     8. 137\n     9. 139\n     10. 149\n     11. 151\n     12. 157\n     13. 163\n     14. 167\n     15. 173\n     16. 179\n     17. 181\n     18. 191\n     19. 193\n     20. 197\n     21. 199\n   - All correct.\n\n6.  **Formulate Response:**\n   - State the answer clearly: 21 prime numbers.\n   - Optionally list them for completeness.\n   - Mention that 100 and 200 are not prime, so \"between\" doesn't change the count.\n   - Keep it concise.\n\n   Draft: There are **21** prime numbers between 100 and 200. They are: 101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, and 199.\n\n7.  **Final Check:** \n   - Question: \"How many prime numbers are between 100 and 200?\"\n   - Answer: 21.\n   - Verification: \u03c0(200) = 46, \u03c0(100) = 25. 46 - 25 = 21. Correct.\n   - List matches known primes.\n   - Response is accurate and direct.\u2705\n   - Ready. \n   - Output matches response.\u2705\n   - Proceed. \n   - [Self-Correction/Refinement during thought]: I should make sure I didn't miss any tricky composites like 121, 143, 169, 187. I already checked them. All good.\n   - The phrasing \"between 100 and 200\" typically excludes endpoints, but since neither is prime, it's unambiguous.\n   - Final answer: 21. Provide list optionally. All set. \n   - Output generation. \n   - [Done.] \n   - *Self-Correction/Verification during process*: \n     Let's quickly recount the list to be 100% sure: \n     101, 103, 107, 109, 113 (5)\n     127, 131, 137, 139, 149 (5) -> total 10\n     151, 157, 163, 167, 173 (5) -> total 15\n     179, 181, 191, 193, 197, 199 (6) -> total 21. \n     Correct. \n     All steps verified. Output matches.\u2705\n   - Final Response Generation. \n   - \"There are **21** prime numbers between 100 and 200. Here they are listed for reference: 101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, and 199.\" (Matches draft) \n   - Proceeds. \n   - [Output Generation] -> *Proceeds* \u2705\n</think>\n\nThere are **21** prime numbers between 100 and 200. \n\nFor reference, here they are listed in order:\n101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, and 199.",
    "prompt_tokens": 26,
    "completion_tokens": 2737
}

The response text contains the thinking trace inline, ending with </think>.

We can now set thinking_token_budget for each request on the two endpoints. The budget limits only the thinking. At the budget, the model stops the thought and then writes its answer (finish_reason stop). max_tokens is different: it stops the full reply (finish_reason length). If thinking is off, the budget has no effect. The value must be a whole number of 1 or more. If not, the server rejects the request.

For the :predict endpoint, the thought stops in the middle of a sentence and a full answer follows. The reply used 230 of the 3000 permitted tokens:

qwen36-27b predict endpoint: with a budget of 100
$ curl -s https://inference.svc.eqiad.wmnet:30443/v1/models/llm-qwen36-27b:predict -X POST \
  -H "Host: llm-qwen36-27b.llm.wikimedia.org" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"How many prime numbers are between 100 and 200?",
       "reasoning":true,"max_tokens":3000,"thinking_token_budget":100}'

{
    "model_name": "llm-qwen36-27b",
    "response": "Here's a thinking process:\n\n1.  **Understand the User's Question**: The user asks for the count of prime numbers between 100 and 200. I need to determine if \"between\" is inclusive or exclusive. Typically, \"between X and Y\" in math contexts can be ambiguous, but I'll assume it means strictly between (101 to 199) or inclusive (100 to 200). Since 1</think>\n\nThere are 21 prime numbers between 100 and 200.\nHere's the list:\n101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199",
    "prompt_tokens": 26,
    "completion_tokens": 230
}

For the chat completions endpoint, the thought stops at the budget (about 100 tokens) and a full answer follows, with finish_reason stop:

qwen36-27b chat completions endpoint: with a budget of 100
$ curl -s https://inference.svc.eqiad.wmnet:30443/openai/v1/chat/completions -X POST \
  -H "Host: llm-qwen36-27b.llm.wikimedia.org" \
  -H "Content-Type: application/json" \
  -d '{"model":"llm-qwen36-27b",
       "messages":[{"role":"user","content":"How many prime numbers are between 100 and 200?"}],
       "max_tokens":3000,
       "chat_template_kwargs":{"enable_thinking":true},
       "thinking_token_budget":100}'

{
    "id": "9ecc3ee6f4f8b7f6",
    "object": "chat.completion",
    "created": 1786627350,
    "model": "llm-qwen36-27b",
    "choices": [
        {
            "index": 0,
            "message": {
                "role": "assistant",
                "content": "\n\nThere are **21** prime numbers between 100 and 200.\n\nHere they are listed:\n101, 103, 107, 109, 113, 127, 131, 137, 139, 149, 151, 157, 163, 167, 173, 179, 181, 191, 193, 197, 199.\n\n*(Note: This count includes numbers strictly greater than 100 and strictly less than 200. Both 100 and 200 are composite, so the count remains 21 regardless of whether the endpoints are included or excluded.)*",
                "refusal": null,
                "annotations": null,
                "audio": null,
                "function_call": null,
                "tool_calls": [],
                "reasoning": "Here's a thinking process:\n\n1.  **Understand the User's Question**: The user asks for the count of prime numbers between 100 and 200.\n\n2.  **Define Prime Numbers**: A prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself.\n\n3.  **Identify the Range**: The range is \"between 100 and 200\". Typically, this"
            },
            "logprobs": null,
            "finish_reason": "stop",
            "stop_reason": null,
            "token_ids": null,
            "routed_experts": null
        }
    ],
    "service_tier": null,
    "system_fingerprint": null,
    "usage": {
        "prompt_tokens": 26,
        "total_tokens": 313,
        "completion_tokens": 287,
        "prompt_tokens_details": null
    },
    "prompt_logprobs": null,
    "prompt_token_ids": null,
    "prompt_text": null,
    "kv_transfer_params": null
}

Keep max_tokens as the general limit. The budget limits only the thinking.