-
-
Notifications
You must be signed in to change notification settings - Fork 22.7k
All issues
Issue creation is restricted in this repository
- #48168 · simon-mo opened
on Jul 9, 2026 4 - #50001 · ywang96 opened
on Jul 27, 2026 14 - #57448 · WoosukKwon opened
on Sep 17, 2026 5
Issues
is:issue state:open
is:issue state:open
Search results
- Status: Open.#58807 In vllm-project/vllm;
[Performance][Bug]: Tiered Offloading
bugSomething isn't workingSomething isn't workingStatus: Open.#58804 In vllm-project/vllm;- Status: Open.#58799 In vllm-project/vllm;
- Status: Open.#58789 In vllm-project/vllm;
- Status: Open.#58776 In vllm-project/vllm;
- Status: Open.#58774 In vllm-project/vllm;
- Status: Open.#58751 In vllm-project/vllm;
[Bug]: MFU/MBU silently drops attention and FFN for GPTQ, AWQ, gpt-oss, DeepSeek-V4 FP8 and online-quantized models
deepseekRelated to DeepSeek modelsRelated to DeepSeek modelsgpt-ossRelated to GPT-OSS modelsRelated to GPT-OSS modelsStatus: Open.#58742 In vllm-project/vllm;[Bug]: DeepSeek-V4-Flash + DFlash speculator fails at startup (fp8_ds_mla draft KV dtype; 16384 vs 4096 hidden states)
bugSomething isn't workingSomething isn't workingdeepseekRelated to DeepSeek modelsRelated to DeepSeek modelsStatus: Open.#58733 In vllm-project/vllm;[Bug]: GLM MXFP4 returns 0 GSM8K on GB200 (DEP8+EP8, vLLM 0.30.0)
bugSomething isn't workingSomething isn't workingStatus: Open.#58729 In vllm-project/vllm;- Status: Open.#58728 In vllm-project/vllm;
[Bug][Frontend]: Anthropic /v1/messages always hoists inline system messages because merge detection checks the CLI --chat-template (None) instead of the model's template
bugSomething isn't workingSomething isn't workingStatus: Open.#58727 In vllm-project/vllm;