Vllm

Vllm

85 Schwachstellen gefunden.

Hinweis: Diese Liste kann unvollständig sein. Daten werden ohne Gewähr im Ursprungsformat bereitgestellt.
  • EPSS 0.45%
  • Veröffentlicht 21.09.2026 22:04:10
  • Zuletzt bearbeitet 29.09.2026 12:47:29

vLLM versions through 0.29.0 contain a denial of service vulnerability in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. Attackers can send requests with incomplete kv_transfer_params dictionary entries to trigge...

  • EPSS 0.2%
  • Veröffentlicht 19.09.2026 22:58:10
  • Zuletzt bearbeitet 28.09.2026 18:38:31

vLLM through 0.29.0 fails to properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer(). Attackers can supply out-of-bounds token indices that corrupt logits memory of concurrent re...

  • EPSS 0.24%
  • Veröffentlicht 18.09.2026 19:06:06
  • Zuletzt bearbeitet 28.09.2026 18:35:40

vLLM through 0.29.0 contains a memory corruption vulnerability in the Triton _bincount_kernel where prompt token IDs index the penalty prompt-presence bitset without bounds checking against vocabulary size. Attackers can submit multimodal audio reque...

  • EPSS 0.25%
  • Veröffentlicht 18.09.2026 19:06:06
  • Zuletzt bearbeitet 28.09.2026 18:36:00

vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead of model output logits width in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary that pass validation, causing Log...

Exploit
  • EPSS 0.38%
  • Veröffentlicht 18.09.2026 13:20:05
  • Zuletzt bearbeitet 28.09.2026 18:37:06

vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID...

  • EPSS 0.54%
  • Veröffentlicht 17.09.2026 22:15:31
  • Zuletzt bearbeitet 28.09.2026 18:37:42

vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound unti...

Exploit
  • EPSS 0.46%
  • Veröffentlicht 16.09.2026 17:49:20
  • Zuletzt bearbeitet 07.10.2026 14:16:57

vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards that choice t...

  • EPSS 0.66%
  • Veröffentlicht 16.09.2026 16:34:39
  • Zuletzt bearbeitet 07.10.2026 14:24:22

vLLM is an inference and serving engine for large language models. Prior to 0.24.0, the input_audio handling path for /v1/chat/completions calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing VLLM_MAX_AUDIO_DECODE_DURATION_S to the...

  • EPSS 0.27%
  • Veröffentlicht 12.09.2026 12:08:58
  • Zuletzt bearbeitet 16.09.2026 17:31:06

vLLM versions before 0.28.0 fail to validate audio sample rate headers in the transcription endpoint, allowing authenticated clients to bypass duration checks. Attackers can submit forged FLAC headers with inflated sample rates to trigger excessive m...

  • EPSS 0.11%
  • Veröffentlicht 12.09.2026 12:08:57
  • Zuletzt bearbeitet 24.09.2026 20:28:13

vLLM versions >=0.10.2 and <0.28.0 do not apply any audio decode-size or duration limit when extracting audio from video input for NanoNemotronVL models. In nano_nemotron_vl.py, _extract_audio_from_videos calls load_audio_pyav(BytesIO(video_bytes)) w...