Vllm-project

Vllm

94 Schwachstellen gefunden.

Hinweis: Diese Liste kann unvollständig sein. Daten werden ohne Gewähr im Ursprungsformat bereitgestellt.
  • EPSS 0.21%
  • Veröffentlicht 05.10.2026 22:32:59
  • Zuletzt bearbeitet 08.10.2026 01:29:51

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cache_salt value, placing the con...

Exploit
  • EPSS 0.66%
  • Veröffentlicht 30.09.2026 16:45:13
  • Zuletzt bearbeitet 06.10.2026 17:14:37

A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of the component Gemma4UnifiedParser. Executing a manipulation can lead to denial of service. The attac...

  • EPSS 0.31%
  • Veröffentlicht 26.09.2026 13:23:23
  • Zuletzt bearbeitet 07.10.2026 17:33:53

vLLM before 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints but validates only that the values are integers, not that each token id is within the model vocabulary/log...

  • EPSS 0.28%
  • Veröffentlicht 26.09.2026 13:23:22
  • Zuletzt bearbeitet 30.09.2026 18:18:01

vLLM is an inference and serving engine for large language models. In versions from 0.22.1 through 0.28.0, the operator-supplied model revision pin (--revision / --code-revision) is not propagated to several Hugging Face artifact loads for the FunAud...

Exploit
  • EPSS 0.33%
  • Veröffentlicht 26.09.2026 13:23:21
  • Zuletzt bearbeitet 06.10.2026 19:22:01

vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit requests with min_tokens great...

Exploit
  • EPSS 0.31%
  • Veröffentlicht 26.09.2026 13:23:21
  • Zuletzt bearbeitet 06.10.2026 19:30:25

vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate. When the request contains a 'features' (multimodal) payload, vllm/entrypoints/serve/disagg/serving.py builds a multimod...

Exploit
  • EPSS 0.61%
  • Veröffentlicht 26.09.2026 13:23:20
  • Zuletzt bearbeitet 06.10.2026 19:32:45

vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and the per-modality --limit-mm-per-prompt item li...

  • EPSS 0.31%
  • Veröffentlicht 26.09.2026 13:23:19
  • Zuletzt bearbeitet 06.10.2026 19:37:35

vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select different sampler subclasses in video ...

  • EPSS 0.33%
  • Veröffentlicht 26.09.2026 13:23:18
  • Zuletzt bearbeitet 06.10.2026 19:40:20

vllm before 0.29.0 fails to enforce VLLM_MAX_AUDIO_CLIP_FILESIZE_MB limit in multimodal chat audio decoding, allowing unauthenticated clients to bypass file size restrictions. Attackers can submit oversized audio files through chat endpoints to consu...

  • EPSS 0.31%
  • Veröffentlicht 26.09.2026 13:23:18
  • Zuletzt bearbeitet 06.10.2026 19:46:06

vLLM versions before 0.29.0 contain a denial-of-service vulnerability in the cache_salt parameter accepted on OpenAI-compatible and Anthropic API endpoints, which lacks maximum length validation and is processed on the single EngineCore scheduler thr...