Vllm

Vllm

85 Schwachstellen gefunden.

Hinweis: Diese Liste kann unvollständig sein. Daten werden ohne Gewähr im Ursprungsformat bereitgestellt.
  • EPSS 0.31%
  • Veröffentlicht 05.10.2026 22:54:53
  • Zuletzt bearbeitet 08.10.2026 03:16:35

vLLM is an inference and serving engine for large language models. From 0.24.0 until 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for the media_io_kwargs.video.max_frames and media_io_kwargs.video.fps fi...

  • EPSS 0.31%
  • Veröffentlicht 05.10.2026 22:52:05
  • Zuletzt bearbeitet 08.10.2026 01:28:00

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a gramm...

Medienbericht
  • EPSS 0.31%
  • Veröffentlicht 05.10.2026 22:49:59
  • Zuletzt bearbeitet 08.10.2026 01:28:23

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, OpenAI-compatible request models accept a non-empty cache_salt value without enforcing the character and length restrictions required by the IPCCacheServerKey consume...

  • EPSS 0.2%
  • Veröffentlicht 05.10.2026 22:47:54
  • Zuletzt bearbeitet 08.10.2026 01:28:50

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker's query_key value from the caller-controlled X-Request-Id header. A concurrent ...

  • EPSS 0.27%
  • Veröffentlicht 05.10.2026 22:46:03
  • Zuletzt bearbeitet 08.10.2026 03:16:35

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors in the features.kwargs_data field, cache identifiers in the fe...

  • EPSS 0.43%
  • Veröffentlicht 05.10.2026 22:37:19
  • Zuletzt bearbeitet 08.10.2026 01:29:33

vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engi...

  • EPSS 0.21%
  • Veröffentlicht 05.10.2026 22:32:59
  • Zuletzt bearbeitet 08.10.2026 01:29:51

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cache_salt value, placing the con...

Exploit
  • EPSS 0.66%
  • Veröffentlicht 30.09.2026 16:45:13
  • Zuletzt bearbeitet 06.10.2026 17:14:37

A flaw has been found in vllm-project vLLM up to 0.26.0. This vulnerability affects unknown code of the file rust/src/parser/src/unified/gemma4.rs of the component Gemma4UnifiedParser. Executing a manipulation can lead to denial of service. The attac...

  • EPSS 0.31%
  • Veröffentlicht 26.09.2026 13:23:23
  • Zuletzt bearbeitet 07.10.2026 17:33:53

vLLM before 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints but validates only that the values are integers, not that each token id is within the model vocabulary/log...

Exploit
  • EPSS 0.33%
  • Veröffentlicht 26.09.2026 13:23:21
  • Zuletzt bearbeitet 06.10.2026 19:22:01

vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit requests with min_tokens great...