Vllm-project

Vllm

94 Schwachstellen gefunden.

Hinweis: Diese Liste kann unvollständig sein. Daten werden ohne Gewähr im Ursprungsformat bereitgestellt.
Exploit
  • EPSS 0.3%
  • Veröffentlicht 06.10.2026 14:45:15
  • Zuletzt bearbeitet 06.10.2026 18:16:50

A security flaw has been discovered in vllm-project vLLM up to 0.31.0. This impacts the function get_token_bin_counts_and_mask of the file vllm/model_executor/layers/utils.py of the component Penalty Handler. Performing a manipulation results in deni...

Exploit
  • EPSS 0.3%
  • Veröffentlicht 06.10.2026 05:45:11
  • Zuletzt bearbeitet 06.10.2026 15:04:52

A security vulnerability has been detected in vllm-project vLLM up to 0.31.0. This impacts the function conv_ssm_forward of the file vllm/model_executor/layers/mamba/mamba_mixer2.py of the component Completions Request Handler. The manipulation leads...

  • EPSS 0.3%
  • Veröffentlicht 05.10.2026 23:01:54
  • Zuletzt bearbeitet 06.10.2026 15:17:15

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level media_io_kwargs field to select the GLMGA video backend and supply large values for the fps and max_frames options without a strict...

  • EPSS 0.33%
  • Veröffentlicht 05.10.2026 22:58:01
  • Zuletzt bearbeitet 06.10.2026 18:16:46

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the Rust frontend's track_http_metrics middleware records the raw HTTP method token as a Prometheus label for requests reaching registered routes. An unauthenticated ...

  • EPSS 0.31%
  • Veröffentlicht 05.10.2026 22:54:53
  • Zuletzt bearbeitet 08.10.2026 03:16:35

vLLM is an inference and serving engine for large language models. From 0.24.0 until 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for the media_io_kwargs.video.max_frames and media_io_kwargs.video.fps fi...

  • EPSS 0.31%
  • Veröffentlicht 05.10.2026 22:52:05
  • Zuletzt bearbeitet 08.10.2026 01:28:00

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a gramm...

Medienbericht
  • EPSS 0.31%
  • Veröffentlicht 05.10.2026 22:49:59
  • Zuletzt bearbeitet 08.10.2026 01:28:23

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, OpenAI-compatible request models accept a non-empty cache_salt value without enforcing the character and length restrictions required by the IPCCacheServerKey consume...

  • EPSS 0.2%
  • Veröffentlicht 05.10.2026 22:47:54
  • Zuletzt bearbeitet 08.10.2026 01:28:50

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker's query_key value from the caller-controlled X-Request-Id header. A concurrent ...

  • EPSS 0.27%
  • Veröffentlicht 05.10.2026 22:46:03
  • Zuletzt bearbeitet 08.10.2026 03:16:35

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors in the features.kwargs_data field, cache identifiers in the fe...

  • EPSS 0.43%
  • Veröffentlicht 05.10.2026 22:37:19
  • Zuletzt bearbeitet 08.10.2026 01:29:33

vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engi...