High [CVE-2026-56340] Denial of service and potential arbitrary code execution via malformed multimodal embedding requests
This high-severity Red Hat Linux advisory covers CVE-2026-56340 affecting Red Hat AI Inference Server.
Android app · Google Play
Monitor future Red Hat Linux CVEs from your phone.
Choose a whole vendor or a precise platform, then receive matching security advisories by phone notification, email, or both. Coverage follows 32 official vendor sources and 160+ reviewed platform categories.
Summary
vLLM versions >= 0.10.2 and < 0.13.0 are missing sparse tensor validation in multimodal embeddings processing.
Because PyTorch disables sparse tensor invariant checks by default, an attacker can submit crafted embedding requests with malformed (negative or out-of-bounds) tensor indices, when the prompt-embeds feature is enabled, to trigger crashes or resource exhaustion (denial of service), with potential for out-of-bounds/write-what-where memory corruption.
This continues CVE-2025-62164, whose prior fix only disabled the feature by default rather than addressing the root cause. A flaw was found in vLLM.
Red Hat rates this issue as having Important impact for affected Red Hat AI Inference Server images shipping vLLM 0.10.2 through 0.13.x when prompt-embeds multimodal embedding support is enabled. Versions outside this range, Red Hat OpenShift AI KServe sidecars, and Red Hat Enterprise Linux AI 3.4 bootc images (vLLM 0.17+/0.18+) are not affected.
Red Hat severity: Important — CVSS 8.8 (CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H). Weakness: CWE-787.
Red Hat lists Red Hat AI Inference Server; Red Hat Enterprise Linux AI (RHEL AI) 3; Red Hat OpenShift AI (RHOAI) as not affected. Will not fix / out of support: Red Hat AI Inference Server.
Red Hat does not currently list a fixing RHSA for this CVE.
Affected versions
No affected-version range was extracted from the source record. The vendor advisory is authoritative — check it before change work.
Official advisory · high-confidence parse· fetched 16 days ago·verify at source
Fixed versions
No fixed release is recorded yet. That does not prove no patch exists — confirm against the vendor advisory.
Official advisory · high-confidence parse· fetched 16 days ago·verify at source
Mitigation checklist
- Disable prompt-embeds if not required. Restrict who can submit multimodal embedding requests. Apply authentication and rate limits on inference APIs. Upgrade to a fixed vLLM build when available from Red Hat.
Official advisory · high-confidence parse· fetched 16 days ago·verify at source
Discussion(0)
No comments yet. Share field notes, upgrade gotchas, or questions — verify against the vendor advisory before acting on community advice.
Sign in to join the discussion.