CVE-2026-105753: vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core
vLLM’s default multimodal cache (mm_processor_cache_type="lru") mirrors state across two processes: the frontend (P0) holds only metadata (MultiModalProcessorSenderCache) while the engine core (P1) holds the real payload (MultiModalReceiverCache). The design invariant is that get_and_update() runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer “is this cached in P1?” without talking to P1.
That invariant breaks when a request is rejected after P0 has rendered and hashed the multimodal input (populating the P0 cache) but before P1 receives the item — for example, an oversized chat prompt rejected on max_model_len after rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends None instead of the payload, and P1 — which has nothing cached — trips assert mm_item is not None, f"Expected a cached item for {mm_hash=}".
This is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration.
References
- github.com/advisories/GHSA-ph3r-5jfg-f84f
- github.com/vllm-project/vllm/commit/396204230423b7cc6798300926b8fa30190d26a9
- github.com/vllm-project/vllm/pull/46747
- github.com/vllm-project/vllm/pull/51897
- github.com/vllm-project/vllm/releases/tag/v0.28.0
- github.com/vllm-project/vllm/security/advisories/GHSA-ph3r-5jfg-f84f
- nvd.nist.gov/vuln/detail/CVE-2026-105753
Code Behaviors & Features
Detect and mitigate CVE-2026-105753 with GitLab Dependency Scanning
Secure your software supply chain by verifying that all open source dependencies used in your projects contain no disclosed vulnerabilities. Learn more about Dependency Scanning →