vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core
vLLM's default multimodal cache (mm_processor_cache_type="lru") mirrors state across two processes: the frontend (P0) holds only metadata (MultiModalProcessorSenderCache) while the engine core (P1) holds the real payload (MultiModalReceiverCache). The design invariant is that get_and_update() runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1. That invariant breaks when a request is rejected after …