LMCache 0.3.9 Through 0.5.6 Exposes CVSS 9.8 Unpatched RCE Vulnerability; vLLM Clusters Face Unauthenticated Intrusion

JFrog security researchers disclosed CVE-2026-105192, a CVSS 9.8 unauthenticated remote code execution flaw affecting LMCache 0.3.9 through 0.5.6, with no

On October 7, 2026, the JFrog security research team disclosed CVE-2026-105192, a CVSS 9.8 unauthenticated remote code execution vulnerability in LMCache versions 0.3.9 through 0.5.6.

Factual Account

According to the vulnerability details published by JFrog, in multiprocess mode LMCache by default exposes a ZeroMQ ROUTER socket on port 5555 and does not enable any CURVE, ZAP, password, or message authentication mechanism. The socket is intended only for communication between similar LMCache processes, but when an operator binds a routable address via the --host parameter, the port becomes externally accessible. The official container image runs the process as the root user, so if exploited, an attacker can obtain the highest privileges.

The root cause lies in the message handling path: during msgpack decoding, extension code 1 is registered, corresponding to the DeviceIPCWrapper.Deserialize method, which directly calls pickle.loads. Decoding occurs during processing of the REGISTER_KV_CACHE parameter, so an attacker only needs to send one crafted DEALER message to trigger deserialization execution. Affected versions include the latest PyPI release 0.5.5, the 0.5.6 release candidate up to rc3, and the development branch as of October 7, 2026; no fixed version is available.

Mechanism Breakdown

Multiprocess mode is the default way LMCache implements cross-node KV cache sharing. If --host 0.0.0.0 is specified when http_server starts, the ZMQ socket listens on all interfaces. The message format is msgpack, and extension type 1 maps directly to pickle deserialization logic; that logic executes before parameter validation, so a single message is enough to complete exploitation. In the PoC demonstration, the attacker constructs an EvilPayload object and writes a file via os.system; the server then only logs a type error and can no longer stop code execution.

On a single host with the default binding to localhost, the port is not externally reachable; however, multi-node production deployments must expose a routable address, at which point the vulnerability is directly exposed. The configuration recommended in the official Kubernetes guide falls precisely into this high-risk scenario.

Industry Impact

As a general-purpose acceleration component for the vLLM production inference layer, LMCache’s vulnerability directly affects inference clusters that use multiprocess mode. Any deployment relying on the component for KV cache sharing faces remote code execution risk as long as the port can be accessed from external or untrusted networks. Official containers run as root, further amplifying the potential damage.

A public PoC has now been released, and the barrier to exploitation is low. If vLLM users have not changed the default configuration or restricted network access, production environments may be compromised. In the short term, limiting the port to listen only on localhost or a trusted cluster network, together with firewall policies, is the only mitigation that can be implemented immediately.

Strategic Assessment

(The following is analysis, not fact.) Looking at historical cases of similar pickle deserialization vulnerabilities, AI infrastructure components that by default trust internal communication and ignore authentication can easily become an attack surface after large-scale deployment. The LMCache vulnerability exposes the disconnect between the inference acceleration layer and security boundaries. Future components of this kind should prioritize a zero-trust model in their design; otherwise, multi-node scaling will continue to carry high risk. Enterprises should quickly assess whether their deployments use affected versions and monitor upstream fix progress.