High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

inference kv-cache llm vllm
19 Open Issues Need Help Last updated: Jul 16, 2026

Open Issues Need Help

View All on GitHub

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm
enhancement good first issue

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm
documentation good first issue

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm
enhancement good first issue

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

AI Summary: The issue reports that workers are initializing sequentially rather than in parallel, despite the expectation for parallel initialization across different devices. The provided logs show significant time delays (several seconds) between the `PeagflowRadixCache` registration events for various workers (TP0-TP7), indicating a bottleneck or serialization in the initialization process.

Complexity: 3/5
good first issue

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm
good first issue

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm

AI Summary: This chore issue requests adding the block count to the save logs, similar to how load logs already include both layer and block counts. The goal is to make the log outputs consistent, providing more complete information during save operations.

Complexity: 1/5
good first issue

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust
#inference#kv-cache#llm#vllm