Open Issues Need Help
View All on GitHub vLLM 0.28 compatibility 4 days ago
help wanted feature requested
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
kvcached support matrix 22 days ago
help wanted
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
start mutiple models about 1 month ago
good first issue
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
Prefix caching makes gemma-4 generations diverge from vanilla vLLM (output stays coherent, not garbled) about 1 month ago
bug help wanted
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
Roadmap/Feature requested/TODOs---Start Here for New Contributors about 2 months ago
help wanted
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
enhancement help wanted
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
Error on gpt-oss with vLLM 10 months ago
good first issue
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
Failed to patch kv_cache_coordinator 11 months ago
good first issue
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
good first issue
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm
Question About kvcached Ability to Dynamically Recognize and Utilize Kubernetes Elastic Scaled GPU Memory Resources about 1 year ago
good first issue
ovg-project/kvcached
1.1K
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Python
#elastic-kvcache#gpu-mutiplexing#gpu-sharing#inference-engine#kvcache#kvcache-optimization#kvcached#llm#llm-framework#llm-inference#llm-serving#ollama#online-offline-coserve#serverless#sglang#vllm