Open Issues Need Help
View All on GitHub bug good first issue priority:P1
Self-hosted LLM inference server with automatic GPU elasticity and on-demand model loading — vLLM + Ray Serve + LiteLLM
Python
Self-hosted LLM inference server with automatic GPU elasticity and on-demand model loading — vLLM + Ray Serve + LiteLLM
Self-hosted LLM inference server with automatic GPU elasticity and on-demand model loading — vLLM + Ray Serve + LiteLLM