Self-hosted LLM inference server with automatic GPU elasticity and on-demand model loading — vLLM + Ray Serve + LiteLLM

0 stars 0 forks 0 watchers Python Apache License 2.0
1 Open Issue Need Help Last updated: Jul 29, 2026

Open Issues Need Help

View All on GitHub
bug good first issue priority:P1

Self-hosted LLM inference server with automatic GPU elasticity and on-demand model loading — vLLM + Ray Serve + LiteLLM

Python