DIO: non-invasive LLM gateway that learns per-backend latency online (dual-timescale NLMS) and routes vLLM/Ollama/SGLang/TGI with SLO-aware admission.

3 stars 0 forks 3 watchers Python Apache License 2.0
admission-control ai-infrastructure inference-server llm llm-gateway llm-inference llm-serving llmops load-balancing mcp-server mlops model-context-protocol model-serving ollama openai-api openai-compatible python self-hosted sglang vllm
7 Open Issues Need Help Last updated: Sep 14, 2026

Open Issues Need Help

View All on GitHub

DIO: non-invasive LLM gateway that learns per-backend latency online (dual-timescale NLMS) and routes vLLM/Ollama/SGLang/TGI with SLO-aware admission.

Python
#admission-control#ai-infrastructure#inference-server#llm#llm-gateway#llm-inference#llm-serving#llmops#load-balancing#mcp-server#mlops#model-context-protocol#model-serving#ollama#openai-api#openai-compatible#python#self-hosted#sglang#vllm

DIO: non-invasive LLM gateway that learns per-backend latency online (dual-timescale NLMS) and routes vLLM/Ollama/SGLang/TGI with SLO-aware admission.

Python
#admission-control#ai-infrastructure#inference-server#llm#llm-gateway#llm-inference#llm-serving#llmops#load-balancing#mcp-server#mlops#model-context-protocol#model-serving#ollama#openai-api#openai-compatible#python#self-hosted#sglang#vllm
documentation good first issue

DIO: non-invasive LLM gateway that learns per-backend latency online (dual-timescale NLMS) and routes vLLM/Ollama/SGLang/TGI with SLO-aware admission.

Python
#admission-control#ai-infrastructure#inference-server#llm#llm-gateway#llm-inference#llm-serving#llmops#load-balancing#mcp-server#mlops#model-context-protocol#model-serving#ollama#openai-api#openai-compatible#python#self-hosted#sglang#vllm
enhancement good first issue

DIO: non-invasive LLM gateway that learns per-backend latency online (dual-timescale NLMS) and routes vLLM/Ollama/SGLang/TGI with SLO-aware admission.

Python
#admission-control#ai-infrastructure#inference-server#llm#llm-gateway#llm-inference#llm-serving#llmops#load-balancing#mcp-server#mlops#model-context-protocol#model-serving#ollama#openai-api#openai-compatible#python#self-hosted#sglang#vllm
enhancement good first issue

DIO: non-invasive LLM gateway that learns per-backend latency online (dual-timescale NLMS) and routes vLLM/Ollama/SGLang/TGI with SLO-aware admission.

Python
#admission-control#ai-infrastructure#inference-server#llm#llm-gateway#llm-inference#llm-serving#llmops#load-balancing#mcp-server#mlops#model-context-protocol#model-serving#ollama#openai-api#openai-compatible#python#self-hosted#sglang#vllm
enhancement good first issue

DIO: non-invasive LLM gateway that learns per-backend latency online (dual-timescale NLMS) and routes vLLM/Ollama/SGLang/TGI with SLO-aware admission.

Python
#admission-control#ai-infrastructure#inference-server#llm#llm-gateway#llm-inference#llm-serving#llmops#load-balancing#mcp-server#mlops#model-context-protocol#model-serving#ollama#openai-api#openai-compatible#python#self-hosted#sglang#vllm
enhancement good first issue

DIO: non-invasive LLM gateway that learns per-backend latency online (dual-timescale NLMS) and routes vLLM/Ollama/SGLang/TGI with SLO-aware admission.

Python
#admission-control#ai-infrastructure#inference-server#llm#llm-gateway#llm-inference#llm-serving#llmops#load-balancing#mcp-server#mlops#model-context-protocol#model-serving#ollama#openai-api#openai-compatible#python#self-hosted#sglang#vllm