Open Issues Need Help
View All on GitHub good first issue qwen3 stale
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Rust
#cuda#cuda-kernels#deepseek#gpu#inference#inference-engine#kimi#kimi-k2#kv-cache#llm#llm-inference#llm-serving#model-serving#moe#openai-api#paged-attention#qwen#qwen3#rust#vllm
qwen3: green-context decode overlap creates resources on device 0 regardless of --device-ordinal 20 days ago
good first issue qwen3
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Rust
#cuda#cuda-kernels#deepseek#gpu#inference#inference-engine#kimi#kimi-k2#kv-cache#llm#llm-inference#llm-serving#model-serving#moe#openai-api#paged-attention#qwen#qwen3#rust#vllm
bench: add HTTP serving benchmark coverage for Qwen mixed sampling about 2 months ago
good first issue
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Rust
#cuda#cuda-kernels#deepseek#gpu#inference#inference-engine#kimi#kimi-k2#kv-cache#llm#llm-inference#llm-serving#model-serving#moe#openai-api#paged-attention#qwen#qwen3#rust#vllm
qwen35: RoPE cache covers 4096 positions but config admits 262144 about 2 months ago
good first issue qwen35
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Rust
#cuda#cuda-kernels#deepseek#gpu#inference#inference-engine#kimi#kimi-k2#kv-cache#llm#llm-inference#llm-serving#model-serving#moe#openai-api#paged-attention#qwen#qwen3#rust#vllm
enhancement good first issue
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Rust
#cuda#cuda-kernels#deepseek#gpu#inference#inference-engine#kimi#kimi-k2#kv-cache#llm#llm-inference#llm-serving#model-serving#moe#openai-api#paged-attention#qwen#qwen3#rust#vllm
Preserve streaming usage in the vLLM frontend about 2 months ago
bug good first issue
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Rust
#cuda#cuda-kernels#deepseek#gpu#inference#inference-engine#kimi#kimi-k2#kv-cache#llm#llm-inference#llm-serving#model-serving#moe#openai-api#paged-attention#qwen#qwen3#rust#vllm