Open Issues Need Help
View All on GitHub DGX Spark (SM121) Current Support Audit 1 day ago
good first issue wip op: gemm op: misc arch: sm12x op: linear attention
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
good first issue op: linear attention
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
good first issue needs-triage model: qwen3.5 / 3.6 / 3.8 op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
flashinfer.nvfp4_quantize: make CuTe-DSL backend's `a_global_sf` arg a host side arg (to align CuTe-DSL backend with CUDA backend) about 2 months ago
good first issue op: misc
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
Next fix for gemm-allreduce two-shot 2 months ago
good first issue cute-dsl refactor op: gemm op: comm
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
Potentially superfluous check that disables non gated activations in the cutlass fused moe API 5 months ago
good first issue priority: should have (P1) op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
MoE autotune print a lot failed kernel on SM120 6 months ago
bug good first issue priority: should have (P1) op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
help wanted priority: must have (P0) needs-triage op: moe-routing
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
BF16 hidden_states for trtllm_fp4_block_scale_moe 6 months ago
good first issue op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
Autotuning failsafe fallback to top1 tactic 8 months ago
feature request good first issue op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
Inaccurate API Docstrings for Attention Prefill 10 months ago
documentation good first issue
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch