Open Issues Need Help
View All on GitHubNeural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact
Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact
Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact
Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact
Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact
Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact