Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact

2 stars 0 forks 2 watchers Cuda Apache License 2.0
bf16 compression cuda gpu inference llm llm-inference lossless lossless-compression nvidia pytorch transformers weight-compression
6 Open Issues Need Help Last updated: Sep 19, 2026

Open Issues Need Help

View All on GitHub

Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact

Cuda
#bf16#compression#cuda#gpu#inference#llm#llm-inference#lossless#lossless-compression#nvidia#pytorch#transformers#weight-compression

Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact

Cuda
#bf16#compression#cuda#gpu#inference#llm#llm-inference#lossless#lossless-compression#nvidia#pytorch#transformers#weight-compression
help wanted benchmark

Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact

Cuda
#bf16#compression#cuda#gpu#inference#llm#llm-inference#lossless#lossless-compression#nvidia#pytorch#transformers#weight-compression

Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact

Cuda
#bf16#compression#cuda#gpu#inference#llm#llm-inference#lossless#lossless-compression#nvidia#pytorch#transformers#weight-compression
help wanted roadmap

Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact

Cuda
#bf16#compression#cuda#gpu#inference#llm#llm-inference#lossless#lossless-compression#nvidia#pytorch#transformers#weight-compression
help wanted roadmap benchmark

Neural Weight Compression: lossless BF16 weights (-31 %) with a fused CUDA matvec that decodes in registers, faster than cuBLAS on RTX 4070 and NVIDIA A16, bit-exact

Cuda
#bf16#compression#cuda#gpu#inference#llm#llm-inference#lossless#lossless-compression#nvidia#pytorch#transformers#weight-compression