turbo-attn 0.57.1
Optimized CUDAgraph-enabled kernels and attention backend for vLLM, SGLang and more based on TurboQuant near-lossless KV cache compression.
Sources
- T1turbo-attn 0.57.1PyPI / crates.io / RubyGems / Go index / NuGet
Optimized CUDAgraph-enabled kernels and attention backend for vLLM, SGLang and more based on TurboQuant near-lossless KV cache compression.