turbo-attn 0.59.0
Optimized CUDAgraph-enabled kernels and attention backend for vLLM, SGLang and more based on TurboQuant near-lossless KV cache compression.
Sources
- T1turbo-attn 0.59.0PyPI / crates.io / RubyGems / Go index / NuGet
Optimized CUDAgraph-enabled kernels and attention backend for vLLM, SGLang and more based on TurboQuant near-lossless KV cache compression.