turbo-attn 0.57.0
Optimized CUDAgraph-enabled kernels and attention backend for vLLM, SGLang and more based on TurboQuant near-lossless KV cache compression.
Sources
- T1turbo-attn 0.57.0PyPI / crates.io / RubyGems / Go index / NuGet
- T1turbo-attn 0.57.2PyPI / crates.io / RubyGems / Go index / NuGet