experts4bit-qlora 0.37.4
Train and serve Mixture-of-Experts models that do not fit in VRAM: fused 4-bit experts, QLoRA, CPU/NVMe offload, and fast inference on consumer NVIDIA GPUs.
Sources
- T1experts4bit-qlora 0.37.4PyPI / crates.io / RubyGems / Go index / NuGet