TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

arXiv:2609.26100v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model inference through collaboration between a lightweight draft model and a target verifier. Existing methods mainly improve the draft side, while the target model is typically kept dense and…

aiscience

Sources