Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression
arXiv:2609.26177v1 Announce Type: cross Abstract: Structured pruning of attention heads provides a hardware-friendly way to compress Transformer language models. However, existing methods for measuring head-level importance require calibration data, gradient computation, or Hessian estimation…