Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

arXiv:2609.26486v1 Announce Type: cross Abstract: Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works…

science

Sources