Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing
arXiv:2609.29169v1 Announce Type: cross Abstract: We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation…