Chin-Yun Yu, György Fazekas
Centre for Digital Music, Queen Mary University of London
Knowledge-driven neural vocoders struggle to learn reliable fundamental frequency end-to-end, because spectral objectives provide weak supervision of periodic structure and lack phase information. We address this with a source-filter model whose alias-free additive source makes the instantaneous phase of the glottal cycle explicit; differentiating it yields the instantaneous frequency, and thus $F_0$, without an external tracker. Waveform error supervises only the deterministic harmonic path, while a spectral loss covers the full signal. On M4Singer and LM-SSD, the reconstruction is phase-aligned, reaching a signal-to-reconstruction-error ratio of 8.1 dB, while neural baselines remain negative. However, GOLF, given an external $F_0$, still reaches lower spectral distortion. On LM-SSD, the recovered $F_0$ attains the highest overall accuracy of any method tested, including supervised neural pitch trackers applied off the shelf, and the glottal closure instants come within 0.53 points of REAPER's identification rate, without any $F_0$ label.
Below are some selected audio samples from the M4Singer dataset rendered by the proposed method (GOLF-inv), the baseline given an external $F_0$ (GOLF), and the supervised DDSP+$F_0$. The original recordings (Ground Truth) are also included for reference. All samples are normalised to have the same peak amplitude. A mel-spectrogram of each clip is shown above its audio player.