“A Deep Generative Diffusion-GAN Framework for Disentangled Pitch–Timbre Modeling and Harmonic Spectral Energy Alignment in Unpaired Audio Signals”
DOI:
https://doi.org/10.62643/Abstract
The objective of musical timbre transfer is to convert the instrumental attributes of an audio signal and retain the underlying musical features like pitch, melody, rhythm and expressive dynamics. Current adversarial models, such as CycleGAN models, have reported encouraging performance on unpaired spectrogram translation but can be susceptible to pitchdrift, harmful harmonic smearing, and spectral over-smoothing using real musical signals. Recent diffusion-based generative models have already shown a high quality of perception in the synthesis of audio but are generally not provided with structures that differentiate music content and perceived timbre. To resolve the above drawbacks, a Multi-Instrument PitchDisentangled Diffusion-GAN (MPD-DG) architecture is suggested in this paper as a solution to the high-fidelity unpaired musical timbre transfer. The proposed architecture uses two encoders to explicitly decouple pitch-related musical content and instrument-specific timbre features and selectively transform timbre without preserving the melody structure. The new loss called Harmonic Distribution Matching (HDM) loss is added to preserve the F0- consistent harmonic energy and stability of overtone structures. Also, a diffusion refinement module, which is lightweight, further improves the spectral realism and minimizes artifacts at relatively low cost. The presented framework can be used to address the problem of many-tomany translation between various instruments within a single architecture, bypassing pairwise CycleGAN mappings. Experimental assessments of multi-instrument models have shown that they are much better than baseline models in terms of Mel-Cepstral Distortion (MCD), Log-Spectral Distance (LSD), and Pitch RMSE, as well as, exhibit higher perceptual quality scores in MUSHRA listening tests. The findings suggest that the suggested MPD-DG model is a promising and practical solution to high-quality musical timbre transfer, and it can be utilized in AI-assisted music composition, digital orchestration, and creative audio production.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













