ETRI Knowledge Sharing Platform : Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator

Titles

논문 검색
Type		SCI
Year	~	Keyword

List

Conference Paper Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator

Cited 2 time in scopus

Citation: International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2022, pp.961-965

Abstract: In this paper, we propose a discriminator design scheme for generative adversarial network-based audio signal generation. Unlike conventional discriminators that take an entire signal as input, our discriminator separates the audio signal into harmonic and percussive components and analyzes each component independently. The rationale behind this idea is that conventional discriminators cannot reliably capture subtle distortions in audio signals, which have complicated time-frequency characteristics. By considering the time-frequency resolution of audio signals, our proposed method encourages the generator to better reconstruct harmonic and percussive features, both of which are critical for the quality of the generated signals. Listening tests show that our framework significantly enhances the stability of pitches and generates clearer piano samples compared to a baseline.

KSP Keywords: Audio signal, Audio synthesis, Design Scheme, Network-based, Signal generation, Time-frequency characteristics, generative adversarial network, listening tests, time frequency(T-F), time-frequency resolution

218 Gajeong-ro, Yuseong-gu, Daejeon, 34129, KOREA, Contact: sh.kim@etri.re.kr

Please refrain from automatic collection of e-mail addresses posted on this homepage.