ETRI Knowledge Sharing Platform : Neural Spectral Band Generation for Audio Coding

Titles

논문 검색
Type		SCI
Year	~	Keyword

List

Conference Paper Neural Spectral Band Generation for Audio Coding

Cited 0 time in scopus

Authors: Woongjib Choi, Byeong Hyeon Kim, Hyungseob Lim, Inseon Jang, Hong-Goo Kang

Citation: International Speech Communication Association (INTERSPEECH) 2025, pp.624-628

Abstract: Spectral band replication (SBR) enables bit-efficient coding by generating high-frequency bands from the low-frequency ones. However, it only utilizes coarse spectral features upon a subband-wise signal replication, limiting adaptability to diverse acoustic signals. In this paper, we explore the efficacy of a deep neural network (DNN)-based generative approach for coding the high-frequency bands, which we call neural spectral band generation (n-SBG). Specifically, we propose a DNN-based encoder-decoder structure to extract and quantize the side information related to the high-frequency components and generate the components given both the side information and the decoded core-band signals. The whole coding pipeline is optimized with generative adversarial criteria to enable the generation of perceptually plausible sound. From experiments using AAC as the core codec, we show that the proposed method achieves a better perceptual quality than HE-AAC-v1 with much less side information.

KSP Keywords: Acoustic signal, Audio coding, Deep neural network(DNN), Efficient coding, Encoder and Decoder, High frequency(HF), Low frequency, Perceptual Quality, Spectral Band Replication(SBR), frequency band, high-frequency components

218 Gajeong-ro, Yuseong-gu, Daejeon, 34129, KOREA, Contact: sh.kim@etri.re.kr

Please refrain from automatic collection of e-mail addresses posted on this homepage.