ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Conference Paper How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?
Cited - time in scopus Share share facebook twitter linkedin kakaostory
Authors
Arda Senocak, Sooyoung Park, Tae-Hyun Oh, Joon Chung
Issue Date
2026-06
Citation
Conference on Computer Vision and Pattern Recognition (CVPR) 2026, pp.22931-22940
Publisher
Computer Vision Foundation
Language
English
Type
Conference Paper
Abstract
We present the first scalable framework for training sound source localization (SSL) models using synthetic data from text-to-X models. Although SSL has made notable progress, existing models remain constrained by limited-scale, un- curated real-world datasets that often suffer from seman- tic misalignment. Furthermore, the introduction of new SSL tasks and benchmarks has increased the need for more generalizable models. To address these challenges, we leverage synthetic data to create synthetic clones of the VGGSound dataset, enabling both fully synthetic and hy- brid real–synthetic training. We demonstrate that synthetic data can effectively replace, refine, and scale real train- ing datasets. Extensive experiments across multiple bench- marks show that synthetic data not only matches real data in performance but also enables significant improvements when combined with real samples. Our findings provide the first systematic evidence that synthetic data can serve as a scalable and effective approach for advancing SSL mod- els. Code and data are available at: https://github. com/swimmiing/SyntheticSSL.
KSP Keywords
Audio-visual, Real data, Real samples, Real-world, Scalable framework, Synthetic data, need for, sound source localization