ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Journal Article GraDual: Resolving the Semantic–Frequency Trade-Off via Gradient-Regulated Dual-Branch Multimodal Manipulation Detection
Cited 0 time in scopus Download 21 time Share share facebook twitter linkedin kakaostory
Authors
Donghun Lee, Jang-Ho Choi, Seongho Kim, Sungwon Yi, Nojun Kwak
Issue Date
2026-06
Citation
IEEE Access, v.14, pp.95360-95377
ISSN
2169-3536
Publisher
IEEE
Language
English
Type
Journal Article
DOI
https://dx.doi.org/10.1109/ACCESS.2026.3706107
Abstract
Integrating frequency-domain features into multimodal manipulation detection introduces a trade-off between localization accuracy and semantic understanding. While these features improve localization accuracy by highlighting artifacts, they degrade text grounding performance through gradient interference in unified encoders. Existing methods exhibit substantial performance degradation (15.9–17.8%) performance degradation exceeding 15–18% as manipulation complexity increases. We address this challenge through three key contributions. First, we introduce DGM4-Complex, a comprehensive benchmark of 27k image-text pairs that incorporates auxiliary manipulations to create realistic multi-level scenarios, addressing the critical limitation of existing datasets that restrict samples to single manipulations per modality. Second, we provide a systematic analysis of the semantic–frequency trade-off, revealing how frequency-domain features enhance localization accuracy while compromising text grounding performance due to gradient interference in shared encoders. Third, we present GraDual (Gradient-Regulated Dual-branch), a gradient-controlled framework that decouples semantic and frequency processing through differentiated gradient scaling, mitigating the trade-off and enabling robust performance across detection and grounding tasks. Our method achieves 84.78% AUC on DGM4-Complex with only a 9.1% performance decline from single-level manipulations, which is substantially lower than the 15.9-17.8% drop observed in existing methods, while maintaining stable text grounding performance (0.4% decline compared to 6.4–12.9% in others).
Keyword
Deepfake, deepfake detection, generative model, multimodal
KSP Keywords
Dual-branch, Frequency domain(FD), Grounding performance, Image-text, Multi-level, Robust performance, Systematic analysis, Trade-off, generative model, localization accuracy, manipulation detection
This work is distributed under the term of Creative Commons License (CCL)
(CC BY)
CC BY