ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Conference Paper Architectural Enhancement for Safety of Vision-Language Mode
Cited - time in scopus Share share facebook twitter linkedin kakaostory
Authors
Youngwan Lee, Kangsan Kim, Kwanyong Park, Ilchae Jung, Soojin Jang, Seanie Lee, Yong-Ju Lee, Sung Ju Hwang
Issue Date
2026-07
Citation
Annual Meeting of the Association for Computational Linguistics (ACL) 2026 Workshop : Advances in Language and Vision Research (ALVR) 2026, pp.1-5
Publisher
ACL
Language
English
Type
Conference Paper
Abstract
Despite emerging efforts to enhance the safety of Vision-Language Models (VLMs), prior methods rely primarily on data-centric tuning, with limited architectural enhancements to intrinsically strengthen safety. To bridge this gap, we propose a novel modular framework for enhancing VLM safety with a Visual Guard Module (VGM), designed to assess the harmfulness of input images. This module endows VLMs with dual functionality: they not only learn to generate safer responses but can also provide an interpretable classification of harmfulness to justify their refusal decisions. A significant advantage of this approach is its modularity; the VGM is designed as a plug-and-play component, allowing for seamless integration with diverse pre-trained VLMs across various scales. Extensive experiments demonstrate that our SafeLLaVA outperforms state-of-the-art data-centric methods across multiple VLM safety benchmarks. Crucially, our architectural approach consistently outperforms both data-centric baselines and standalone guard models while strictly preserving conversational helpfulness, providing a robust and integrated solution for multimodal safety.
KSP Keywords
Architectural enhancement, Data-centric, Integrated Solution, Language Models, Plug-and-Play, interpretable classification, seamless integration, state-of-The-Art