ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Conference Paper Query-Driven Scene Selection for Personalized Video Summarization Using Multimodal Semantic Descriptions
Cited - time in scopus Share share facebook twitter linkedin kakaostory
Authors
Hyun-Jeong Yim, Sungjun Ahn, Jung Sun Um, Jae Hyun Seo
Issue Date
2026-07
Citation
International Symposium on Broadband Multimedia Systems and Broadcasting (BMSB) 2026, pp.1-4
Publisher
IEEE
Language
English
Type
Conference Paper
Abstract
This paper presents a query-driven video summarization framework based on scene-level semantic decision making using a language model. Unlike conventional video summarization approaches that rely on visual saliency, keyword matching, or learned importance scores, the proposed method interprets a user query as a high-level semantic constraint that governs scene selection. Video content is first segmented into scenes, and each scene is represented by a unified textual description constructed from visual captions and aligned speech transcripts. A language model is then employed as a semantic decision engine to determine whether each scene satisfies the semantic constraints implied by the user query. Based on the resulting binary relevance decisions, query- relevant scenes are selected and concatenated in temporal order to generate a personalized summary without modifying scene boundaries or synthesizing new content. Experimental results on real-world news and movie videos demonstrate that the proposed approach achieves high recall of query-relevant scenes while effectively excluding unrelated content, particularly for abstract and context-dependent queries. These results indicate that language-model-based semantic reasoning provides an effective and flexible mechanism for personalized video summarization.
Keyword
Query-driven video summarization, semantic reasoning, scene selection, multi-modal video understanding, personalized media services
KSP Keywords
Binary relevance, Decision-making, Flexible mechanism, High recall, Keyword matching, Language Models, Multi-modal, Personalized video, Query-driven, Real-world, Semantic constraints