ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Conference Paper Query-Adaptive Diversity for Long-Form Video Question Answering
Cited - time in scopus Download 0 time Share share facebook twitter linkedin kakaostory
Authors
Jonghee Kim, Jinyoung Moon
Issue Date
2026-09
Citation
European Conference on Computer Vision (ECCV) 2026 Workshop : Multimodal Large Language Models for Unified Comprehension and Generation (MUCG), pp.1-11
Publisher
Multimodal Large Language Models for Unified Comprehension and Generation (MUCG)
Language
English
Type
Conference Paper
Abstract
Agent-based long-form video question answering selects frames iteratively by query relevance, but across successive turns it can repeatedly retrieve the same frames. To reduce this, we accumulate the selected frames into a cross-turn memory pool and score each turn with Maximal Marginal Relevance (MMR) over it — and find that the MMR diversity penalty is relevance-blind: it measures redundancy by frame-to-frame similarity alone, ignoring how relevant each frame is to the query. For questions that require multiple frames of the same recurring event, e.g., counting or temporal ordering, it therefore discards genuine evidence and can perform worse than applying no diversity term. To address this, we propose a query-adaptive diversity penalty: the penalty on a candidate is reduced in proportion to its own query relevance, retaining query-relevant recurrences while suppressing irrelevant duplicates. We evaluate this on MLVU and LVBench under a shared Qwen3.5-9B backbone and a minimal retrieval agent: on both benchmarks the relevance-blind pool falls below plain top-K retrieval, whereas the query-adaptive penalty surpasses it, with the improvement concentrated on counting and ordering questions.
Keyword
Query-Adaptive Diversity, Maximal Marginal Relevance, Long-Form Video Understanding, Frame Selection, Video Question Answering
KSP Keywords
Adaptive penalty, Frame selection, Query Relevance, Question Answering, Temporal ordering, Top-k retrieval, Video understanding, agent-based, memory pool
This work is distributed under the term of Creative Commons License (CCL)
(CC BY)
CC BY