ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Conference Paper 효율적인 체화 질의 응답을 위한 에피소드 수준 멀티모달 KV Caching
Cited - time in scopus Share share facebook twitter linkedin kakaostory
Authors
옹효빈, 장민수
Issue Date
2026-02
Citation
한국로봇학회 종합 학술 대회 2026, pp.129-131
Publisher
한국로봇학회
Language
Korean
Type
Conference Paper
Abstract
Embodied Question Answering (EQA) requires agents to sustain a representation of the world while answering multi-turn queries in real time. A key challenge is how to maintain and update this world model efficiently under resource constraints. Existing approaches repeatedly re-encode visual inputs or apply retrieval-augmented generation, both of which introduce latency that limits interactive use. We propose an episode-level multimodal KV cache that is constructed once from uniformly sampled frames and reused across all queries in the same episode. This cache serves as a lightweight multimodal memory that reduces redundant computation while pre serving relevant context. On the openEQA benchmark, our method achieves up to an 82% reduction in total question-answering time compared to naïve multi-image inference, with only a modest drop in accuracy. These findings demonstrate that reusing an episode-level cache provides an effective mechanism for maintaining and updating world models to achieve efficient reasoning in EQA.
Keyword
Embodied Question Answering, KV Cache
KSP Keywords
Efficient reasoning, Existing Approaches, Multi-image, Question Answering, Relevant context, World model, real time, resource constraints