ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Journal Article Real-Time Anomaly Detection on Edge Devices via VLM Prompt Optimization
Cited - time in scopus Download 18 time Share share facebook twitter linkedin kakaostory
Authors
Sungmin Yu, Jongwon Moon, Hosub Yoon
Issue Date
2026-07
Citation
Electronics (Switzerland), v.15, no.15, pp.1-19
ISSN
2079-9292
Publisher
Multidisciplinary Digital Publishing Institute (MDPI)
Language
English
Type
Journal Article
DOI
https://dx.doi.org/10.3390/electronics15153305
Abstract
Real-time video anomaly detection (VAD) under realistic edge constraints—sub-second latency, ≤25 W power, no cloud dependency, and human-interpretable output—remains an open problem. Existing lightweight video convolutional neural networks (X3D, MoViNets) are bound to closed-set training distributions, while recent vision–language-model-based VAD methods (LAVAD, VERA, Holmes-VAD) achieve 80–89% area under the curve (AUC) but rely on datacenter-grade GPUs and Chain-of-Thought (CoT) reasoning that pushes per-segment latency well above one second. This paper reframes the design target from peak accuracy to practical edge deployability and contributes two tightly coupled designs: (i) an edge-optimized inference stack that compresses Qwen3-VL-2B with 4-bit Activation-aware Weight Quantization (INT4 AWQ) and serves it through a TensorRT-LLM C++ runtime on NVIDIA Jetson Orin NX (16 GB, 25 W); and (ii) a fully automatic, CoT-free verbalized prompt optimization in which an 8B optimizer iteratively refines a natural-language definition block 𝐷𝑡 using class-balanced (stratified) development batches on a disjoint development subset, with no human editing and no runtime cost on the edge device. Three findings support this framing: (a) the inference stack reduces per-segment latency to 0.25 s, a 7.4× speed-up and 55% memory reduction over a Python/PyTorch baseline; (b) verbalized prompt optimization improves zero-shot AUC from 71.82% (manual prompt) to 76.39%, outperforming GPT-4- and Gemini-Pro-generated prompts (74.12% and 74.35%) under the same edge backbone; and (c) single-frame input attains the highest mean AUC among one-, five-, and eight-frame windows—statistically comparable to the five-frame setting—while offering the lowest latency, making it the preferred operating point under the edge budget. While the absolute AUC (76.39%) is below recent server-side methods (CLIP-TSA 87.58%, VadCLIP 88.02%, Holmes-VAD 89.51%), our framework is the only one in this comparison that operates entirely on a ≤25 W edge device, providing a deployment-oriented operating point on the accuracy–feasibility frontier of VLM-based VAD.
Keyword
anomaly detection, edge computing, Jetson Orin NX, streaming processing, TensorRT-LLM, verbalized learning, prompt optimization, vision–language model
KSP Keywords
Class-balanced, Closed-set, Convolution neural network(CNN), Design target, Edge Computing, Edge devices, Language Models, Memory reduction, Open Problem, Operating Point, Real-Time Video
This work is distributed under the term of Creative Commons License (CCL)
(CC BY)
CC BY