ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Conference Paper QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer
Cited - time in scopus Download 39 time Share share facebook twitter linkedin kakaostory
Authors
Sehyeon Oh, Yongin Kwon, Jemin Lee
Issue Date
2026-08
Citation
International Joint Conference on Artificial Intelligence (IJCAI) 2026, pp.1-9
Publisher
International Joint Conferences on Artificial Intelligence Organization (IJCAI)
Language
English
Type
Conference Paper
Abstract
FlashAttention improves efficiency through tiling, but its online softmax still relies on floating-point arithmetic for numerical stability, making full quantization difficult. We identify three main obstacles to integer-only FlashAttention: (1) scale explosion during tile-wise accumulation, (2) inefficient shift-based exponential operations on GPUs, and (3) quantization granularity constraints requiring uniform scales for integer comparison. To address these challenges, we propose \textit{QFlash}, an end-to-end integer FlashAttention design that performs softmax entirely in the integer domain and runs as a single Triton kernel. On seven attention workloads from ViT, DeiT, and Swin models, QFlash achieves up to 6.73$\times$ speedup over I-ViT and up to 8.69$\times$ speedup on Swin, while reducing energy consumption by 18.8\% compared to FP16 FlashAttention, without sacrificing Top-1 accuracy on ViT/DeiT and remaining competitive on Swin under per-tensor quantization. Our code is publicly available at https://github.com/EfficientCompLab/qflash.
KSP Keywords
End to End(E2E), Floating-Point Arithmetic, Memory Efficiency, Numerical Stability, reducing energy consumption
© 2026 International Joint Conferences on Artificial Intelligence Organization (IJCAI). This is the preprint version made available on the official IJCAI-ECAI 2026 website.