| Lightweight Skeleton Feature Fusion for Video Anomaly Detection | |
|---|---|
| 學年 | 115 |
| 學期 | 1 |
| 發表日期 | 2026-09-05 |
| 作品名稱 | Lightweight Skeleton Feature Fusion for Video Anomaly Detection |
| 作品名稱(其他語言) | |
| 著者 | Zih-Syuan Liou, Chen-Chien Hsu, Shao-Kang Huang, and Cheng-Kai Lu |
| 作品所屬單位 | |
| 出版者 | |
| 會議名稱 | ICCE Berlin 2026 |
| 會議地點 | 德國 |
| 摘要 | Video anomaly detection (VAD) aims to detect events that deviate from normal behavior in surveillance videos. Existing RGB-based methods primarily rely on appearance cues and may fail when anomalies are characterized by subtle human motion patterns. In contrast, skeleton-based representations capture human pose geometry, providing complementary behavioral information. This paper presents a lightweight latefusion framework that augments Jigsaw-VAD with skeletonderived cues through a plug-in design. We introduce three physically interpretable features, including Pose Extension, Trajectory Tortuosity, and Pose Velocity Magnitude, to describe behavior in terms of geometric spread, path regularity, and motion intensity, respectively. The framework supports both linear and attention-based fusion without changing the original RGB model. Experiments on ShanghaiTech Campus dataset show that linear fusion improves the AUC from 84.2% to 86.5%, while reducing the cross-fold standard deviation from ±1.76% to ±0.53%. Additional analysis further demonstrates positive complementarity among the proposed features. The attention-based fusion module introduces only about 3.6 kFLOPs, resulting in negligible computational overhead. |
| 關鍵字 | video anomaly detection; skeleton feature fusion; late fusion; self-supervised learning; complementarity analysis |
| 語言 | en_US |
| 收錄於 | |
| 會議性質 | 國際 |
| 校內研討會地點 | 無 |
| 研討會時間 | 20260905~20260907 |
| 通訊作者 | Shao-Kang Huang |
| 國別 | DEU |
| 公開徵稿 | |
| 出版型式 | |
| 出處 | |
| 相關連結 |
機構典藏連結 ( http://tkuir.lib.tku.edu.tw:8080/dspace/handle/987654321/129889 ) |