Learning Suspected Anomalies from Event Prompts for Video Anomaly Detection

摘要

Most models for Weakly Supervised Video Anomaly Detection (WS-VAD) rely on multiple instance learning, aiming to distinguish normal and abnormal snippets without specifying the type of anomaly. However, the ambiguous nature of anomaly definitions across contexts may introduce inaccuracy in discriminating abnormal and normal events. To show the model what is anomalous, a novel framework is proposed to guide the learning of suspected anomalies from event prompts. Given a textual prompt dictionary of potential anomaly events and the captions generated from anomaly videos, the semantic anomaly similarity between them could be calculated to identify the suspected events for each video snippet. It enables a new multi-prompt learning process to constrain the visual-semantic features across all videos, as well as provides a new way to label pseudo anomalies for self-training. Comprehensive experiments and detailed ablation studies are conducted on four datasets, namely XD-Violence, UCF-Crime, TAD, and ShanghaiTech. The proposed model outperforms most state-of-the-art methods and shows promising performance in open-set and cross-dataset cases.

出版物
ACM Transactions on Multimedia Computing, Communications, and Applications
陶晨晨
陶晨晨
硕士生

研究方向为视频异常检测与行为识别,关注特征重构、事件提示学习和大规模监控视频建模在异常检测中的应用。

彭晓浩
彭晓浩
硕士生

研究方向为视频异常检测、行为识别与大模型,关注图时序融合、场景感知边界和多专家模型在开放世界异常检测中的应用。

王 翀
王 翀
副教授

研究兴趣:人机交互、人工智能、计算机视觉、多媒体计算.