Suhang Cai

dblp:388/4517 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
0009-0004-4364-2035ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Video understanding and tracking · 44% Representation and self-supervised learning · 44% Deep learning architectures and training · 13%

Topics — the 3 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › reconstruction-based representation learning
feature reconstruction
0.812024
Feature Reconstruction With Disruption for Unsupervised Video Anomaly Detection · IEEE Trans. Multim. 2024
Computer vision › Video understanding and tracking
video anomaly detection
0.812024
Feature Reconstruction With Disruption for Unsupervised Video Anomaly Detection · IEEE Trans. Multim. 2024
Machine learning › Deep learning architectures and training
transformer
0.212024
Feature Reconstruction With Disruption for Unsupervised Video Anomaly Detection · IEEE Trans. Multim. 2024

Methods — techniques the papers use, named apart from their topics

pseudo-labeling · 0.8memory bank · 0.8cross-attention · 0.8
YearPublicationVenuePosition
2024 Feature Reconstruction With Disruption for Unsupervised Video Anomaly Detection
abstract
Unsupervised video anomaly detection (UVAD) has gained significant attention due to its label-free nature. Typically, UVAD methods can be categorized into two branches, i.e. the one-class classification (OCC) methods and fully UVAD ones. However, the former may suffer from data imbalance and high false alarm rates, while the latter relies heavily on feature representation and pseudo-labels. In this paper, a novel feature reconstruction and disruption model (FRD-UVAD) is proposed for effective feature refinement and better pseudo-label generation in fully UVAD, based on cascade cross-attention transformers, a latent anomaly memory bank and an auxiliary scorer. The clip features are reconstructed using the space-time intra-clip information, as well as cross-inter-clip knowledge. Moreover, instead of blindly reconstructing all training features as OCC methods, a new disruption process is proposed to cooperate with the feature reconstruction simultaneously. Using the collected pseudo anomaly samples, it is able to emphasize the feature differences between normal and abnormal events. Additionally, a pre-trained UVAD scorer is utilized as a different criteria for anomaly prediction, which further refines the pseudo-labels. To demonstrate its effectiveness, comprehensive experiments and detailed ablation studies are conducted on three video benchmarks, namely CUHK Avenue, ShanghaiTech and UCF-Crime. Our proposed model (FRD-UVAD) achieves the best AUC performance (91.23%, 80.14%, and 82.12%) on all three datasets, surpassing other state-of-the-art OCC and fully UVAD methods. Furthermore, it obtains the lowest false alarm rate with a lower scene dependency, compared with other OCC methods. The code is available athttps://github.com/tcc-power/FRD-unsupervised-video-anomaly-detection.
Chenchen Tao, Chong Wang 0001, Sunqi Lin, Suhang Cai, Jiangbo Qian
IEEE Trans. Multim.4