VLDB 2026 Research / reviewers in the wild / expert
Peng Wu 0015
dblp:15/6146-15
· DBLP profile ↗
36ranked-venue papers
14as first author
31since 2021 · last 2026
0000-0003-2938-6798ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 11 first-author · 26 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 13 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TargetVAU: Multimodal Anomaly-Aware Reasoning for Target Behavior Understanding in VideosabstractUnderstanding anomalous human behaviors at a fine-grained level remains a major challenge in complex scenarios. Existing video anomaly understanding (VAU) methods often rely on coarse frame-level cues or overlook structured modeling of individual actions, limiting their capacity for reasoning about human interactions and accountability. To address these challenges, we propose TargetVAU, a multimodal anomaly-aware reasoning framework designed for individual-level anomaly recognition and explanation. TargetVAU first extracts both global-level and human-centric visual features using a frozen Vision Transformer (ViT) encoder. An Anomaly-focused Temporal Sampler is then employed to select behaviorally informative frames via a density-aware strategy guided by predicted anomaly scores. A Spatio-Temporal Interaction Graph is constructed to explicitly model interactions among individuals across time and space. These structured representations are fused with prompt embeddings via a frozen Q-Former to form a unified semantic representation. Finally, a large language model fine-tuned with low-rank adaptation (LoRA) performs instruction-guided reasoning to identify anomalous individuals and generate natural language explanations. Extensive experiments on UCCD and HIVAU-70K demonstrate that TargetVAU significantly outperforms existing methods in both accuracy and interpretability, advancing the state of individual-level anomaly understanding in surveillance videos. Lingru Zhou, Peng Wu 0015, Manqing Zhang, Qingsheng Wang, Guansong Pang, Peng Wang 0015 |
AAAI | 2 |
| 2026 | Single-frame supervision for temporal video anomaly grounding
Yuzhou Long, Peng Wu 0015, Guansong Pang, Peng Wang 0015, Yanning Zhang 0001 |
Neurocomputing | 2 |
| 2026 | Semantic consistency-aware pseudo-temporal framework for multimodal remote sensing image segmentation
Yuejiang Li, Weisheng Dong, Peng Wu 0015, Lichao Mou, Xin Li 0005 |
Neural Networks | 5 |
| 2026 | HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement
Kangbiao Shi, Yixu Feng, Tao Hu 0013, Peng Wu 0015, Guansong Pang, Qingsen Yan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Boosting HDR Image Reconstruction via Semantic Knowledge TransferabstractRecovering High Dynamic Range (HDR) images from multiple Standard Dynamic Range (SDR) images becomes challenging when the SDR images exhibit noticeable degradation and missing content. Leveraging scene-specific semantic priors offers a promising solution for restoring heavily degraded regions. However, these priors are typically extracted from sRGB SDR images, the domain/format gap poses a significant challenge when applying it to HDR imaging. To address this issue, we propose a general framework that transfers semantic knowledge derived from SDR domain via self-distillation to boost existing HDR reconstruction. Specifically, the proposed framework first introduces the Semantic Priors Guided Reconstruction Model (SPGRM), which leverages SDR image semantic knowledge to address ill-posed problems in the initial HDR reconstruction results. Subsequently, we leverage a self-distillation mechanism that constrains the color and content information with semantic knowledge, aligning the external outputs between the baseline and SPGRM. Furthermore, to transfer the semantic knowledge of the internal features, we utilize a Semantic Knowledge Alignment Module (SKAM) to fill the missing semantic contents with the complementary masks. Extensive experiments demonstrate that our framework significantly boosts HDR imaging quality for existing methods without altering the network architecture. Tao Hu 0013, Longyao Wu, Wei Dong 0010, Peng Wu 0015, Jinqiu Sun, Xiaogang Xu 0002, Qingsen Yan, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Ghost-Free HDR Imaging via Latent Low-Frequency Priors and Deformable Attention AlignmentabstractRecovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, showing promising performance, particularly in achieving visually perceptible better results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. These priors are integrated into a carefully designed Dynamic HDR Reconstruction Network (DHRNet), which employs a regression-based approach to produce high-quality HDR images. Furthermore, we introduce the Attention-guided Deformable Alignment Module (ADAM) that utilizes correlation-driven feature matching to learn deformable receptive fields for self-attention, enabling efficient pre-alignment of LDR images by focusing on salient regions. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is $10\times $ faster than previous DM-based methods. Tao Hu 0013, Qingsen Yan, Wei Dong 0010, Peng Wu 0015, Yuankai Qi, Weisi Lin, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Deep Learning for Video Anomaly Detection: A ReviewabstractVideo anomaly detection (VAD) aims to discover behaviors or events deviating from the normality in videos. As a long-standing task in the field of computer vision, VAD has witnessed much good progress. In the era of deep learning, with the explosion of architectures of continuously growing capability and capacity, a great variety of deep learning-based methods are constantly emerging for the VAD task, greatly improving the generalization ability of detection algorithms and broadening the application scenarios. Therefore, such a multitude of methods and a large body of literature make a comprehensive survey a pressing necessity. In this article, we present an extensive and comprehensive research review, covering the spectrum of five different categories, namely, semi-supervised, weakly supervised, fully supervised, unsupervised, and open-set supervised VAD, and we also delve into the latest VAD works based on pretrained large models and open-world learning, remedying the limitations of past reviews in terms of only focusing on semi-supervised VAD and small model-based methods. For the VAD task with different levels of supervision, we construct a well-organized taxonomy, profoundly discuss the characteristics of different types of methods, and show their performance comparisons. In addition, this review involves the public datasets, open-source codes, and evaluation metrics covering all the aforementioned VAD tasks. Finally, we provide several important research directions for the VAD community. Additional details of the survey are available on the project homepage: https://github.com/Roc-Ng/DeepVAD. Peng Wu 0015, Chengyu Pan, Guansong Pang, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2026 | STPrompt\(\boldsymbol{++}\): Prompting Vision-Language Models for Weakly Supervised Video Anomaly Detection and Fine-Grained LocalizationabstractTraditional weakly supervised video anomaly detection (WSVAD) tasks typically rely on coarse-grained frame-level labels for training. Although this approach reduces annotation costs, it results in weak semantic understanding and spatial localization capabilities due to the absence of fine-grained annotations, hindering precise pixel-level anomaly detection and localization. Thanks to the success of vision-language models (VLMs), e.g., CLIP, recent approaches leveraging large VLMs focus on exploiting their strong semantic understanding capabilities, but they typically feed only keyframes or short video segments into the models, without supplying sufficient prior contextual information (e.g., contextual frames around anomalies, zoomed-in anomaly regions, and detailed anomaly descriptions), which restricts the models’ capability for fine-grained anomaly understanding and precise localization. More recently, a few methods leveraging VLMs, attempt to achieve training-free spatial anomaly localization by fusing patch-level visual features with simple textual features. However, these methods employ simplistic textual descriptions, lacking deep semantic comprehension of anomalies, leading to coarse localization results with significant irrelevant background noise. To address these issues, we propose STPrompt \(++\) , a novel weakly supervised spatio-temporal video anomaly detection and localization method based on VLMs. In our work, we systematically leverage preliminary coarse localization regions derived from anomaly scores as spatial priors, together with contextual frames around keyframes, zoomed-in views of suspected anomalous regions, and refined textual descriptions of anomalies. This comprehensive prompting mechanism guides the VLMs toward deep semantic comprehension of video anomalies, enabling accurate pixel-level spatial localization. The proposed STPrompt \(++\) requires no additional training and significantly enhances the precision of anomaly understanding and localization through a carefully designed multi-round and multi-modal prompting mechanism. Extensive experiments on two widely used WSVAD benchmarks, UCF-Crime and UBnormal, show that our method achieves state-of-the-art spatial localization performance and competitive temporal anomaly detection results. Notably, on UCF-Crime dataset, our approach improves spatial localization accuracy (in TIoU) by 5.61% over the current best method (from 23.90% to 29.51%), underscoring its superior capabilities in precise anomaly localization and semantic understanding. Peng Wu 0015, Chengyu Pan, Guansong Pang, Xiangteng He, Zhiwei Yang 0013, Peng Wang 0015, Yanning Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | VarCMP: Adapting Cross-Modal Pre-Training Models for Video Anomaly RetrievalabstractVideo anomaly retrieval (VAR) aims to retrieve pertinent abnormal or normal videos from collections of untrimmed and long videos through cross-modal requires such as textual descriptions and synchronized audios. Cross-modal pre-training (CMP) models, by pre-training on large-scale cross-modal pairs, e.g., image and text, can learn the rich associations between different modalities, and this cross-modal association capability gives CMP an advantage in conventional retrieval tasks. Inspired by this, how to utilize the robust cross-modal association capabilities of CMP in VAR to search crucial visual component from these untrimmed and long videos becomes a critical research problem. Therefore, this paper proposes a VAR method based on CMP models, named VarCMP. First, a unified hierarchical alignment strategy is proposed to constrain the semantic and spatial consistency between video and text, as well as the semantic, temporal, and spatial consistency between video and audio. It fully leverages the efficient cross-modal association capabilities of CMP models by considering cross-modal similarities at multiple granularities, enabling VarCMP to achieve effective all-round information matching for both video-text and video-audio VAR tasks. Moreover, to further solve the problem of untrimmed and long video alignment, an anomaly-biased weighting is devised in the fine-grained alignment, which identifies key segments in untrimmed long videos using anomaly priors, giving them more attention, thereby discarding irrelevant segment information, and achieving more accurate matching with cross-modal queries. Extensive experiments demonstrates high efficacy of VarCMP in both video-text and video-audio VAR tasks, achieving significant improvements on both text-video (UCFCrime-AR) and audio-video (XDViolence-AR) datasets against the best competitors by 5.0% and 5.3% R@1. Peng Wu 0015, Wanshun Su, Xiangteng He, Peng Wang 0015, Yanning Zhang 0001 |
AAAI | 1 |
| 2025 | HVI: A New Color Space for Low-light Image EnhancementabstractLow-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce color bias and brightness artifacts due to inherent high color sensitivity in sRGB. While converting the images using Hue, Saturation and Value (HSV) color space helps resolve the brightness issue, it introduces significant red and black noise artifacts. To address this issue, we propose a new color space for LLIE, namely Horizontal/Vertical-Intensity (HVI), defined by polarized HS maps and learnable intensity. The former enforces small distances for red coordinates to remove the red artifacts, while the latter compresses the low-light regions to remove the black artifacts. To fully leverage the chromatic and intensity information, a novel Color and Intensity Decoupling Network (CIDNet) is further introduced to learn accurate photometric mapping function under different lighting conditions in the HVI space. Comprehensive results from benchmark and ablation experiments show that the proposed HVI color space with CIDNet outperforms the state-of-the-art methods on 10 datasets. The code is available at https://github.com/Fediory/HVI-CIDNet. Qingsen Yan, Yixu Feng, Guansong Pang, Kangbiao Shi, Peng Wu 0015, Wei Dong 0010, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 6 |
| 2025 | Text-Visual Semantic Constrained AI-Generated Image Quality Assessment
Qingsen Yan, Haojian Huang, Peng Wu 0015, Haokui Zhang, Yanning Zhang 0001 |
ACM Multimedia | 4 |
| 2025 | Distilling Hierarchical Knowledge From Multimodal Fusion for Unimodal Image SegmentationabstractThe application of multimodal image fusion has become increasingly widespread across various fields in the era of deep learning. Existing fusion methods integrate infrared and visible images to provide complementary content and enhance the robustness of complex real-world scenes for high-level visual tasks, such as semantic segmentation and object detection. In return, high-level visual tasks facilitate the fusion of infrared and visible by providing mid-level semantic information. However, such frameworks rely heavily on multimodal data and require strict registration of images from different modalities before fusion, seriously limiting their practical applications due to the common realistic situations of missing modalities or misregistration. To move beyond this limitation, we propose a novel hierarchical knowledge distillation (HKD) framework tailored for unimodal image segmentation with the guidance of multi-modality. This framework aims to retain as much diverse information from multimodal image fusion as possible, thereby enhancing downstream high-level visual tasks when only the unimodal images are available during the inference phase. Our proposed method is two-stage, and we construct a robust multimodal fusion and segmentation interaction network in the first stage as a powerful teacher model. In the second stage, we design a hierarchical distillation method to transfer the fused and segmented multi-layer knowledge from the multimodal teacher model to the unimodal student model. Extensive experimental results on two public datasets, i.e., MFNet and FMB, demonstrate that the proposed hierarchical knowledge distillation framework can effectively transfuse multimodal knowledge into the unimodal student model for image enhancement and segmentation under incomplete multimodal conditions, and achieves considerably competitive results compared to multimodal image fusion and segmentation models. Weisheng Dong, Shuaibo Wang, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Efficient Image Enhancement With a Diffusion-Based Frequency PriorabstractDue to the lack of appropriate priors, generating the content of dark regions remains a challenge in low-light image enhancement tasks. Currently, diffusion models employ robust image generation capabilities for enhancing low-light images. However, diffusion models require multiple iterations at the image feature level to generate details and content, which limits the speed. Moreover, the diffusion-based methods tend to generate unexpected artifacts in the degraded regions. To address these issues, we propose a Frequency Priors-guided Image Enhancement (FPIE) network, including a frequency prior generation network and an image restoration network. FPIE significantly accelerates inference by learning abstract prior with frequency domain constraints. Concretely, to learn compacted priors at the frequency domain, we introduce a joint training approach for the prior generation and restoration models to constrain the distribution of priors. Furthermore, to better utilize frequency-domain features for enhancing the network’s generation capabilities, a wavelet-based transformer block is introduced to produce intricate details and avoid the artifacts of the output. Extensive experimental results on the commonly used benchmarks demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images. Qingsen Yan, Tao Hu 0013, Peng Wu 0015, Duwei Dai, Shuhang Gu, Wei Dong 0010, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | From Dynamic to Static: Stepwisely Generate HDR Image for Ghost RemovalabstractGenerating high-quality high dynamic range (HDR) images in dynamic scenes is particularly challenging due to the influence of large motion. Despite the effectiveness of existing deep learning methods, they still suffer from ghosting artifacts when saturation and motion coexist. Inspired by fusion on static scenes, we propose an inpainting and fusion strategy to enhance the quality of the generated HDR images. The proposed method consists of pseudo-static LDR generation and detail-guided HDR generation, which creates pseudo-static images and then generates ghost-free HDR images. Specifically, the pseudo-static LDR generation network utilizes semantic information to identify the motion regions, and employs a diffusion model-based inpainting approach to produce pseudo-static LDR images that closely resemble real scenes. In the detail-guided HDR generation network, we employ a detail enhancement module to refine diverse high-frequency features with detailed information extracted from pseudo-static LDR images, which effectively enhances the visual quality. Extensive experiments on four public datasets demonstrate the superiority of the proposed method, both quantitatively and qualitatively. Qingsen Yan, Kangzhen Yang, Tao Hu 0013, Genggeng Chen, Kexin Dai, Peng Wu 0015, Wenqi Ren, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Incomplete Modalities Restoration via Hierarchical Adaptation for Robust Multimodal SegmentationabstractMultimodal semantic segmentation has significantly advanced the field of semantic segmentation by integrating data from multiple sources. However, this task often encounters missing modality scenarios due to challenges such as sensor failures or data transmission errors, which can result in substantial performance degradation. Existing approaches to addressing missing modalities predominantly involve training separate models tailored to specific missing scenarios, typically requiring considerable computational resources. In this paper, we propose a Hierarchical Adaptation framework to Restore Missing Modalities for Multimodal segmentation (HARM3), which enables frozen pretrained multimodal models to be directly applied to missing-modality semantic segmentation tasks with minimal parameter updates. Central to HARM3 is a text-instructed missing modality prompt module, which learns multimodal semantic knowledge by utilizing available modalities and textual instructions to generate prompts for the missing modalities. By incorporating a small set of trainable parameters, this module effectively facilitates knowledge transfer between high-resource domains and low-resource domains where missing modalities are more prevalent. Besides, to further enhance the model's robustness and adaptability, we introduce adaptive perturbation training and an affine modality adapter. Extensive experimental results demonstrate the effectiveness and robustness of HARM3 across a variety of missing modality scenarios. Weisheng Dong, Peng Wu 0015, Mingtao Feng, Xin Li 0005, Guangming Shi |
IEEE Trans. Image Process. | 3 |
| 2024 | VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly DetectionabstractThe recent contrastive language-image pre-training (CLIP) model has shown great success in a wide range of image-level tasks, revealing remarkable ability for learning powerful visual representations with rich semantics. An open and worthwhile problem is efficiently adapting such a strong model to the video domain and designing a robust video anomaly detector. In this work, we propose VadCLIP, a new paradigm for weakly supervised video anomaly detection (WSVAD) by leveraging the frozen CLIP model directly without any pre-training and fine-tuning process. Unlike current works that directly feed extracted features into the weakly supervised classifier for frame-level binary classification, VadCLIP makes full use of fine-grained associations between vision and language on the strength of CLIP and involves dual branch. One branch simply utilizes visual features for coarse-grained binary classification, while the other fully leverages the fine-grained language-image alignment. With the benefit of dual branch, VadCLIP achieves both coarse-grained and fine-grained video anomaly detection by transferring pre-trained knowledge from CLIP to WSVAD task. We conduct extensive experiments on two commonly-used benchmarks, demonstrating that VadCLIP achieves the best performance on both coarse-grained and fine-grained WSVAD, surpassing the state-of-the-art methods by a large margin. Specifically, VadCLIP achieves 84.51% AP and 88.02% AUC on XD-Violence and UCF-Crime, respectively. Code and features are released at https://github.com/nwpu-zxr/VadCLIP. Peng Wu 0015, Xuerong Zhou, Guansong Pang, Lingru Zhou, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001 |
AAAI | 1 |
| 2024 | Open-Vocabulary Video Anomaly DetectionabstractCurrent video anomaly detection (VAD) approaches with weak supervisions are inherently limited to a closed-set setting and may struggle in open-world applications where there can be anomaly categories in the test data unseen during training. A few recent studies attempt to tackle a more realistic setting, open-set VAD, which aims to de-tect unseen anomalies given seen anomalies and normal videos. However, such a setting focuses on predicting frame anomaly scores, having no ability to recognize the specific categories of anomalies, despite the fact that this ability is essential for building more informed video surveillance systems. This paper takes a step further and explores open-vocabulary video anomaly detection (OVVAD), in which we aim to leverage pretrained large models to detect and cate-gorize seen and unseen anomalies. To this end, we propose a model that decouples OVVAD into two mutually comple-mentary tasks - class-agnostic detection and class-specific classification - and jointly optimizes both tasks. Particu-larly, we devise a semantic knowledge injection module to introduce semantic knowledge from large language models for the detection task, and design a novel anomaly synthesis module to generate pseudo unseen anomaly videos with the help of large vision generation models for the classification task. These semantic knowledge and synthesis anomalies substantially extend our model's capability in detecting and categorizing a variety of seen and unseen anomalies. Exten-sive experiments on three widely-used benchmarks demonstrate our model achieves state-of-the-art performance on OVVAD task. Peng Wu 0015, Xuerong Zhou, Guansong Pang, Jing Liu 0006, Peng Wang 0015, Yanning Zhang 0001 |
CVPR | 1 |
| 2024 | Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection (WSVAD) is a challenging task. Generating fine-grained pseudo-labels based on weak-label and then self-training a classifier is currently a promising solution. However, since the existing methods use only RGB visual modality and the utilization of category text information is neglected, thus limiting the generation of more accurate pseudo-labels and affecting the performance of self-training. Inspired by the manual labeling process based on the event description, in this paper, we propose a novel pseudo-label generation and self-training framework based on Text Prompt with Normality Guidance (TPWNG) for WSVAD. Our idea is to transfer the rich language-visual knowledge of the contrastive language-image pre-training (CLIP) model for aligning the video event description text and corresponding video frames to generate pseudo-labels. Specifically, We first fine-tune the CLIP for domain adaptation by designing two ranking losses and a distributional inconsistency loss. Further, we propose a learnable text prompt mechanism with the assist of a normality visual prompt to further improve the matching accuracy of video event description text and video frames. Then, we design a pseudo-label generation module based on the normality guidance to infer reliable frame-level pseudo-labels. Finally, we introduce a temporal context self-adaptive learning module to learn the temporal dependencies of different video events more flexibly and accurately. Extensive experiments show that our method achieves state-of-the-art performance on two benchmark datasets, UCF-Crime and XD-Violence, demonstrating the effectiveness of our proposed method. Zhiwei Yang 0013, Jing Liu 0006, Peng Wu 0015 |
CVPR | 3 |
| 2024 | Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsabstractCurrent weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global features from full-resolution video frames and training frame-level classifiers to detect anomalies in the temporal dimension. However, most anomalous events tend to occur in localized spatial regions rather than the entire video frames, which implies existing frame-level feature based works may be misled by the dominant background information and lack the interpretation of the detected anomalies. To address this dilemma, this paper introduces a novel method called STPrompt that learns spatio-temporal prompt embeddings for weakly supervised video anomaly detection and localization (WSVADL) based on pre-trained vision-language models (VLMs). Our proposed method employs a two-stream network structure, with one stream focusing on the temporal dimension and the other primarily on the spatial dimension. By leveraging the learned knowledge from pre-trained VLMs and incorporating natural motion priors from raw videos, our model learns prompt embeddings that are aligned with spatio-temporal regions of videos (e.g., patches of individual frames) for identify specific local regions of anomalies, enabling accurate video anomaly detection while mitigating the influence of background information. Without relying on detailed spatio-temporal annotations or auxiliary object detection/tracking, our method achieves state-of-the-art performance on three public benchmarks for the WSVADL task. Peng Wu 0015, Xuerong Zhou, Guansong Pang, Zhiwei Yang 0013, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 1 |
| 2024 | A Transformer-based visual object tracker via learning immediate appearance changeabstractTransformer has shown its great strength in visual object tracking due to its effective attention mechanism , but most prevailing transformer-based trackers only explore temporal information frame by frame, thus overlooking the rich context information inherent in videos. To alleviate this problem, we propose a transformer-based tracker via learning immediate appearance change information in videos, called IAC-tracker. The proposed tracker enhances the perception of the immediate motion state to improve the performance of single target tracking . IAC-tracker contains three key components: a spatial information extractor (SIE) with a superior attention mechanism to progressively extract spatial information, a temporal information extractor (TIE) with a designed temporal attention mechanism to progressively learn target immediate appearance change, and a novel spatial–temporal context enhanced fusion module integrating the information from SIE and TIE to prepare for the final prediction head. Comparison experiments with state-of-the-art trackers on six challenging datasets demonstrate the superior performance of IAC-tracker with real-time running speed. Dian Yuan, Jiaoying Wang, Peng Wu 0015, Jing Liu 0006 |
Pattern Recognit. | 5 |
| 2024 | Fine-Granularity Alignment for Text-Based Person Retrieval Via Semantics-Centric Visual DivisionabstractText-based Person Retrieval aims to search the target pedestrian image from video surveillance or a large image database with a text description. Previous works have recognized the significance of mining local information in images and descriptions and performing fine-grained alignment. These approaches adopt hard division or auxiliary networks for locating local visual regions. However, the two existing ways are not flexible enough for various images and may even bring noise. Meanwhile, the Vision-Language Pre-training models like CLIP exhibit strong generalization and zero-shot abilities, which provide an available way to this issue. In this paper, we propose a novel Fine-Granularity Alignment model with Semantics-Centric Visual Division (SCVD). Our method contains a Semantics Deconstructor (SD), a Cross-modal Guided Interaction (CGI) module, and a Dynamic Focus Alignment (DFA) module. The SD aims to extract fine-grained semantic prompts from the raw description which is easy-understand for CLIP. In CGI, we propose a Text-Guided Visual Localization (TVL) module to generate local visual representations according to the semantic prompts and a Vision-Guided Semantics Reconstruction (VSR) module to integrate the prompts into the textual representation. The DFA is used finally to align vision-text fine-grained information. The extensive experiments demonstrate that our proposed framework significantly outperforms current state-of-the-art methods in terms of Rank@1 metric on three benchmarks by an absolute gain of 6.56%, 8.93%, and 11.53%, respectively. Our code is available in https://github.com/tujun233/SCVD.git. Zhimin Wei, Peng Wu 0015, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Toward Video Anomaly Retrieval From Video Anomaly Detection: New Benchmarks and ModelabstractVideo anomaly detection (VAD) has been paid increasing attention due to its potential applications, its current dominant tasks focus on online detecting anomalies, which can be roughly interpreted as the binary or multiple event classification. However, such a setup that builds relationships between complicated anomalous events and single labels, e.g., "vandalism", is superficial, since single labels are deficient to characterize anomalous events. In reality, users tend to search a specific video rather than a series of approximate videos. Therefore, retrieving anomalous events using detailed descriptions is practical and positive but few researches focus on this. In this context, we propose a novel task called Video Anomaly Retrieval (VAR), which aims to pragmatically retrieve relevant anomalous videos by cross-modalities, e.g., language descriptions and synchronous audios. Unlike the current video retrieval where videos are assumed to be temporally well-trimmed with short duration, VAR is devised to retrieve long untrimmed videos which may be partially relevant to the given query. To achieve this, we present two large-scale VAR benchmarks and design a model called Anomaly-Led Alignment Network (ALAN) for VAR. In ALAN, we propose an anomaly-led sampling to focus on key segments in long untrimmed videos. Then, we introduce an efficient pretext task to enhance semantic associations between video-text fine-grained representations. Besides, we leverage two complementary alignments to further match cross-modal contents. Experimental results on two benchmarks reveal the challenges of VAR task and also demonstrate the advantages of our tailored method. Captions are publicly released at https://github.com/Roc-Ng/VAR. Peng Wu 0015, Jing Liu 0006, Xiangteng He, Yuxin Peng 0001, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Human-Centric Behavior Description in Videos: New Benchmark and ModelabstractIn the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each individual's behavior provides more detailed situational analysis, enabling accurate assessment and response to potential risks, ensuring the safety and harmony of public places. Currently, video-level captioning datasets cannot provide fine-grained descriptions for each individual's specific behavior. However, mere descriptions at the video-level fail to provide an in-depth interpretation of individual behaviors, making it challenging to accurately determine the specific identity of each individual. To address this challenge, we construct a human-centric video surveillance captioning dataset, which provides detailed descriptions of the dynamic behaviors of 7,820 individuals. Specifically, we have labeled several aspects of each person, such as location, clothing, and interactions with other elements in the scene, and these people are distributed across 1,012 videos. Based on this dataset, we can link individuals to their respective behaviors, allowing for further analysis of each person's behavior in surveillance videos. Besides the dataset, we propose a novel video captioning approach that can describe individual behavior in detail on a person-level basis, achieving state-of-the-art results. Lingru Zhou, Yiqi Gao, Manqing Zhang, Peng Wu 0015, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | EOGT: Video Anomaly Detection with Enhanced Object Information and Global Temporal DependencyabstractVideo anomaly detection (VAD) aims to identify events or scenes in videos that deviate from typical patterns. Existing approaches primarily focus on reconstructing or predicting frames to detect anomalies and have shown improved performance in recent years. However, they often depend highly on local spatio-temporal information and face the challenge of insufficient object feature modeling. To address the above issues, this article proposes a video anomaly detection framework with E nhanced O bject Information and G lobal T emporal Dependencies (EOGT) and the main novelties are: (1) A L ocal O bject A nomaly S tream (LOAS) is proposed to extract local multimodal spatio-temporal anomaly features at the object level. LOAS integrates two modules: a D iffusion-based O bject R econstruction N etwork (DORN) with multimodal conditions detects anomalies with object RGB information; and an O bject P ose A nomaly Refiner (OPA) discovers anomalies with human pose information. (2) A G lobal T emporal S trengthening S tream (GTSS) with video-level temporal dependencies is proposed, which leverages video-level temporal dependencies to identify long-term and video-specific anomalies effectively. Both streams are jointly employed in EOGT to learn multimodal and multi-scale spatio-temporal anomaly features for VAD, and we finally fuse the anomaly features and scores to detect anomalies at the frame level. Extensive experiments are conducted to verify the performance of EOGT on three public datasets: ShanghaiTech Campus, CUHK Avenue, and UCSD Ped2. Ruoyan Pi, Peng Wu 0015, Xiangteng He, Yuxin Peng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Video Event Restoration Based on Keyframes for Video Anomaly DetectionabstractVideo anomaly detection (VAD) is a significant computer vision problem. Existing deep neural network (DNN) based VAD methods mostly follow the route of frame reconstruction or frame prediction. However, the lack of mining and learning of higher-level visual features and temporal context relationships in videos limits the further performance of these two approaches. Inspired by video codec theory, we introduce a brand-new VAD paradigm to break through these limitations: First, we propose a new task of video event restoration based on keyframes. Encouraging DNN to infer missing multiple frames based on video keyframes so as to restore a video event, which can more effectively motivate DNN to mine and learn potential higher-level visual features and comprehensive temporal context relationships in the video. To this end, we propose a novel U-shaped Swin Transformer Network with Dual Skip Connections (USTN-DSC) for video event restoration, where a cross-attention and a temporal upsampling residual skip connection are introduced to further assist in restoring complex static and dynamic motion object features in the video. In addition, we propose a simple and effective adjacent frame difference loss to constrain the motion consistency of the video sequence. Extensive experiments on benchmarks demonstrate that USTN-DSC outperforms most existing methods, validating the effectiveness of our method. Zhiwei Yang 0013, Jing Liu 0006, Zhaoyang Wu, Peng Wu 0015 |
CVPR | 4 |
| 2023 | WSAD-Net: Weakly Supervised Anomaly Detection in Untrimmed Surveillance Videos
Peng Wu 0015, Yanning Zhang 0001 |
ICIG (5) | 1 |
| 2023 | Weakly Supervised Audio-Visual Violence DetectionabstractViolence detection in videos is very promising in practical applications due to the emergence of massive videos in recent years. Most previous works define violence detection as a simple video classification task and use the single modality of small-scale datasets, e.g., visual signal. However, such solutions are undersupplied. To mitigate this problem, we study weakly supervised violence detection on the large-scale audio-visual violence data, and first introduce two complementary tasks, i.e., coarse-grained violent frame detection and fine-grained violent event detection, to advance the simple violence video classification to frame-level violent event localization, which aims to accurately locate the violent events on untrimmed videos. We then propose a novel network that takes as input audio-visual data and contains three parallel branches to capture different relationships among video snippets and further integrate features, where similarity branch and proximity branch capture long-range dependencies using similarity prior and proximity prior, respectively, and score branch dynamically captures the closeness of predicted score. In both coarse-grained and fine-grained tasks, our approach outperforms other state-of-the-art approaches on two public datasets. Moreover, experiment results also show the positive effect of audio-visual input and relationship modeling. Peng Wu 0015, Jing Liu 0006 |
IEEE Trans. Multim. | 1 |
| 2022 | Dynamic Local Aggregation Network with Adaptive Clusterer for Anomaly Detection
Zhiwei Yang 0013, Peng Wu 0015, Jing Liu 0006 |
ECCV (4) | 2 |
| 2022 | Exploiting foreground and background separation for prohibited item detection in overlapping X-Ray images
Fangtao Shao, Jing Liu 0006, Peng Wu 0015, Zhiwei Yang 0013, Zhaoyang Wu |
Pattern Recognit. | 3 |
| 2021 | HANet: Hierarchical Alignment Networks for Video-Text RetrievalabstractVideo-text retrieval is an important yet challenging task in vision-language understanding, which aims to learn a joint embedding space where related video and text instances are close to each other. Most current works simply measure the video-text similarity based on video-level and text-level embeddings. However, the neglect of more fine-grained or local information causes the problem of insufficient representation. Some works exploit the local details by disentangling sentences, but overlook the corresponding videos, causing the asymmetry of video-text representation. To address the above limitations, we propose a Hierarchical Alignment Network (HANet) to align different level representations for video-text matching. Specifically, we first decompose video and text into three semantic levels, namely event (video and text), action (motion and verb), and entity (appearance and noun). Based on these, we naturally construct hierarchical representations in the individual-local-global manner, where the individual level focuses on the alignment between frame and word, local level focuses on the alignment between video clip and textual context, and global level focuses on the alignment between the whole video and text. Different level alignments capture fine-to-coarse correlations between video and text, as well as take the advantage of the complementary information among three semantic levels. Besides, our HANet is also richly interpretable by explicitly learning key semantic concepts. Extensive experiments on two public datasets, namely MSR-VTT and VATEX, show the proposed HANet outperforms other state-of-the-art methods, which demonstrates the effectiveness of hierarchical representation and alignment. Our code is publicly available at https://github.com/Roc-Ng/HANet. Peng Wu 0015, Xiangteng He, Mingqian Tang, Yiliang Lv, Jing Liu 0006 |
ACM Multimedia | 1 |
| 2021 | Learning Causal Temporal Relation and Feature Discrimination for Anomaly DetectionabstractWeakly supervised anomaly detection is a challenging task since frame-level labels are not given in the training phase. Previous studies generally employ neural networks to learn features and produce frame-level predictions and then use multiple instance learning (MIL)-based classification loss to ensure the interclass separability of the learned features; all operations simply take into account the current time information as input and ignore the historical observations. According to investigations, these solutions are universal but ignore two essential factors, i.e., the temporal cue and feature discrimination. The former introduces temporal context to enhance the current time feature, and the latter enforces the samples of different categories to be more separable in the feature space. In this article, we propose a method that consists of four modules to leverage the effect of these two ignored factors. The causal temporal relation (CTR) module captures local-range temporal dependencies among features to enhance features. The classifier (CL) projects enhanced features to the category space using the causal convolution and further expands the temporal modeling range. Two additional modules, namely, compactness (CP) and dispersion (DP) modules, are designed to learn the discriminative power of features, where the compactness module ensures the intraclass compactness of normal features, and the dispersion module enhances the interclass dispersion. Extensive experiments on three public benchmarks demonstrate the significance of causal temporal relations and feature discrimination for anomaly detection and the superiority of our proposed method. Peng Wu 0015, Jing Liu 0006 |
IEEE Trans. Image Process. | 1 |
| 2020 | Not only Look, But Also Listen: Learning Multimodal Violence Detection Under Weak Supervision
Peng Wu 0015, Jing Liu 0006, Yujia Shi, Fangtao Shao, Zhaoyang Wu, Zhiwei Yang 0013 |
ECCV (30) | 1 |
| 2020 | Fast sparse coding networks for anomaly detection in videos
Peng Wu 0015, Jing Liu 0006, Fang Shen |
Pattern Recognit. | 1 |
| 2020 | Evolutionary Network Embedding Preserving Both Local Proximity and Community StructureabstractThe complex network is an important tool to represent relational data in nature and human society, which has been widely applied in various real-world application scenarios. A key issue for analyzing the features of networks is to represent the characteristic information in the network with rationality. Network embedding, attracting plenty of attention recently, aims to convert network information into a low-dimensional space while maintaining the structure and properties of the network maximally. Most of the existing network embedding methods intend to preserve the pairwise relationship or similarity between nodes, but the community structure, which is one of the most important features of complex networks, is largely ignored. In this article, we propose a novel network embedding method based on evolutionary algorithm (EA), termed as EA-NECommunity, which can preserve both the local proximity of nodes and the community structure of the network by optimizing a carefully designed objective function. The number of communities in the network can be automatically determined without any prior knowledge. Moreover, taking the intrinsic properties of the network embedding problems in mind, we design a local search operator based on multidirectional search which can effectively find feasible solutions. In the experiments, we first visualize the embedding representation obtained by different algorithms, and then use the problems of node clustering, node classification, and link prediction to further validate the quality of the embedding representation obtained. The experimental results show that EA-NECommunity outperforms other state-of-the-art algorithms on both the real life and synthetic networks. Jing Liu 0006, Peng Wu 0015, Xiangyi Teng |
IEEE Trans. Evol. Comput. | 3 |
| 2020 | A Deep One-Class Neural Network for Anomalous Event Detection in Complex ScenesabstractHow to build a generic deep one-class (DeepOC) model to solve one-class classification problems for anomaly detection, such as anomalous event detection in complex scenes? The characteristics of existing one-class labels lead to a dilemma: it is hard to directly use a multiple classifier based on deep neural networks to solve one-class classification problems. Therefore, in this article, we propose a novel DeepOC neural network, termed as DeepOC, which can simultaneously learn compact feature representations and train a DeepOC classifier. Only with the given normal samples, we use the stacked convolutional encoder to generate their low-dimensional high-level features and train a one-class classifier to make these features as compact as possible. Meanwhile, for the sake of the correct mapping relation and the feature representations' diversity, we utilize a decoder in order to reconstruct raw samples from these low-dimensional feature representations. This structure is gradually established using an adversarial mechanism during the training stage. This mechanism is the key to our model. It organically combines two seemingly contradictory components and allows them to take advantage of each other, thus making the model robust and effective. Unlike methods that use handcrafted features or those that are separated into two stages (extracting features and training classifiers), DeepOC is a one-stage model using reliable features that are automatically extracted by neural networks. Experiments on various benchmark data sets show that DeepOC is feasible and achieves the state-of-the-art anomaly detection results compared with a dozen existing methods. Peng Wu 0015, Jing Liu 0006, Fang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Double Complete D-LBP with Extreme Learning Machine Auto-Encoder and Cascade Forest for Facial Expression AnalysisabstractAlthough the obtained accuracy on some lab-controlled facial expression datasets has been very high, the recognition of facial expressions in wild environments is still a challenging problem. Local Binary Patterns (LBP) is a widely used operator in facial expression recognition. However, there are few variations of LBP operators specifically designed for facial expression recognition. In this paper, we propose a novel representation approach called the Double Complete d-LBP (Double Cd-LBP) according to the characteristics of facial expressions. Two d-LBP are employed to represent details and the contour of faces separately, and complete LBP is used to take sign and magnitude components into account. Moreover, multi-scale LBP is exploited to obtain local texture and global information. We then use the extreme learning machine auto-encoder (ELM-AE) as the feature selection approach to learn the discriminative feature. Cascade forest is employed as the final decision classifier. Experiments conducted on the six facial expression databases, including both lab-controlled and wild environments databases, show that our method outperforms or on par with state-of-the-arts. Fang Shen, Jing Liu 0006, Peng Wu 0015 |
ICIP | 3 |