EDBT 2026 Demo / reviewers in the wild / expert
Kewei Wu
dblp:00/8818
· DBLP profile ↗
21ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-7332-5653ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Confidence-Aware Prototypes for Weakly-Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection aims to identify abnormal snippets in untrimmed videos. Existing methods learn prototypes to describe global representation of snippet distributions. But, in weakly-labeled videos, the normal snippets in abnormal video may take high-uncertainty labels for distribution modeling. Without confidence-aware modeling, abnormal/normal prototype distributions may overlap with each other, leading to inaccurate predictions. In this work, we propose the Unified Confident Prototype (UCP) model, which contains a feature extractor, a confidence-aware prototype learner, and a local-global prototype unifier. The prototype learning is designed to ensure proper separability, stability, and representation.First, after learning the weight of each snippet’s loss, snippets with high-uncertainty labels may take small weights. These snippets tend to lie in the overlap between abnormal/normal distributions, hindering their separation. We design uncertainty-aware sampling, which removes high-uncertainty snippets in the small-weight snippets to ensure separable prototype learning.Second, snippets with high-uncertainty labels tend to be far from the prototype center, thus falling in the low-confidence region. These snippets may enlarge the distribution’s variation, resulting in unstable prototype learning. We design confidence-aware sampling, which removes low-confidence snippets to ensure stable prototype learning.Third, after assigning pseudo labels to prototypes, we measure the prototype representation with the distribution’s purity. We design prototype distribution purification, which penalizes normal snippets in the abnormal-majority distribution with purity loss to ensure representative prototype learning.Fourth, beyond prototype learning, prototypes can be enhanced by local/global temporal semantics. We further introduce the local-global prototype unifier to learn the relations across local-global durations, thereby enhancing the semantics for anomaly detection. For weakly-supervised anomaly detection, experiments demonstrate that our method achieves state-of-the-art performance on the UCF-Crime, ShanghaiTech, and XD-Violence datasets. Moreover, to further verify the generality of our method, we further conduct experiments on THUMOS’14 for weakly-supervised temporal action localization. Zhao Xie, Jinkang Luo, Kewei Wu, Zhehan Kan, Dan Guo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly DetectionabstractVideo anomaly detection (VAD) has been extensively researched due to its potential for intelligent video systems. However, most existing methods based on CNNs and transformers still suffer from substantial computational burdens and have room for improvement in learning spatial-temporal normality. Recently, Mamba has shown great potential for modeling long-range dependencies with linear complexity, providing an effective solution to the above dilemma. To this end, we propose a lightweight and effective Mamba-based network named STNMamba, which incorporates carefully designed Mamba modules to enhance the learning of spatial-temporal normality. Firstly, we develop a dual-encoder architecture, where the spatial encoder equipped with Multi-Scale Vision Space State Blocks (MS-VSSB) extracts multi-scale appearance features, and the temporal encoder employs Channel-Aware Vision Space State Blocks (CA-VSSB) to capture significant motion patterns. Secondly, a Spatial-Temporal Interaction Module (STIM) is introduced to integrate spatial and temporal information across multiple levels, enabling effective modeling of intrinsic spatial-temporal consistency. Within this module, the Spatial-Temporal Fusion Block (STFB) is proposed to fuse the spatial and temporal features into a unified feature space, and the memory bank is utilized to store spatial-temporal prototypes of normal patterns, restricting the model's ability to represent anomalies. Extensive experiments on three benchmark datasets demonstrate that our STNMamba achieves competitive performance with fewer parameters and lower computational costs than existing methods. Zhangxun Li, Mengyang Zhao 0002, Yang Liu 0246, Jiamu Sheng, Xinhua Zeng, Tian Wang 0002, Kewei Wu, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 8 |
| 2026 | Mask-Aware Kernel Learning for Action RecognitionabstractAction recognition aims to identify an action from video frames. The actions are usually surrounded by irrelevant backgrounds. The action/background information is diverse in different video frames, which hinders learning the implicit action patterns. In this work, we propose a Mask-aware Kernel Model (MKM), which ensures implicit action pattern learning by integrating kernel learning with proper cluster relations. The MKM provides novel cluster-aware kernels to enhance the action representation for frame patches. The MKM is deployed on a temporal Vision Transformer, and introduces a kernel clustering learner, kernel masking filter, and a kernel attention selector.First, to learn temporal features, the temporal Vision Transformer uses temporal correlation to ensure the action features for kernel learning.Second, to analyze the action kernels for frame patches, we design a kernel clustering learner module. This module learns cluster relations with patch- wise convolutions to describe the common action among patches. The cluster relations are learned in each frame, which ensures cluster-aware kernel learning with input frame adaptivity.Third, to analyze the action kernels with spatial adaptivity, we design a kernel masking filter module. This module introduces a location mask by analyzing the region patterns with spatial convolution. The patch-level mask ensures the kernel learning with region-aware selection.Fourth, after learning multiple channel features by convolution with multiple kernels, we design a kernel attention selector module. This module excites kernel-aware features by learning channel- wise attention with channel- wise convolutions, which ensures the kernel learning with channel- wise selection for effective action representation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on Something-Something V1 & V2, Kinetics-400, UAV-human, and Diving 48 datasets. Kewei Wu, Chongjia Zhu, Zhao Xie, Kun Shao, Dan Guo 0001 |
IEEE Trans. Multim. | 1 |
| 2026 | Transition-aware Path and Direction Variation Modeling for Gaze Target Detection in VideoabstractGaze target detection aims to localize a person’s gaze target. During gaze transition in video, the absence of accurate temporal variation modeling (TVM) may lead to errors in gaze target localization. In this work, we propose a Transition-aware Gaze Model (TGM), which focuses on analyzing temporal differences to achieve accurate location variation modeling. The TGM contains four key components: a frame gaze model, and three transition-aware modules (path variation, direction variation, and fusion). First , the frame Transformer extracts gaze location and direction features. Second , to analyze the feature difference among transition frames, we introduce TVM guided by transition-aware loss. TVM analyzes the location features to capture the moving trajectory of targets (defined as path variation ), which facilitates the search for target locations near the path. Third , TVM also analyzes the direction features to capture the transition-aware direction area (defined as direction variation ), which facilitates the search for target locations within this area. Fourth , since gaze directions dynamically adjust to track gaze targets, path variation, and direction variation are inherently aligned with the natural movement of a person’s gaze. Thus, these two variations are fused into a unified transition-aware feature, which helps cover all potential target locations. To search for accurate target locations, we embed this transition-aware feature into frame features with cross-attention, which can enhance gaze target detection in transition frames. Extensive experiments demonstrate that our method achieves state-of-the-art performance on two datasets, namely VideoAttentionTarget and VideoCoAtt. Xingming Yang, Kewei Wu, Zhao Xie, Chongjia Zhu, Dan Guo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Understand and Detect: Multi-step zero-shot detection with image-level specific prompt
Miaotian Guo, Kewei Wu, Zhuqing Jiang, Haiying Wang 0005, Aidong Men |
Knowl. Based Syst. | 2 |
| 2025 | Instructive Probabilistic Transformer for Complex Action RecognitionabstractComplex action recognition aims to identify multiple actions over a long time. Multiple actions may occur at the same time (defined as simultaneous actions), and may occur after each other (defined as each action) Complex action recognition may suffer from two challenges. (1)Temporal repeated bias.The same action may repeat in a temporal duration. In this duration, the prediction may be biased to the majority of actions, which occur repeatedly in the past temporal frames. (2)Epistemic uncertainty of multiple actions.When there are multiple simultaneous actions in one frame, this frame's feature may result in the distribution of multiple actions overlapping each other. Without modeling proper relations between actions, the model may hinder accurately explaining certain categories in multiple actions (defined as the model's epistemic uncertainty). In this work, we propose anInstructive Probabilistic Transformer, which contains a probabilistic temporal memorizer, and a probabilistic prototype Transformer.First, to alleviate temporal repeated bias, we design a probabilistic temporal memory module, which learns probabilistic temporal gates to localize each action. The probabilistic gates instruct the selective memory of each action in long-term frames.Second, we cluster features to capture common action semantics among features (defined as action prototypes). To alleviate the epistemic uncertainty of multiple actions, we design a probabilistic prototype Transformer module. This module learns probabilistic relations depending on each prototype, which can ensure the separation between different prototypes.Third, to ensure the proper probabilistic relations depending on each prototype, we extend action loss with distribution loss to learn uncertainty-aware action loss. In uncertainty-aware action loss, the distribution loss measures the consistency between probabilistic relations and prototype relation distribution. The prediction uncertainty is learned by analyzing the entropy of multiple predictions, and helps to ensure the effect between action loss and distribution loss. Extensive experiments demonstrate that our method achieves state-of-the-art performance on Charades, Breakfast Actions, and MultiTHUMOS. Zhao Xie, Longsheng Lu, Kewei Wu, Zhehan Kan, Xingming Yang, Dan Guo 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Ensemble Prototype Network For Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (TAL) aims to localize the action instances in untrimmed videos using only video-level action labels. Without snippet-level labels, this task should be hard to distinguish all snippets with accurate action/background categories. The main difficulties are the large variations brought by the unconstraint background snippets and multiple subactions in action snippets. The existing prototype model focuses on describing snippets by covering them with clusters (defined as prototypes). In this work, we argue that the clustered prototype covering snippets with simple variations still suffers from the misclassification of the snippets with large variations. We propose an ensemble prototype network (EPNet), which ensembles prototypes learned with consensus-aware clustering. The network stacks a consensus prototype learning (CPL) module and an ensemble snippet weight learning (ESWL) module as one stage and extends one stage to multiple stages in an ensemble learning way. The CPL module learns the consensus matrix by estimating the similarity of clustering labels between two successive clustering generations. The consensus matrix optimizes the clustering to learn consensus prototypes, which can predict the snippets with consensus labels. The ESWL module estimates the weights of the misclassified snippets using the snippet-level loss. The weights update the posterior probabilities of the snippets in the clustering to learn prototypes in the next stage. We use multiple stages to learn multiple prototypes, which can cover the snippets with large variations for accurate snippet classification. Extensive experiments show that our method achieves the state-of-the-art weakly supervised TAL methods on two benchmark datasets, that is, THUMOS'14, ActivityNet v1.2, and ActivityNet v1.3 datasets. Kewei Wu, Zhao Xie, Dan Guo 0001, Zhao Zhang 0001, Richang Hong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Towards Understanding Future: Consistency Guided Probabilistic Modeling for Action AnticipationabstractAction anticipation aims to infer the action in the unobserved segment (future segment) with the observed segment (past segment). Existing methods focus on learning key past semantics to predict the future, but they do not model the temporal continuity between the past and the future. However, past actions are always highly uncertain in anticipating the unobserved future. The absence of temporal continuity smoothing in the video's past-and-future segments may result in an inconsistent anticipation of future action. In this work, we aim to smooth the global semantics changes in the past and future segments. We propose a Consistency-guided Probabilistic Model (CPM), which focuses on learning the globally temporal probabilistic consistency to inhibit the unexpected temporal consistency. The CPM is deployed on the Transformer architecture, which includes three modules of future semantics estimation, global semantics estimation, and global distribution estimation involving the learning of past-to-future semantics, past-and-future semantics, and semantically probabilistic distributions. To achieve the smoothness of temporal continuity, we follow the principle of variational analysis and describe two probabilistic distributions, i.e., a past-aware distribution and a global-aware distribution, which help to estimate the evidence lower bound of future anticipation. In this study, we maximize the evidence lower bound of future semantics by reducing the distribution distance between the above two distributions for model optimization. Extensive experiments demonstrate that the effectiveness of our method and the CPM achieves state-of-the-art performance on Epic-Kitchen100, Epic-Kitchen55, and EGTEA-GAZE. Zhao Xie, Yadong Shi, Kewei Wu, Yaru Cheng, Dan Guo 0001 |
AAAI | 3 |
| 2024 | Active Factor Graph Network for Group Activity RecognitionabstractGroup activity recognition aims to identify a consistent group activity from different actions performed by respective individuals. Most existing methods focus on learning the interaction between each two individuals (i.e., second-order interaction). In this work, we argue that the second-order interactive relation is insufficient to address this task. We propose a third-order active factor graph network, which models the third-order interaction in each pair of three active individuals. At first, to alleviate the noisy individual actions, we select active individuals by measuring each individual's influence. The individuals with the top-k largest influence weights are selected as active individuals. Then, for each three-individuals pair, we build a new factor node and contact the factor node with these individual nodes. In other words, we extend the base second-order interactive graph to a new third-order interactive graph, which is defined as factor graph. Next, we design a two-branch factor graph network, in which one branch is to consider all individuals (denoted as full factor graph) and the other one takes the active individuals into consideration (denoted as active factor graph). We leverage both the active and full factor graphs comprehensively for group activity recognition. Besides, to enforce group consistency, a consistency-aware reasoning module is designed with two penalty terms, which describe the inconsistency between individual actions and group activity respectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance on four benchmark datasets, i.e., Volleyball, Collective Activity, Collective Activity Extended, and SoccerNet-v3 datasets. Visualization results further validate the interpretability of our method. Zhao Xie, Jiao Chang, Kewei Wu, Dan Guo 0001, Richang Hong |
IEEE Trans. Image Process. | 3 |
| 2023 | An Actor-centric Causality Graph for Asynchronous Temporal Inference in Group ActivityabstractThe causality relation modeling remains a challenging task for group activity recognition. The causality relations describe the influence on the centric actor (effect actor) from its correlative actors (cause actors). Most existing graph models focus on learning the actor relation with synchronous temporal features, which is insufficient to deal with the causality relation with asynchronous temporal features. In this paper, we propose an Actor-Centric Causality Graph Model, which learns the asynchronous temporal causality relation with three modules, i.e., an asynchronous temporal causality relation detection module, a causality feature fusion module, and a causality relation graph inference module. First, given a centric actor and its correlative actor, we analyze their influences to detect causality relation. We estimate the self influence of the centric actor with self regression. We estimate the correlative influence from the correlative actor to the centric actor with correlative regression, which uses asynchronous features at different timestamps. Second, we synchronize the two action features by estimating the temporal delay between the cause action and the effect action. The synchronized features are used to enhance the feature of the effect action with a channel-wise fusion. Third, we describe the nodes (actors) with causality features and learn the edges by fusing the causality relation with the appearance relation and distance relation. The causality relation graph inference provides crucial features of effect action, which are complementary to the base model using synchronous relation inference. Experiments show that our method achieves state-of-the-art performance on the Volleyball dataset and Collective Activity dataset. Zhao Xie, Kewei Wu, Jiao Chang |
CVPR | 3 |
| 2023 | Attentive spatial-temporal contrastive learning for self-supervised video representation
Xingming Yang, Sixuan Xiong, Kewei Wu, Dongfeng Shan, Zhao Xie |
Image Vis. Comput. | 3 |
| 2023 | Global Temporal Difference Network for Action RecognitionabstractTemporal modeling still remains as a challenge for action recognition. Most existing temporal models focus on learning local variation between neighbor frames. There exists obvious deviations between local and global variations, such as subtle and notable motion variations. In this paper, we propose a global temporal difference module for action recognition, which consists of two sub-modules,i.e., a global aggregation module and a global difference module. These two sub-modules cooperate following the idea of using prior knowledge from the global view (i.e., global motion variation) to guide local learning at each moment. In the global aggregation module, the global prior knowledge is learned by aggregating the visual feature sequence of video into a global vector. In the global difference module, we prepare the difference vector sequence of video by subtracting each local vector from the global vector. Our method performs as a contextual guidance with a global view. The sequential dependency between these difference vectors is exploited with a channel-wise self-attention operation. Finally, the difference vectors at each timestamp are further used to enhance the semantics of the original local features. The enhanced features endow the action recognition has less deviation to understand the variation in the video globally. We instantiate the global temporal difference module into the ResNet block to form a global temporal difference network (GTDNet). Exhaustive experiments are conducted and our method achieves competitive performance at small FLOPs on Something-Something V1 & V2 and Kinetics-400. Zhao Xie, Jiansong Chen, Kewei Wu, Dan Guo 0001, Richang Hong |
IEEE Trans. Multim. | 3 |
| 2021 | Distilling Dynamic Spatial Relation Network for Human Pose Estimation
Kewei Wu, Zhao Xie, Dan Guo 0001 |
BMVC | 1 |
| 2021 | DDFPN: Context enhanced network for object detection
Kewei Wu, Zhao Xie, Dan Guo 0001 |
Future Gener. Comput. Syst. | 1 |
| 2021 | Deep social force network for anomaly event detectionabstractAbstract Anomaly event detection is vital in surveillance video analysis. However, how to learn the discriminative motion in the crowd scene is still not tackled. Here, a deep social force network by exploiting both social force extracting and deep motion coding is proposed. Given a grid of particles with velocity provided by the optical flow, the interaction force in the crowd scene is investigated and a social force module is embedded in a deep network. A deep motion convolution was further designed with a 3D (DMC‐3D) module. The DMC‐3D not only eliminates the noise motion in the crowd scene with a spatial encoder–decoder but also learns the 3D feature with a spatio‐temporal encoder. The deep social force coding is modelled with multiple features, in which each feature can describe specific anomaly motion. The experiments on UCF‐Crime and ShanghaiTech datasets demonstrate that our method can predict the temporal localization of anomaly events and outperform the state‐of‐the‐art methods. Xingming Yang, Kewei Wu, Zhao Xie, Jinkui Hou |
IET Image Process. | 3 |
| 2021 | Learning continuous temporal embedding of videos using pattern theory
Zhao Xie, Kewei Wu, Xingming Yang, Jinkui Hou |
Pattern Recognit. Lett. | 2 |
| 2021 | Faceted Text Segmentation via Multitask LearningabstractText segmentation is a fundamental step in natural language processing (NLP) and information retrieval (IR) tasks. Most existing approaches do not explicitly take into account the facet information of documents for segmentation. Text segmentation and facet annotation are often addressed as separate problems, but they operate in a common input space. This article proposes FTS, which is a novel model for faceted text segmentation via multitask learning (MTL). FTS models faceted text segmentation as an MTL problem with text segmentation and facet annotation. This model employs the bidirectional long short-term memory (Bi-LSTM) network to learn the feature representation of sentences within a document. The feature representation is shared and adjusted with common parameters by MTL, which can help an optimization model to learn a better-shared and robust feature representation from text segmentation to facet annotation. Moreover, the text segmentation is modeled as a sequence tagging task using LSTM with a conditional random fields (CRFs) classification layer. Extensive experiments are conducted on five data sets from five domains: data structure, data mining, computer network, solid mechanics, and crystallography. The results indicate that the FTS model outperforms several highly cited and state-of-the-art approaches related to text segmentation and facet annotation. Bei Wu 0003, Bifan Wei, Jun Liu 0002, Kewei Wu, Meng Wang 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | A deep generative directed network for scene depth ordering
Kewei Wu, Yongxuan Sun, Zhao Xie |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Jointly social grouping and identification in visual dynamics with causality-induced hierarchical Bayesian model
Zhao Xie, Tianfu Wu 0001, Xingming Yang, Kewei Wu |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Monocular relative depth reordering by propagating confidence of local and global cues
Kewei Wu |
Multim. Tools Appl. | 1 |
| 2008 | FCM in novel application of science and technology progress monitor systemabstractThis paper focuses on the issues about the complex relations in large-scale FCM, and then proposes a promising method for weight global optimization with local inference to analyze and predict indexes in Anhui sci-tech progress monitor system. Firstly, a new concept, unbalanced degree, is introduced for standard evaluation in FCM model to modify the weight assessment factors and result in the satisfied convergence rate. Secondly, relations between unbalanced degree and convergence error are also presented for further analysis with training error and guarantee on perfect condition in model. Thirdly, local inference in FCM is discussed to enhance prediction accuracy rate. Finally, experimental result reveals successful application of FCM in large-scale complex sci-tech systems. Kewei Wu, Zhao Xie, Jun Gao 0006, Wengang Feng |
FUZZ-IEEE | 1 |