EDBT 2026 Demo / reviewers in the wild / expert
Zhao Xie
dblp:32/6328
· DBLP profile ↗
25ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0001-9834-4730ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Confidence-Aware Prototypes for Weakly-Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection aims to identify abnormal snippets in untrimmed videos. Existing methods learn prototypes to describe global representation of snippet distributions. But, in weakly-labeled videos, the normal snippets in abnormal video may take high-uncertainty labels for distribution modeling. Without confidence-aware modeling, abnormal/normal prototype distributions may overlap with each other, leading to inaccurate predictions. In this work, we propose the Unified Confident Prototype (UCP) model, which contains a feature extractor, a confidence-aware prototype learner, and a local-global prototype unifier. The prototype learning is designed to ensure proper separability, stability, and representation.First, after learning the weight of each snippet’s loss, snippets with high-uncertainty labels may take small weights. These snippets tend to lie in the overlap between abnormal/normal distributions, hindering their separation. We design uncertainty-aware sampling, which removes high-uncertainty snippets in the small-weight snippets to ensure separable prototype learning.Second, snippets with high-uncertainty labels tend to be far from the prototype center, thus falling in the low-confidence region. These snippets may enlarge the distribution’s variation, resulting in unstable prototype learning. We design confidence-aware sampling, which removes low-confidence snippets to ensure stable prototype learning.Third, after assigning pseudo labels to prototypes, we measure the prototype representation with the distribution’s purity. We design prototype distribution purification, which penalizes normal snippets in the abnormal-majority distribution with purity loss to ensure representative prototype learning.Fourth, beyond prototype learning, prototypes can be enhanced by local/global temporal semantics. We further introduce the local-global prototype unifier to learn the relations across local-global durations, thereby enhancing the semantics for anomaly detection. For weakly-supervised anomaly detection, experiments demonstrate that our method achieves state-of-the-art performance on the UCF-Crime, ShanghaiTech, and XD-Violence datasets. Moreover, to further verify the generality of our method, we further conduct experiments on THUMOS’14 for weakly-supervised temporal action localization. Zhao Xie, Jinkang Luo, Kewei Wu, Zhehan Kan, Dan Guo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Mask-Aware Kernel Learning for Action RecognitionabstractAction recognition aims to identify an action from video frames. The actions are usually surrounded by irrelevant backgrounds. The action/background information is diverse in different video frames, which hinders learning the implicit action patterns. In this work, we propose a Mask-aware Kernel Model (MKM), which ensures implicit action pattern learning by integrating kernel learning with proper cluster relations. The MKM provides novel cluster-aware kernels to enhance the action representation for frame patches. The MKM is deployed on a temporal Vision Transformer, and introduces a kernel clustering learner, kernel masking filter, and a kernel attention selector.First, to learn temporal features, the temporal Vision Transformer uses temporal correlation to ensure the action features for kernel learning.Second, to analyze the action kernels for frame patches, we design a kernel clustering learner module. This module learns cluster relations with patch- wise convolutions to describe the common action among patches. The cluster relations are learned in each frame, which ensures cluster-aware kernel learning with input frame adaptivity.Third, to analyze the action kernels with spatial adaptivity, we design a kernel masking filter module. This module introduces a location mask by analyzing the region patterns with spatial convolution. The patch-level mask ensures the kernel learning with region-aware selection.Fourth, after learning multiple channel features by convolution with multiple kernels, we design a kernel attention selector module. This module excites kernel-aware features by learning channel- wise attention with channel- wise convolutions, which ensures the kernel learning with channel- wise selection for effective action representation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on Something-Something V1 & V2, Kinetics-400, UAV-human, and Diving 48 datasets. Kewei Wu, Chongjia Zhu, Zhao Xie, Kun Shao, Dan Guo 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Transition-aware Path and Direction Variation Modeling for Gaze Target Detection in VideoabstractGaze target detection aims to localize a person’s gaze target. During gaze transition in video, the absence of accurate temporal variation modeling (TVM) may lead to errors in gaze target localization. In this work, we propose a Transition-aware Gaze Model (TGM), which focuses on analyzing temporal differences to achieve accurate location variation modeling. The TGM contains four key components: a frame gaze model, and three transition-aware modules (path variation, direction variation, and fusion). First , the frame Transformer extracts gaze location and direction features. Second , to analyze the feature difference among transition frames, we introduce TVM guided by transition-aware loss. TVM analyzes the location features to capture the moving trajectory of targets (defined as path variation ), which facilitates the search for target locations near the path. Third , TVM also analyzes the direction features to capture the transition-aware direction area (defined as direction variation ), which facilitates the search for target locations within this area. Fourth , since gaze directions dynamically adjust to track gaze targets, path variation, and direction variation are inherently aligned with the natural movement of a person’s gaze. Thus, these two variations are fused into a unified transition-aware feature, which helps cover all potential target locations. To search for accurate target locations, we embed this transition-aware feature into frame features with cross-attention, which can enhance gaze target detection in transition frames. Extensive experiments demonstrate that our method achieves state-of-the-art performance on two datasets, namely VideoAttentionTarget and VideoCoAtt. Xingming Yang, Kewei Wu, Zhao Xie, Chongjia Zhu, Dan Guo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Instructive Probabilistic Transformer for Complex Action RecognitionabstractComplex action recognition aims to identify multiple actions over a long time. Multiple actions may occur at the same time (defined as simultaneous actions), and may occur after each other (defined as each action) Complex action recognition may suffer from two challenges. (1)Temporal repeated bias.The same action may repeat in a temporal duration. In this duration, the prediction may be biased to the majority of actions, which occur repeatedly in the past temporal frames. (2)Epistemic uncertainty of multiple actions.When there are multiple simultaneous actions in one frame, this frame's feature may result in the distribution of multiple actions overlapping each other. Without modeling proper relations between actions, the model may hinder accurately explaining certain categories in multiple actions (defined as the model's epistemic uncertainty). In this work, we propose anInstructive Probabilistic Transformer, which contains a probabilistic temporal memorizer, and a probabilistic prototype Transformer.First, to alleviate temporal repeated bias, we design a probabilistic temporal memory module, which learns probabilistic temporal gates to localize each action. The probabilistic gates instruct the selective memory of each action in long-term frames.Second, we cluster features to capture common action semantics among features (defined as action prototypes). To alleviate the epistemic uncertainty of multiple actions, we design a probabilistic prototype Transformer module. This module learns probabilistic relations depending on each prototype, which can ensure the separation between different prototypes.Third, to ensure the proper probabilistic relations depending on each prototype, we extend action loss with distribution loss to learn uncertainty-aware action loss. In uncertainty-aware action loss, the distribution loss measures the consistency between probabilistic relations and prototype relation distribution. The prediction uncertainty is learned by analyzing the entropy of multiple predictions, and helps to ensure the effect between action loss and distribution loss. Extensive experiments demonstrate that our method achieves state-of-the-art performance on Charades, Breakfast Actions, and MultiTHUMOS. Zhao Xie, Longsheng Lu, Kewei Wu, Zhehan Kan, Xingming Yang, Dan Guo 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Ensemble Prototype Network For Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (TAL) aims to localize the action instances in untrimmed videos using only video-level action labels. Without snippet-level labels, this task should be hard to distinguish all snippets with accurate action/background categories. The main difficulties are the large variations brought by the unconstraint background snippets and multiple subactions in action snippets. The existing prototype model focuses on describing snippets by covering them with clusters (defined as prototypes). In this work, we argue that the clustered prototype covering snippets with simple variations still suffers from the misclassification of the snippets with large variations. We propose an ensemble prototype network (EPNet), which ensembles prototypes learned with consensus-aware clustering. The network stacks a consensus prototype learning (CPL) module and an ensemble snippet weight learning (ESWL) module as one stage and extends one stage to multiple stages in an ensemble learning way. The CPL module learns the consensus matrix by estimating the similarity of clustering labels between two successive clustering generations. The consensus matrix optimizes the clustering to learn consensus prototypes, which can predict the snippets with consensus labels. The ESWL module estimates the weights of the misclassified snippets using the snippet-level loss. The weights update the posterior probabilities of the snippets in the clustering to learn prototypes in the next stage. We use multiple stages to learn multiple prototypes, which can cover the snippets with large variations for accurate snippet classification. Extensive experiments show that our method achieves the state-of-the-art weakly supervised TAL methods on two benchmark datasets, that is, THUMOS'14, ActivityNet v1.2, and ActivityNet v1.3 datasets. Kewei Wu, Zhao Xie, Dan Guo 0001, Zhao Zhang 0001, Richang Hong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Towards Understanding Future: Consistency Guided Probabilistic Modeling for Action AnticipationabstractAction anticipation aims to infer the action in the unobserved segment (future segment) with the observed segment (past segment). Existing methods focus on learning key past semantics to predict the future, but they do not model the temporal continuity between the past and the future. However, past actions are always highly uncertain in anticipating the unobserved future. The absence of temporal continuity smoothing in the video's past-and-future segments may result in an inconsistent anticipation of future action. In this work, we aim to smooth the global semantics changes in the past and future segments. We propose a Consistency-guided Probabilistic Model (CPM), which focuses on learning the globally temporal probabilistic consistency to inhibit the unexpected temporal consistency. The CPM is deployed on the Transformer architecture, which includes three modules of future semantics estimation, global semantics estimation, and global distribution estimation involving the learning of past-to-future semantics, past-and-future semantics, and semantically probabilistic distributions. To achieve the smoothness of temporal continuity, we follow the principle of variational analysis and describe two probabilistic distributions, i.e., a past-aware distribution and a global-aware distribution, which help to estimate the evidence lower bound of future anticipation. In this study, we maximize the evidence lower bound of future semantics by reducing the distribution distance between the above two distributions for model optimization. Extensive experiments demonstrate that the effectiveness of our method and the CPM achieves state-of-the-art performance on Epic-Kitchen100, Epic-Kitchen55, and EGTEA-GAZE. Zhao Xie, Yadong Shi, Kewei Wu, Yaru Cheng, Dan Guo 0001 |
AAAI | 1 |
| 2024 | Active Factor Graph Network for Group Activity RecognitionabstractGroup activity recognition aims to identify a consistent group activity from different actions performed by respective individuals. Most existing methods focus on learning the interaction between each two individuals (i.e., second-order interaction). In this work, we argue that the second-order interactive relation is insufficient to address this task. We propose a third-order active factor graph network, which models the third-order interaction in each pair of three active individuals. At first, to alleviate the noisy individual actions, we select active individuals by measuring each individual's influence. The individuals with the top-k largest influence weights are selected as active individuals. Then, for each three-individuals pair, we build a new factor node and contact the factor node with these individual nodes. In other words, we extend the base second-order interactive graph to a new third-order interactive graph, which is defined as factor graph. Next, we design a two-branch factor graph network, in which one branch is to consider all individuals (denoted as full factor graph) and the other one takes the active individuals into consideration (denoted as active factor graph). We leverage both the active and full factor graphs comprehensively for group activity recognition. Besides, to enforce group consistency, a consistency-aware reasoning module is designed with two penalty terms, which describe the inconsistency between individual actions and group activity respectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance on four benchmark datasets, i.e., Volleyball, Collective Activity, Collective Activity Extended, and SoccerNet-v3 datasets. Visualization results further validate the interpretability of our method. Zhao Xie, Jiao Chang, Kewei Wu, Dan Guo 0001, Richang Hong |
IEEE Trans. Image Process. | 1 |
| 2023 | An Actor-centric Causality Graph for Asynchronous Temporal Inference in Group ActivityabstractThe causality relation modeling remains a challenging task for group activity recognition. The causality relations describe the influence on the centric actor (effect actor) from its correlative actors (cause actors). Most existing graph models focus on learning the actor relation with synchronous temporal features, which is insufficient to deal with the causality relation with asynchronous temporal features. In this paper, we propose an Actor-Centric Causality Graph Model, which learns the asynchronous temporal causality relation with three modules, i.e., an asynchronous temporal causality relation detection module, a causality feature fusion module, and a causality relation graph inference module. First, given a centric actor and its correlative actor, we analyze their influences to detect causality relation. We estimate the self influence of the centric actor with self regression. We estimate the correlative influence from the correlative actor to the centric actor with correlative regression, which uses asynchronous features at different timestamps. Second, we synchronize the two action features by estimating the temporal delay between the cause action and the effect action. The synchronized features are used to enhance the feature of the effect action with a channel-wise fusion. Third, we describe the nodes (actors) with causality features and learn the edges by fusing the causality relation with the appearance relation and distance relation. The causality relation graph inference provides crucial features of effect action, which are complementary to the base model using synchronous relation inference. Experiments show that our method achieves state-of-the-art performance on the Volleyball dataset and Collective Activity dataset. Zhao Xie, Kewei Wu, Jiao Chang |
CVPR | 1 |
| 2023 | Attentive spatial-temporal contrastive learning for self-supervised video representation
Xingming Yang, Sixuan Xiong, Kewei Wu, Dongfeng Shan, Zhao Xie |
Image Vis. Comput. | 5 |
| 2023 | Global Temporal Difference Network for Action RecognitionabstractTemporal modeling still remains as a challenge for action recognition. Most existing temporal models focus on learning local variation between neighbor frames. There exists obvious deviations between local and global variations, such as subtle and notable motion variations. In this paper, we propose a global temporal difference module for action recognition, which consists of two sub-modules,i.e., a global aggregation module and a global difference module. These two sub-modules cooperate following the idea of using prior knowledge from the global view (i.e., global motion variation) to guide local learning at each moment. In the global aggregation module, the global prior knowledge is learned by aggregating the visual feature sequence of video into a global vector. In the global difference module, we prepare the difference vector sequence of video by subtracting each local vector from the global vector. Our method performs as a contextual guidance with a global view. The sequential dependency between these difference vectors is exploited with a channel-wise self-attention operation. Finally, the difference vectors at each timestamp are further used to enhance the semantics of the original local features. The enhanced features endow the action recognition has less deviation to understand the variation in the video globally. We instantiate the global temporal difference module into the ResNet block to form a global temporal difference network (GTDNet). Exhaustive experiments are conducted and our method achieves competitive performance at small FLOPs on Something-Something V1 & V2 and Kinetics-400. Zhao Xie, Jiansong Chen, Kewei Wu, Dan Guo 0001, Richang Hong |
IEEE Trans. Multim. | 1 |
| 2021 | Distilling Dynamic Spatial Relation Network for Human Pose Estimation
Kewei Wu, Zhao Xie, Dan Guo 0001 |
BMVC | 3 |
| 2021 | DDFPN: Context enhanced network for object detection
Kewei Wu, Zhao Xie, Dan Guo 0001 |
Future Gener. Comput. Syst. | 3 |
| 2021 | Deep social force network for anomaly event detectionabstractAbstract Anomaly event detection is vital in surveillance video analysis. However, how to learn the discriminative motion in the crowd scene is still not tackled. Here, a deep social force network by exploiting both social force extracting and deep motion coding is proposed. Given a grid of particles with velocity provided by the optical flow, the interaction force in the crowd scene is investigated and a social force module is embedded in a deep network. A deep motion convolution was further designed with a 3D (DMC‐3D) module. The DMC‐3D not only eliminates the noise motion in the crowd scene with a spatial encoder–decoder but also learns the 3D feature with a spatio‐temporal encoder. The deep social force coding is modelled with multiple features, in which each feature can describe specific anomaly motion. The experiments on UCF‐Crime and ShanghaiTech datasets demonstrate that our method can predict the temporal localization of anomaly events and outperform the state‐of‐the‐art methods. Xingming Yang, Kewei Wu, Zhao Xie, Jinkui Hou |
IET Image Process. | 4 |
| 2021 | Learning continuous temporal embedding of videos using pattern theory
Zhao Xie, Kewei Wu, Xingming Yang, Jinkui Hou |
Pattern Recognit. Lett. | 1 |
| 2019 | A deep generative directed network for scene depth ordering
Kewei Wu, Yongxuan Sun, Zhao Xie |
J. Vis. Commun. Image Represent. | 6 |
| 2019 | Jointly social grouping and identification in visual dynamics with causality-induced hierarchical Bayesian model
Zhao Xie, Tianfu Wu 0001, Xingming Yang, Kewei Wu |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Camera-Assisted Video Saliency Prediction and Its ApplicationsabstractVideo saliency prediction is an indispensable yet challenging technique which can facilitate various applications, such as video surveillance, autonomous driving, and realistic rendering. Based on the popularity of embedded cameras, we in this paper predict region-level saliency from videos by leveraging human gaze locations recorded using a camera, (e.g., those equipped on an iMAC and laptop PC). Our proposed camera-assisted mechanism improves saliency prediction by discovering human attended regions inside a video clip. It is orthogonal to the current saliency models, i.e., any existing video/image saliency model can be boosted by our mechanism. First of all, the spatial-and temporal-level visual features are exploited collaboratively for calculating an initial saliency map. We notice that the current saliency models are not sufficiently adaptable to the variations in lighting, different view angles, and complicated backgrounds. Therefore, assisted by a camera tracking human gaze movements, a non-negative matrix factorization algorithm is designed to accurately localize the semantically/visually salient video regions perceived by humans. Finally, the learned human gaze locations as well as the initial saliency map are integrated to optimize video saliency calculation. Empirical results thoroughly demonstrated that: 1) our approach achieves the state-of-the-art video saliency prediction accuracy by outperforming 11 mainstream algorithms considerably and 2) our method can conveniently and successfully enhance video retargeting, quality estimation, and summarization. Xiao Sun 0003, Yuxing Hu, Ping Li 0006, Zhao Xie, Zhenguang Liu |
IEEE Trans. Cybern. | 6 |
| 2017 | Learning universal multiview dictionary for human action recognition
Zhiyong Wang 0001, Zhao Xie, Jun Gao 0006, David Dagan Feng |
Pattern Recognit. | 3 |
| 2016 | Discriminative sequential association latent dirichlet allocation for visual recognition
Zhao Xie, Jun Gao 0006 |
Pattern Anal. Appl. | 2 |
| 2015 | Robust tracking with per-exemplar support vector machineabstractThe authors extend exemplar representation to the field of tracking and propose a robust tracking algorithm with per‐exemplar support vector machine (SVM) classifiers. First, the authors train the simple yet effective exemplar SVM classifier using the target object as the single positive and mining its surroundings as hard negatives. Second, the authors propose an online ensemble tracker, which integrates the useful ‘key historical templates’ of the target to refine the current template, leading to better discriminative power of tracker and effectively decreasing the risk of drift. Experiments on challenging sequences demonstrate that the tracker performs well in accuracy and robustness, especially under the sequences with strong illumination variation and scale variation, as well as pose change and partial occlusion in the long‐time sequence. Rongmei Shi, Jun Zhang 0017, Zhao Xie, Jun Gao 0006, Xinxiang Zheng |
IET Comput. Vis. | 3 |
| 2015 | Geometric structure-constraint tracking with confident parts
Zhao Xie, Yongxuan Sun |
Signal Process. Image Commun. | 1 |
| 2012 | Cooperative Sparse Representation in Two Opposite Directions for Semi-Supervised Image AnnotationabstractRecent studies have shown that sparse representation (SR) can deal well with many computer vision problems, and its kernel version has powerful classification capability. In this paper, we address the application of a cooperative SR in semi-supervised image annotation which can increase the amount of labeled images for further use in training image classifiers. Given a set of labeled (training) images and a set of unlabeled (test) images, the usual SR method, which we call forward SR, is used to represent each unlabeled image with several labeled ones, and then to annotate the unlabeled image according to the annotations of these labeled ones. However, to the best of our knowledge, the SR method in an opposite direction, that we call backward SR to represent each labeled image with several unlabeled images and then to annotate any unlabeled image according to the annotations of the labeled images which the unlabeled image is selected by the backward SR to represent, has not been addressed so far. In this paper, we explore how much the backward SR can contribute to image annotation, and be complementary to the forward SR. The co-training, which has been proved to be a semi-supervised method improving each other only if two classifiers are relatively independent, is then adopted to testify this complementary nature between two SRs in opposite directions. Finally, the co-training of two SRs in kernel space builds a cooperative kernel sparse representation (Co-KSR) method for image annotation. Experimental results and analyses show that two KSRs in opposite directions are complementary, and Co-KSR improves considerably over either of them with an image annotation performance better than other state-of-the-art semi-supervised classifiers such as transductive support vector machine, local and global consistency, and Gaussian fields and harmonic functions. Comparative experiments with a nonsparse solution are also performed to show that the sparsity plays an important role in the cooperation of image representations in two opposite directions. This paper extends the application of SR in image annotation and retrieval. Zhong-Qiu Zhao, Hervé Glotin, Zhao Xie, Jun Gao 0006, Xindong Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2009 | Regional category parsing in undirected graphical models
Zhao Xie, Jun Gao 0006, Xindong Wu 0001 |
Pattern Recognit. Lett. | 1 |
| 2008 | FCM in novel application of science and technology progress monitor systemabstractThis paper focuses on the issues about the complex relations in large-scale FCM, and then proposes a promising method for weight global optimization with local inference to analyze and predict indexes in Anhui sci-tech progress monitor system. Firstly, a new concept, unbalanced degree, is introduced for standard evaluation in FCM model to modify the weight assessment factors and result in the satisfied convergence rate. Secondly, relations between unbalanced degree and convergence error are also presented for further analysis with training error and guarantee on perfect condition in model. Thirdly, local inference in FCM is discussed to enhance prediction accuracy rate. Finally, experimental result reveals successful application of FCM in large-scale complex sci-tech systems. Kewei Wu, Zhao Xie, Jun Gao 0006, Wengang Feng |
FUZZ-IEEE | 2 |
| 2007 | Generic object recognition with regional statistical models and layer joint boosting
Jun Gao 0006, Zhao Xie, Xindong Wu 0001 |
Pattern Recognit. Lett. | 2 |