Aichun Zhu

dblp:154/1314 · DBLP profile ↗
← Back
43ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0001-6972-5534ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Agent Consultation and Uncertainty-guided Voting for text-to-image person retrieval
Mingcheng Ni, Zijie Wang 0003, Aichun Zhu, Jingyi Xue, Guannan Dong, Yong Cheng 0001
Eng. Appl. Artif. Intell.3
2026 Integrating Historical Rules and Curriculum-Enhanced Embeddings for Temporal Knowledge Graph Forecasting
abstract
Temporal knowledge graph forecasting (TKGF) has become a crucial tool for predicting future facts based on temporal knowledge graphs. Traditional methods face inherent limitations: rule-based reasoning approaches are restricted to exploring events that are related to historical rules and fail to fully utilize global graph information. In contrast, although embedding-based methods have the ability to capture global graph information through vector representations, they struggle to effectively incorporate historical rules. To address these issues, this article proposes a novel framework that combines historical rule-based reasoning with embedding-based prediction enhanced by curriculum learning (CL) to predict unknown events. The framework consists of two complementary modes: the rule mode, which extracts temporal rules through time-filtered weighted sampling and applies them for prediction, and the embedding mode, which progressively represents entities, relations, and time as vectors from simple to complex through CL and generates predictions based on these embedding vectors. The framework effectively integrates rules with the entire graph information by combining the predictions from two models, resulting in a more comprehensive prediction. Extensive experiments on multiple public datasets demonstrate that the proposed method outperforms state-of-the-art approaches in terms of performance, validating the effectiveness of its hybrid approach, and highlighting the significant potential of CL in TKGF.
Weizhou Wang, Aichun Zhu, Tianyuan Hu, Shiping Wang, Yuanfei Dai
IEEE Trans. Comput. Soc. Syst.2
2026 Taking Astray Domain Back Home for Single-Source Domain Generalizable Text-to-Image Person Retrieval
abstract
Given a query sentence, text-to-image person retrieval aims to identify matched pedestrian images from a large gallery. Most of the existing methods are designed for the unified domain setting, which is operated under the assumption that the training and test data are drawn from the same distribution. However, this assumption is difficult to guarantee in real application scenes, as data is often collected from various surveillance scenarios. To this end, in this paper, we introduce the concept of single-source domain generalization into the context of text-to-image person retrieval and propose a novel task called single-source domain generalizable text-to-image person retrieval (SSDG-TIPR). This task is applicable in real-world scenarios but poses significant challenges due to the limitation of accessible training data. Intuitively, a trained model is the most familiar with the domain on which it was trained, that is, the source domain. Therefore, to handle this SSDG-TIPR task, we propose a new method to infinitely close astray features from unseen target domains to the source domain, namely, to take it home (TIME), allowing the model to handle the features in a familiar manner. The proposed TIME method comprises three main modules: the Domain Astray Leading (DAL) module, the Domain Invariant Feature Extract (DIFE) module and the Domain Home Taking (DoT) module. We evaluated TIME on 3 benchmark datasets, namely CUHK-PEDES, ICFG-PEDES and RSTPReid, and demonstrated its superior performance on 10 SSDG-TIPR sub-tasks as well as on 3 conventional TIPR sub-tasks, establishing a new state-of-the-art in both settings.
Guan-Nan Dong, Zijie Wang 0003, Aichun Zhu, Yuanfei Dai, Tian Wang 0002, Hichem Snoussi
IEEE Trans. Image Process.4
2025 Unveiling Local Well-posedness Influence for Cross-modal Person Re-Identification
abstract
The existing cross-modal retrieval methods trend toward the conventional multi-modal alignment while ignoring the localization bias caused by visual hallucination, including color pollution and appearance-like occlusion due to uncontrollable factors such as weather, illumination, and occlusion. This feature blinding misleads the model to lock in the pseudo-real position and further leads to local unmatched. To this end, we discuss cross-modal local alignment well-posedness by making a phased local modal-masking to calibrate the undisturbed actual local alignment from entity, attribute, and appearance. Specifically, we introduce a mask-based local well-posedness modeling (MLWM) strategy, including text-based entity masking (TEM), text-based attribute-specific masking (TAM), and image-based appearance masking (IAM) to phased collaboratively consider image prompting-based text entities, image prompting-based text attributes, and text prompting-based appearance inference contrast, respectively. Finally, we dynamically optimize the weights of positively correlated image-text pairs by comparing the similarity between original and reconstructed features. Experimental results demonstrate that our method is effective on three public datasets.
Guannan Dong, Aichun Zhu, Mingcheng Ni, Yifeng Li 0002
ICASSP3
2025 Dynamic, Multi-scale, and Noise-Aware Modeling for Skeleton Action Prediction
Ran Cui, Aichun Zhu
KSEM (3)2
2025 Balance Orthogonal Projection for Prompt in Continual Learning
Junjian Ren, Tian Wang 0002, Aichun Zhu, Chuanyun Wang, Nadia Bali, Hichem Snoussi
PRCV (2)3
2025 Style-Texture Collaborative Learning for Face Forgery Detection
Pengcheng Jia, Guannan Dong, Shubo Wang, Aichun Zhu
PRCV (15)5
2025 Understanding the Dimensional Need of Noncontrastive Learning
abstract
Noncontrastive self-supervised learning methods offer an effective alternative to contrastive approaches by avoiding the need for negative samples to avoid representation collapse. Noncontrastive learning methods explicitly or implicitly optimize the representation space, yet they often require large representation dimensions, leading to dimensional inefficiency. To provide negative samples, contrastive learning methods often require large batch sizes, thus regarded as sample inefficient, while noncontrastive learning methods require large representation dimensions, thus regarded as dimension inefficient. Although we have some understanding of the noncontrastive learning method, theoretical analysis of such phenomenon still remains largely unexplored. We present a theoretical analysis of the dimensional need for noncontrastive learning. We investigate the transfer between upstream representation learning and downstream tasks' performance, demonstrating how noncontrastive methods implicitly increase interclass distances within the representation space and how the distance affects the model performance of evaluation performance. We prove that the performance of noncontrastive methods is affected by the output dimension and the number of latent classes, and illustrate why performance degrades significantly when the output dimension is substantially smaller than the number of latent classes. We demonstrate our findings through experiments on image classification experiments, and enrich the verification in audio, graph and text modalities. We also perform empirical evaluation for image models on extensive detection and segmentation tasks beyond classification that show satisfactory correspondence to our theorem.
Zhexiao Cao, Lei Huang 0015, Tian Wang 0002, Yinquan Wang, Jingang Shi, Aichun Zhu, Tianyun Shi, Hichem Snoussi
IEEE Trans. Cybern.6
2025 Improving Text-Based Person Retrieval by Excavating All-Round Information Beyond Color
abstract
Text-based person retrieval is the process of searching a massive visual resource library for images of a particular pedestrian, based on a textual query. Existing approaches often suffer from a problem of color (CLR) over-reliance, which can result in a suboptimal person retrieval performance by distracting the model from other important visual cues such as texture and structure information. To handle this problem, we propose a novel framework to Excavate All-round Information Beyond Color for the task of text-based person retrieval, which is therefore termed EAIBC. The EAIBC architecture includes four branches, namely an RGB branch, a grayscale (GRS) branch, a high-frequency (HFQ) branch, and a CLR branch. Furthermore, we introduce a mutual learning (ML) mechanism to facilitate communication and learning among the branches, enabling them to take full advantage of all-round information in an effective and balanced manner. We evaluate the proposed method on three benchmark datasets, including CUHK-PEDES, ICFG-PEDES, and RSTPReid. The experimental results demonstrate that EAIBC significantly outperforms existing methods and achieves state-of-the-art (SOTA) performance in supervised, weakly supervised, and cross-domain settings.
Aichun Zhu, Zijie Wang 0003, Jingyi Xue, Xili Wan, Jing Jin 0002, Tian Wang 0002, Hichem Snoussi
IEEE Trans. Neural Networks Learn. Syst.1
2024 TVPR: Text-to-Video Person Retrieval and a New Benchmark
abstract
Most existing methods for text-based person retrieval focus on text-to-image person retrieval. Nevertheless, due to the lack of dynamic information provided by isolated frames, the performance is hampered when the person is obscured or variable motion details are missed in isolated frames. To overcome this, we propose a novel Text-to-Video Person Retrieval (TVPR) task. Since there is no dataset or benchmark that describes person videos with natural language, we construct a large-scale cross-modal person video dataset containing detailed natural language annotations, termed as Text-to-Video Person Reidentification (TVPReid) dataset. In this paper, we introduce a Multielement Feature Guided Fragments Learning (MFGF) strategy, which leverages the cross-modal text-video representations to provide strong text-visual and text-motion matching information to tackle uncertain occlusion conflicting and variable motion details. Specifically, we establish two potential cross-modal spaces for text and video feature collaborative learning to progressively reduce the semantic difference between text and video. To evaluate the effectiveness of the proposed MFGF, extensive experiments have been conducted on TVPReid dataset. To the best of our knowledge, MFGF is the first successful attempt to use video for text-based person retrieval task and has achieved state-of-the-art performance on TVPReid dataset. The TVPReid dataset will be publicly available to benefit future research.
Xu Zhang 0075, Fan Ni, Guannan Dong, Aichun Zhu, Mingcheng Ni, Hui Liu 0026
ACM Multimedia4
2024 Hierarchical Discrepancy-Aware Interaction Network for Face Forgery Detection
Pengcheng Jia, Guannan Dong, Aichun Zhu
PRCV (15)5
2024 EESSO: Exploiting Extreme and Smooth Signals via Omni-frequency learning for Text-based Person Retrieval
Jingyi Xue, Zijie Wang 0003, Guannan Dong, Aichun Zhu
Image Vis. Comput.4
2024 Potential source-information dominated learning for composed cross-modal person re-identification
Xiangyun Zhang, Guannan Dong, Kunyu Wu, Yong Cheng 0001, Aichun Zhu
Multim. Tools Appl.6
2024 A flow-based multi-scale learning network for single image stochastic super-resolution
Qianyu Wu, Zhongqian Hu, Aichun Zhu, Jiaxin Zou, Yan Xi, Yang Chen 0008
Signal Process. Image Commun.3
2023 CT-Mixer: Exploiting Multiscale Design for Local-Global Representations Learning
Xili Wan, Xinjie Guan, Aichun Zhu
ICA3PP (1)4
2023 C2SFormer: Rethinking the Local-Global Design for Efficient Visual Recognition Model
abstract
Vision Transformers are born with the property of data-dependent and long-range dependencies, accomplishing a number of astonishing results against their contemporary competitor CNNs. To alleviate the excessive computational burden, previous methods apply the local operation (e.g., convolution, local attention) in the high-resolution stages. Although these designs are efficient for local relations learning, especially for the high redundancy stages, they inevitably lead to the losses of non-locality and are constrained by the limited receptive field. In this paper, we present an effective hybrid-style vision backbone that is explicitly built with dynamic convolution and self-attention to respectively undertake both local and global interaction, dubbed C2SFormer. We adopt two homogeneous modules whose structure follows the typical Transformers. For local relations learning, we take the parallel multi-scale design and additive aggregation as simple but effective ideas, named MS-SCDC. For the global context modeling, we leverage the efficient factorized self-attention mechanism proposed in CoaT and apply it with the MS-SCDC in a cross-stacking manner over the high-resolution stages. Additionally, we further introduce a general approach for multi-scale learning of transformer-based modules, named MS-MHSA. The experiments conducted on a variety of general-purpose vision tasks demonstrate the superiority of the proposed model.
Xili Wan, Yaping Wu, Xinjie Guan, Aichun Zhu
IJCNN5
2023 Deformable multi-scale fusion network for non-uniform single image deblurring
Yang Chen 0008, Aichun Zhu, Hanxi Liu
Multim. Tools Appl.3
2023 Towards Optimal Application Offloading in Heterogeneous Edge-Cloud Computing
abstract
Application offloading plays a crucial role in application deployment in edge-cloud computing. However, finding the optimal solution for application offloading is challenging due to the heterogeneous resources, computing dependency of tasks, and complex network. Existing research on application offloading problems often assumes that the communication delay between the edge and the cloud (inter-side) is symmetrical or that within the cloud or edge (intra-side) can be omitted. However, this assumption is not practical considering the distinct features of the clouds and the edge clusters. Therefore, we study application offloading in the heterogeneous edge-cloud environment by considering both intra-side communication delay between tasks assigned to the same side and asymmetry inter-side communication delay between edge and cloud sides. We first focus on the specific circumstances with a boundary condition that lead to an optimal offloading solution in a heterogeneous edge-cloud environment. Then we study the general case by designing an iterative algorithm with maximum gain technique to solve it. Furthermore, considering various bandwidths within each side and resource capacities of physical nodes, we develop two efficient algorithms by combining minimum cut and maximum gain approaches. Both simulations and real trace-based evaluations are conducted to validate that the proposed algorithms outperform existing solutions.
Tingxiang Ji, Xili Wan, Xinjie Guan, Aichun Zhu, Feng Ye 0002
IEEE Trans. Computers4
2023 Synchronous Spatiotemporal Graph Transformer: A New Framework for Traffic Data Prediction
abstract
Modeling the spatiotemporal relationship (STR) of traffic data is important yet challenging for existing graph networks. These methods usually capture features separately in temporal and spatial dimensions or represent the spatiotemporal data by adopting multiple local spatial-temporal graphs. The first kind of method mentioned above is difficult to capture potential temporal-spatial relationships, while the other is limited for long-term feature extraction due to its local receptive field. To handle these issues, the Synchronous Spatio-Temporal grAph Transformer (S2TAT) network is proposed for efficiently modeling the traffic data. The contributions of our method include the following: 1) the nonlocal STR can be synchronously modeled by our integrated attention mechanism and graph convolution in the proposed S2TAT block; 2) the timewise graph convolution and multihead mechanism designed can handle the heterogeneity of data; and 3) we introduce a novel attention-based strategy in the output module, being able to capture more valuable historical information to overcome the shortcoming of conventional average aggregation. Extensive experiments are conducted on PeMS datasets that demonstrate the efficacy of the S2TAT by achieving a top-one accuracy but less computational cost by comparing with the state of the art.
Tian Wang 0002, Jinhu Lü 0001, Aichun Zhu, Hichem Snoussi, Baochang Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2022 Bi-level Doubly Variational Learning for Energy-based Latent Variable Models
abstract
Energy-based latent variable models (EBLVMs) are more expressive than conventional energy-based models. However, its potential on visual tasks are limited by its training process based on maximum likelihood estimate that requires sampling from two intractable distributions. In this paper, we propose Bi-level doubly variational learning (BiDVL), which is based on a new bi-level optimization framework and two tractable variational distributions to facilitate learning EBLVMs. Particularly, we lead a decoupled EBLVM consisting of a marginal energy-based distribution and a structural posterior to handle the difficulties when learning deep EBLVMs on images. By choosing a symmetric KL divergence in the lower level of our framework, a compact BiDVL for visual tasks can be obtained. Our model achieves impressive image generation performance over related works. It also demonstrates the significant capacity of testing image reconstruction and out-of-distribution detection.
Ge Kan, Jinhu Lü 0001, Tian Wang 0002, Baochang Zhang 0001, Aichun Zhu, Lei Huang 0015, Guodong Guo, Hichem Snoussi
CVPR5
2022 Look Before You Leap: Improving Text-based Person Retrieval by Learning A Consistent Cross-modal Common Manifold
abstract
The core problem of text-based person retrieval is how to bridge the heterogeneous gap between multi-modal data. Many previous approaches contrive to learning a latent common manifold mapping paradigm following a cross-modal distribution consensus prediction (CDCP) manner. When mapping features from distribution of one certain modality into the common manifold, feature distribution of the opposite modality is completely invisible. That is to say, how to achieve a cross-modal distribution consensus so as to embed and align the multi-modal features in a constructed cross-modal common manifold all depends on the experience of the model itself, instead of the actual situation. With such methods, it is inevitable that the multi-modal data can not be well aligned in the common manifold, which finally leads to a sub-optimal retrieval performance. To overcome this CDCP dilemma, we propose a novel algorithm termed LBUL to learn a Consistent Cross-modal Common Manifold (C3 M) for text-based person retrieval. The core idea of our method, just as a Chinese saying goes, is to 'san si er hou xing', namely, to Look Before yoU Leap (LBUL). The common manifold mapping mechanism of LBUL contains a looking step and a leaping step. Compared to CDCP-based methods, LBUL considers distribution characteristics of both the visual and textual modalities before embedding data from one certain modality into C3 M to achieve a more solid cross-modal distribution consensus, and hence achieve a superior retrieval accuracy. We evaluate our proposed method on two text-based person retrieval datasets CUHK-PEDES and RSTPReid. Experimental results demonstrate that the proposed LBUL outperforms previous methods and achieves the state-of-the-art performance.
Zijie Wang 0003, Aichun Zhu, Jingyi Xue, Xili Wan, Tian Wang 0002, Yifeng Li 0002
ACM Multimedia2
2022 CAIBC: Capturing All-round Information Beyond Color for Text-based Person Retrieval
abstract
Given a natural language description, text-based person retrieval aims to identify images of a target person from a large-scale person image database. Existing methods generally face a color over-reliance problem, which means that the models rely heavily on color information when matching cross-modal data. Indeed, color information is an important decision-making accordance for retrieval, but the over-reliance on color would distract the model from other key clues (e.g. texture information, structural information, etc.), and thereby lead to a sub-optimal retrieval performance. To solve this problem, in this paper, we propose to Capture All-round Information Beyond Color (CAIBC) via a jointly optimized multi-branch architecture for text-based person retrieval. CAIBC contains three branches including an RGB branch, a grayscale (GRS) branch and a color (CLR) branch. Besides, with the aim of making full use of all-round information in a balanced and effective way, a mutual learning mechanism is employed to enable the three branches which attend to varied aspects of information to communicate with and learn from each other. Extensive experimental analysis is carried out to evaluate our proposed CAIBC method on the CUHK-PEDES and RSTPReid datasets in both supervised and weakly supervised text-based person retrieval settings, which demonstrates that CAIBC significantly outperforms existing methods and achieves the state-of-the-art performance on all the three tasks.
Zijie Wang 0003, Aichun Zhu, Jingyi Xue, Xili Wan, Tian Wang 0002, Yifeng Li 0002
ACM Multimedia2
2022 ASPD-Net: Self-aligned part mask for improving text-based person re-identification with adversarial representation learning
Zijie Wang 0003, Jingyi Xue, Xili Wan, Aichun Zhu, Yifeng Li 0002, Xiaomei Zhu, Fangqiang Hu
Eng. Appl. Artif. Intell.4
2022 HARNet: Hierarchical adaptive regression with location recovery for crowd counting
Na Ni, Guangping Xie, Aichun Zhu, Yingna Wu
Neurocomputing4
2022 Frequency-driven channel attention-augmented full-scale temporal modeling network for skeleton-based action recognition
Fanjia Li, Aichun Zhu, Juanjuan Li, Yonggang Xu, Yandong Zhang, Hongsheng Yin 0001, Gang Hua 0002
Knowl. Based Syst.2
2022 SUM: Serialized Updating and Matching for text-based person retrieval
Zijie Wang 0003, Aichun Zhu, Jingyi Xue, Dai-Hong Jiang, Yifeng Li 0002, Fangqiang Hu
Knowl. Based Syst.2
2022 CACrowdGAN: Cascaded Attentional Generative Adversarial Network for Crowd Counting
abstract
Crowd counting is a valuable technology for extremely dense scenes in the transportation. Existing methods generally have higher-order inconsistencies between ground truth density maps and generated density maps. To address this issue, we incorporate an attentional discriminator to take charge of checking the density map between the generator and the ground truth. Thus, a Cascaded Attentional Generative Adversarial Network (CACrowdGAN) is proposed that enables the attentional-driven discriminator to distinguish implausible density maps and simultaneously to guide the generator to deliver fine-grained high quality density maps. The proposed CACrowdGAN consists of two components: an attentional generator and a cascaded attentional discriminator. The attentional generator has an attention module and a density module. The attention module is developed for the generator to focus on the crowd regions of the input images, while the density module is used to provide the attentional input of the discriminator. In addition, a cascaded attentional discriminator is proposed to synthesize attentional-driven fine-grained details at different crowd regions of the input image and compute a per-pixel fine-grained loss for training generator. The proposed CACrowdGAN achieves the state-of-the-art performance on five popular crowd counting datasets (ShanghaiTech, WorldEXPO’10, UCSD, UCF_CC_50 and UCF_QNRF), which demonstrates the effectiveness and robustness of the proposed approach in the complex scenes.
Aichun Zhu, Yaoying Huang, Tian Wang 0002, Jing Jin 0002, Fangqiang Hu, Gang Hua 0002, Hichem Snoussi
IEEE Trans. Intell. Transp. Syst.1
2021 DSSL: Deep Surroundings-person Separation Learning for Text-based Person Retrieval
abstract
Many previous methods on text-based person retrieval tasks are devoted to learning a latent common space mapping, with the purpose of extracting modality-invariant features from both visual and textual modality. Nevertheless, due to the complexity of high-dimensional data, the unconstrained mapping paradigms are not able to properly catch discriminative clues about the corresponding person while drop the misaligned information. Intuitively, the information contained in visual data can be divided into person information (PI) and surroundings information (SI), which are mutually exclusive from each other. To this end, we propose a novel Deep Surroundings-person Separation Learning (DSSL) model in this paper to effectively extract and match person information, and hence achieve a superior retrieval accuracy. A surroundings-person separation and fusion mechanism plays the key role to realize an accurate and effective surroundings-person separation under a mutually exclusion constraint. In order to adequately utilize multi-modal and multi-granular information for a higher retrieval accuracy, five diverse alignment paradigms are adopted. Extensive experiments are carried out to evaluate the proposed DSSL on CUHK-PEDES, which is currently the only accessible dataset for text-base person retrieval task. DSSL achieves the state-of-the-art performance on CUHK-PEDES. To properly evaluate our proposed DSSL in the real scenarios, a Real Scenarios Text-based Person Reidentification (RSTPReid) dataset is constructed to benefit future research on text-based person retrieval, which will be publicly available.
Aichun Zhu, Zijie Wang 0003, Yifeng Li 0002, Xili Wan, Jing Jin 0002, Tian Wang 0002, Fangqiang Hu, Gang Hua 0002
ACM Multimedia1
2021 AMEN: Adversarial Multi-space Embedding Network for Text-Based Person Re-identification
Zijie Wang 0003, Jingyi Xue, Aichun Zhu, Yifeng Li 0002, Chongliang Zhong
PRCV (2)3
2021 An enhanced 3DCNN-ConvLSTM for spatiotemporal multimedia data analysis
abstract
Summary At present, human action recognition is a challenging and complex task in the field of computer vision. The combination of CNN and RNN is a common and effective network structure for this task. Especially, we use 3DCNN in CNN part and ConvLSTM in RNN part. We divide the video into multiple temporal segments by average and compress each segment into one feature map by pooling layer. Adding the pooling layer, dropout layer, and batch normalization layer into ConvLSTM is our groundbreaking work. We test our model on KTH, UCF‐11, and HMDB51 datasets and achieve a high accuracy of action recognition.
Tian Wang 0002, Aichun Zhu, Hichem Snoussi, Chang Choi
Concurr. Comput. Pract. Exp.4
2021 A feature binding model in computer vision for object detection
Jing Jin 0002, Aichun Zhu, James Wright
Multim. Tools Appl.2
2021 Pose-Guided Inflated 3D ConvNet for action recognition in videos
Qianyu Wu, Aichun Zhu, Ran Cui, Tian Wang 0002, Fangqiang Hu, Yaping Bao, Hichem Snoussi
Signal Process. Image Commun.2
2021 CDADNet: Context-guided dense attentional dilated network for crowd counting
Aichun Zhu, Guoxiu Duan, Xiaomei Zhu, Yaoying Huang, Gang Hua 0002, Hichem Snoussi
Signal Process. Image Commun.1
2021 RecapNet: Action Proposal Generation Mimicking Human Cognitive Process
abstract
Generating action proposals in untrimmed videos is a challenging task, since video sequences usually contain lots of irrelevant contents and the duration of an action instance is arbitrary. The quality of action proposals is key to action detection performance. The previous methods mainly rely on sliding windows or anchor boxes to cover all ground-truth actions, but this is infeasible and computationally inefficient. To this end, this article proposes a RecapNet-a novel framework for generating action proposal, by mimicking the human cognitive process of understanding video content. Specifically, this RecapNet includes a residual causal convolution module to build a short memory of the past events, based on which the joint probability actionness density ranking mechanism is designed to retrieve the action proposals. The RecapNet can handle videos with arbitrary length and more important, a video sequence will need to be processed only in one single pass in order to generate all action proposals. The experiments show that the proposed RecapNet outperforms the state of the art under all metrics on the benchmark THUMOS14 and ActivityNet-1.3 datasets. The code is available publicly at https://github.com/tianwangbuaa/RecapNet.
Tian Wang 0002, Yang Chen 0030, Zhiwei Lin 0002, Aichun Zhu, Yong Li 0025, Hichem Snoussi, Hui Wang 0001
IEEE Trans. Cybern.4
2020 Abnormal event detection via the analysis of multi-frame optical flow information
Tian Wang 0002, Meina Qiao, Aichun Zhu, Guangcun Shan, Hichem Snoussi
Frontiers Comput. Sci.3
2020 Skeleton-based attention-aware spatial-temporal model for action detection and recognition
abstract
Action detection and recognition are popular subjects of research in the field of computer vision. The task of action detection can be regarded as the sum of action location and recognition. Action features described by using information concerning the human skeleton have the advantages of robustness against external factors and requiring a small amount of calculation. This study proposes a skeleton‐based action analysis model based on a recurrent neural network framework. The model learns action features by modelling static and dynamic features of skeleton joints and the importance of different video frames by introducing an attention module. For action location, conditional random field loss function is introduced to establish the context dependency of output labels. In the aspect of action recognition, the hierarchical training mechanism with triple loss models action features at coarse‐grained and fine‐grained levels. The authors’ proposed method delivers state‐of‐the‐art results on action location and recognition tasks.
Ran Cui, Aichun Zhu, Jingran Wu, Gang Hua 0002
IET Comput. Vis.2
2020 Exploring a rich spatial-temporal dependent relational model for skeleton-based action recognition by bidirectional LSTM-CNN
Aichun Zhu, Qianyu Wu, Ran Cui, Tian Wang 0002, Wenlong Hang, Gang Hua 0002, Hichem Snoussi
Neurocomputing1
2019 Multiple human upper bodies detection via candidate-region convolutional neural network
Aichun Zhu, Tian Wang 0002
Multim. Tools Appl.1
2018 Multi-source Learning for Skeleton -based Action Recognition Using Deep LSTM Networks
abstract
Skeleton-based action recognition is widely concerned because skeletal information of human body can express action features simply and clearly, and it is not affected by physical features of the human body. Therefore, in this paper, the method of action recognition is based on skeletal information extracted from RGBD video. Since the skeleton coordinates we studied are two-dimensional, our method can be applied to RGB video directly. The recently proposed method based on the deep network only focuses on the temporal dynamic of action and ignores spatial configuration. In this paper, a Multi-source model is proposed based on the fusion of the temporal and spatial models. The temporal model is divided into three branches, which perceive the global-level, local-level, and detail-level information respectively. The spatial model is used to perceive the relative position information of skeleton joints. The fusion of the two models is beneficial to improve the recognition accuracy. The proposed method is compared with the state-of-the-art methods on a large scale dataset. The experimental results demonstrate the effectiveness of our method.
Ran Cui, Aichun Zhu, Gang Hua 0002
ICPR2
2018 Exposing image resampling forgery by using linear parametric model
Aichun Zhu, Florent Retraint
Multim. Tools Appl.2
2018 Abnormal event detection via covariance matrix for optical flow based feature
Tian Wang 0002, Meina Qiao, Aichun Zhu, Yida Niu, Ce Li 0001, Hichem Snoussi
Multim. Tools Appl.3
2016 Detection of Abnormal Event in Complex Situations Using Strong Classifier Based on BP Adaboost
Tian Wang 0002, Meina Qiao, Aichun Zhu, Ce Li 0001, Hichem Snoussi
ICIC (2)4
2015 Articulated pose estimation via multiple mixture parts model
abstract
State-of-the-art methods for articulated human pose estimation are based on pictorial structures model (PS). Most of these methods predict the pose directly in part-based models and only consider rigid parts guided by human anatomy. In this paper, we propose a new framework for human pose estimation which is composed of two stages: pre-estimation and estimation. The first stage includes three steps: upper body detection, upper body categorization, and model selection. In the second stage, a new upper body category based multiple mixture parts (MMP) model is proposed. We present quantitative results demonstrating that our model significantly improves the accuracy of the pose estimation.
Aichun Zhu, Hichem Snoussi, Abel Cherouat
AVSS1