VLDB 2026 Research / reviewers in the wild / expert
Zhao Pei
dblp:96/7949
· DBLP profile ↗
28ranked-venue papers
10as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A novel multi-modal attentional collaborative learning framework with semantic enhancement for audio-visual question answering
Miao Ma, Zhao Pei, Longjiang Guo |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Dual-branch non-negative matrix factorization guided by information decoupling for multi-view clustering
Mingxia Gong, Chengcai Leng, Zhao Pei, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Neurocomputing | 4 |
| 2026 | MEMA-ConvLSTM: Spatiotemporal prediction via multi-scale autocorrelation memory and hierarchical fusion
Chengcai Leng, Huaiping Yan, Zhao Pei, Jinye Peng 0001 |
Inf. Sci. | 4 |
| 2026 | DGA-GCN: Dynamic Global Adaptive Graph Convolutional Networks for skeleton-based action recognition
Zhao Pei, Yanni Xue, Zhichao Ren, Chengcai Leng, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2025 | Dual graph-regularized low-rank representation for hyperspectral image denoising
Chengcai Leng, Mingpei Tang, Zhao Pei, Jinye Peng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | MFEL-YOLO for small object detection in UAV aerial images
Ting Hou, Chengcai Leng, Zhao Pei, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Expert Syst. Appl. | 4 |
| 2025 | SAGM-Net: Spatio-temporal action graph modeling network for weakly-supervised temporal action localization
Miao Ma, Zhao Pei |
Neurocomputing | 4 |
| 2025 | Multi-view data representation via adaptive label propagation nonnegative matrix factorization
Chengcai Leng, Jinye Peng 0001, Zhao Pei, Anup Basu |
Inf. Sci. | 4 |
| 2024 | Incremental semi-supervised graph learning NMF with block-diagonal
Xue Lv, Chengcai Leng, Jinye Peng 0001, Zhao Pei, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Feature matching based on Gaussian kernel convolution and minimum relative motion
Chengcai Leng, Huaiping Yan, Jinye Peng 0001, Zhao Pei, Anup Basu |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Autofocusing for Synthetic Aperture Imaging Based on Pedestrian Trajectory PredictionabstractOcclusions and complex backgrounds are common factors that hinder many computer vision applications. In a street scene, the challenge of accurately predicting pedestrian trajectories comes from the complexity of human behavior and the diversity of the external environment. It is difficult, if not impossible, to extract relevant information to accurately predict pedestrian trajectories in dynamic scenes. Synthetic aperture imaging (SAI) uses an array of cameras to mimic a camera with a large virtual convex lens by projecting images of a scene from different views onto a virtual focal plane. It is commonly used to reconstruct occluded objects, and in a street scene, can provide observation of pedestrians occluded by other objects and pedestrians. In this paper, we propose a joint prediction method based on autofocusing of SAI to predict pedestrian trajectories in dynamic scenes. The main contributions of this paper include: 1) The task of pedestrian trajectory prediction in dynamic scenarios is redefined as pedestrian trajectory prediction and SAI autofocusing from a practical but more challenging perspective. 2) The proposed method is based on an existing SAI-based method to extract information in heavily occluded views, which can obtain more accurate results but with less computational cost and without using other sensors such as LiDAR or depth cameras. 3) A new pedestrian trajectory prediction model, an attention-based trajectory prediction variational autoencoder (ATP-VAE), is proposed to extract complex human behavior and social interactions in dynamic scenes through a new Intention Attention Unit. The experimental results on multiple public datasets show that the proposed method achieves state-of-the-art results in the first-person perspective and in aerial view. Zhao Pei, Jianing Wang 0003, Yee-Hong Yang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | SSTA-Net: Self-supervised Spatio-Temporal Attention Network for Action Recognition
Zhao Pei |
ICIG (2) | 3 |
| 2023 | FeatsFlow: Traceable representation learning based on normalizing flows
Zhao Pei, Fei-Yue Wang 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | An ensemble belief rule base model for pathologic complete response prediction in gastric cancer
Jie Wu 0016, Miao Ma, Zhao Pei, Yingshi Sun |
Expert Syst. Appl. | 5 |
| 2023 | Long-Short Term Spatio-Temporal Aggregation for Trajectory PredictionabstractPedestrian trajectory prediction in crowd scenes plays a significant role in intelligent transportation systems. The main challenges are manifested in learning motion patterns and addressing future uncertainty. Typically, trajectory prediction is considered in two dimensions, including temporal dynamics modeling and social interactions capturing. For temporal dependencies, although existing models based on recurrent neural networks (RNNs) or convolutional neural networks (CNNs) achieve high performance on short-term prediction, they still suffer from limited scalability for long sequences. For social interactions, previous graph-based methods only consider fixed features but ignore dynamic interactions between pedestrians. Considering that the transformer network has a strong capability of capturing spatial and long-term temporal dynamics, we propose Long-Short Term Spatio-Temporal Aggregation (LSSTA) network for human trajectory prediction. First, a modern variant of graph neural networks, named spatial encoder, is presented to characterize spatial interactions between pedestrians. Second, LSSTA utilizes a transformer network to handle long-term temporal dependencies and aggregates the spatial and temporal features with a temporal convolution network (TCN). Thus, TCN is combined with the transformer to form a long-short term temporal dependency encoder. Additionally, multi-modal prediction is an efficient way to address future uncertainty. Existing auto-encoder modules are extended with static scene information and future ground truth for multi-modal trajectory prediction. Experimental results on complex scenes demonstrate the superior performance of our method in comparison to existing approaches. Cuiliu Yang, Zhao Pei |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Find Small Objects in UAV Images by Feature Mining and AttentionabstractWith the increasing popularity of Unmanned Aerial Vehicles (UAVs), the accuracy of detecting small objects in large-view images is also expected to increase. However, accurate small object detection is still a challenging problem. Currently, Image Pyramid Network, Feature Pyramid Network (FPN), rich training strategies and data augmentation are widely used to address this problem. To accurately detect small objects, the most important thing is to mine for more feature information. We propose Widened Residual Block (WRB) to break through the bottleneck of residual information gain to extract more feature information. The second is to emphasize or suppress features to prevent small objects from being overwhelmed by a broad background. We introduce an attention mechanism into PANet and propose Enhanced Attention PANet (EA-PANet), which consists of two parts: Context Attention Module (COAM) and Attention Enhancement Module (AEM). COAM outputs attention heatmaps with context, and AEM fuses features from the channel attention module (CAM) and COAM to avoid distraction from a vast background. In addition, we design a lightweight Decoupled Attention Head (DA-head) to dynamically compute important regions for specific tasks and achieve reliable predictions. Experiments show that our method outperforms state-of-the-art (SOTA) detectors. The source code for this work is available at https://github.com/liuxiaolei111/FindSmallObjects. Chengcai Leng, Xiaoming Niu, Zhao Pei, Irene Cheng 0001, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Remote Sensing Image Registration Based on Local Affine Constraint With Circle DescriptorabstractMany methods have been developed to improve the performance of image registration. In this letter, we introduce a novel method based on a local affine constraint for remote sensing image registration, which can be widely used in image processing and pattern recognition. Our algorithm has three components. First, we exploit the scale invariant feature transform (SIFT) method to extract feature points and calculate the gradient magnitude to establish feature descriptors with a circular instead of square neighborhood. Second, an initial matching is implemented by the nearest neighbor distance ratio (NNDR) and the fast sample consensus (FSC) algorithm. Finally, fine registration is established using more correct matches obtained by the local affine transformation circular region search algorithm. Experimental results show that the proposed method achieves subpixel accuracy. In addition, both the correct matching rate and registration demonstrate the effectiveness and efficiency of our method. Chengcai Leng, Guo-Rong Cai, Zhao Pei, Naigong Yu, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Alzheimer's disease diagnosis based on long-range dependency mechanism using convolutional neural network
Zhao Pei, Yuanshuai Gou, Miao Ma, Chengcai Leng, Jun Li 0009 |
Multim. Tools Appl. | 1 |
| 2022 | Multi-scale attention-based pseudo-3D convolution neural network for Alzheimer's disease diagnosis using structural MRI
Zhao Pei, Zhiyang Wan, Yanning Zhang 0001, Miao Wang 0008, Chengcai Leng, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2021 | Human Trajectory Prediction Using Stacked Temporal Convolutional NetworkabstractHuman trajectory prediction is essential for avoiding collisions in crowded environments in the navigation of autonomous driving system. In the past few years, human trajectory prediction based on deep learning methods has been extensively studied and significant progress has been made. However, both the social interactions among pedestrians and the intention of the pedestrians are unpredictable. They cause great difficulty in forecasting the future paths. Previous methods use recurrent architectures which will produce lots of inefficient parameters in training process and lead to the low accuracy. In this paper, we propose a novel method to predict human trajectory by using stacked temporal convolutional network. First, we use a kernel function to embed the social interactions among pedestrians with the adjacency matrix. Next, we represent a spatial temporal representation of all pedestrian in the scene. Then, we conduct a spatial temporal embedding for observed trajectories. Finally, the stacked temporal convolutional network receives the features extracted by spatial temporal network and predicts the future human trajectories. The proposed approach is a data-driven approach which combines the advantages of spatial temporal graph convolutional network and stacked temporal convolutional network. Experimental results on two public datasets demonstrate that our stacked temporal method has high data efficiency and outperforms the existing methods. Jinli Ma, Zhao Pei |
TrustCom | 3 |
| 2021 | GEA-net: Global embedded attention neural network for image classificationabstractRecently, it is generally acknowledged that the receptive field size of visual cortical neurons are regulated by the stimulus in the neuroscience community. Thus, once the global receptive field is obtained, the network performance can be greatly improved. Unfortunately, the larger receptive field method has been rarely considered in constructing CNNs. Recent studies on network design have demonstrated that the key to improving model performance is channel attention. However, they usually neglect the location information. Hence, it is difficult to capture the long-term dependency of location information. In particular, the location information is important for generating spatially selective attention blocks. Therefore, in this paper, we propose a novel attention neural network termed “GEA” by embedding global context information into channel attention. Firstly, instead of channel attention, which is transformed from feature tensor to single feature vector via 2D pooling, our method decomposes the channel attention into two 1D feature encoding processes that aggregate features along two spatial directions. In particular, the long-term dependency can be captured by using one spatial direction as well as preserving accurate location information in the opposite direction. Then, we concatenate the results of the two directions, while carry out batch normalization and use relu activation function. In addition, the cross channel soft attention is used to adaptively select different spatial scales of the information. Finally, our method is simple and can be flexibly inserted into classic networks, such as ResNet and EfficientNet, with limited computational overhead. A large number of experiments show that our method is not only conducive to the classification of COCO and ImageNet, but also demonstrates promising performance for 3D features, such as Magnetic Resonance Imaging(MRI). Zhiyang Wan, Zhao Pei, Chengcai Leng |
TrustCom | 3 |
| 2021 | All-in-focus synthetic aperture imaging using generative adversarial network-based semantic inpainting
Zhao Pei, Yanning Zhang 0001, Miao Ma, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2019 | Learner posture recognition via a fusing model based on improved SILTP and LDP
Yuqian Kuang, Zhao Pei |
Multim. Tools Appl. | 4 |
| 2019 | Human trajectory prediction in crowded scene using social-affinity Long Short-Term Memory
Zhao Pei, Xiaoning Qi, Yanning Zhang 0001, Miao Ma, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2018 | All-In-Focus Synthetic Aperture Imaging Using Image Mattingabstract“Seeing through” occluders is one of the most important effects that can be achieved with synthetic aperture imaging. As well, the occlusion problem, a challenging task for many computer vision applications, can be easily handled. Synthetic aperture imaging takes advantage of the property that only objects on the focal plane are sharp. The resulting image that is obtained by averaging images from different views consists of blurry objects away from the focal plane and sharp objects on the focal plane. Removing the blurriness caused by defocusing in synthetic aperture images to achieve an all-in-focus “seeing through” image is a challenging research problem. In this paper, we propose a novel method to improve the image quality of synthetic aperture imaging using image matting via energy minimization by estimating the foreground and the background. In particular, we first estimate the out-of-focus region by focusing on the background objects in each camera view using energy minimization. Next, we utilize a labeling method to create a sharp “see through” synthetic aperture image of the hidden objects. Then, image matting is used to extract the alpha matte of the hidden objects. Finally, by compositing the hidden objects with the estimated background regions, a sharp “see through” synthetic aperture image is created. The experimental results show that the proposed method outperforms the traditional synthetic aperture imaging method [1] as well as its improved versions [2]-[4], which simply dim and blur the area in the image that is out of focus, and a recent all-in-focus method [5]. We show that both the occluded objects and the background can be combined using our method to create a sharp synthetic aperture image. Zhao Pei, Xida Chen, Yee-Hong Yang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Synthetic aperture imaging using pixel labeling via energy minimization
Zhao Pei, Yanning Zhang 0001, Xida Chen, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2012 | A novel multi-object detection method in complex scene using synthetic aperture imaging
Zhao Pei, Yanning Zhang 0001, Tao Yang 0006, Xiuwei Zhang 0001, Yee-Hong Yang |
Pattern Recognit. | 1 |
| 2009 | A Method of Image Processing Algorithm Evaluation Based on Orthogonal Experimental DesignabstractOrthogonal experimental design, an effective method of multi-factors researches, is capable of reducing the times of experiments, determining the sequence of priority of influencing factors rapidly and figuring out the best parameters and confidence degree. The novel characteristic of this paper is to apply the orthogonal experimental design into the evaluation research of image processing algorithms. It can determine the sequence of priority of influencing factors and confidence degree by analyzing the range and variance of the results with experimental scheme. The experimental results demonstrate that the proposed method can accurately measure the sequence of priority of influencing factors by fewer experiments accurately, and also reflect the algorithm performance scientifically. Zhao Pei, Zenggang Lin |
ICIG | 1 |