Ping Zhang 0023

dblp:13/4682-23 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
16since 2021 · last 2027
0000-0002-8330-2164ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2027 ViS2T: A vision and scene graph to text model for captive giant panda video captioning
Chenyu Ma, Chang Duan, Ke Zhang 0022, Mengnan He, Zhongrong Wang, Ping Zhang 0023, Ce Zhu
Expert Syst. Appl.7
2026 Flow-accelerated diffusion model for trajectory generation and optimization in offline reinforcement learning
He Diao, Xianglin Chen, Ping Zhang 0023, Zhenyu Feng, Bei Peng 0002
Eng. Appl. Artif. Intell.4
2026 GATOC: Learning temporal abstraction with the option transition graph attention mechanism
He Diao, Jingkui Zhang, Ping Zhang 0023, Gang Wang 0020, Zhenyu Feng, Bei Peng 0002
Expert Syst. Appl.6
2026 LongSF: Long state fusion with SSMs for multimodal 3D object detection
Pan Gao 0008, Xiwen Ren, Chun Fei, Ping Zhang 0023
Expert Syst. Appl.6
2026 A three-stage model for infrared small target detection with spatial and semantic feature fusion
Sixiang Ji, Haofei Zhang, Jingmin Zhang, Chun Fei, Xiaoyang Wang 0005, Juanxiu Liu, Ping Zhang 0023
Expert Syst. Appl.7
2026 HAVEN: Hierarchical diffusion and value-based trajectory selection for offline safe reinforcement learning
Erlie Wang, He Diao, Xianglin Chen, Jingkui Zhang, Xiaofeng Chai, Qiang Qi, Ping Zhang 0023
Neurocomputing7
2026 MPCF: Multi-Phase Consolidated Fusion for Multi-Modal 3D Object Detection With Pseudo Point Cloud
abstract
Pseudo points offer dimensional alignment between camera and LiDAR data for multi-modal 3D object detection. However, the pseudo points are irregularly and densely distributed, while LiDAR point distribution becomes sparse. Because the aliasing of different distributions can lead to redundancy or deficiency in objects, thereby causing size errors. Existing methods neglect the factor of consistent distribution representation across modalities, which compromises performance. In this paper, we propose a novel multi-phase consolidated fusion (MPCF) framework, a multimodal network favoring consistent feature distribution for 3D object detection. Specifically, we propose a region fusion strategy with cross-modal distribution alignment. This strategy utilizes matrix interactions and parameter sharing to align the inherent spatial distributions of each modality. Subsequently, multiple identical layers are employed to refine the channel distribution. This process organizes the features distribution from different modalities into a consistent representation. Furthermore, we restructured the pseudo-point branch, where conservative color weighting and extraction minimize interference with the feature distribution. Extensive experiments conducted on the KITTI and nuScenes benchmarks have demonstrated the remarkable performance and efficiency of MPCF. In the KITTI 3D detection test leaderboard of car, the MPCF has achieved 1st place among all single-use data methods. The code is publicly available at: https://github.com/ELOESZHANG/MPCF--3d_object_detection.
Pan Gao 0008, Xiwen Ren, Ping Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.5
2025 Coordinate-aware thermal infrared tracking via natural language modeling
Miao Yan, Ping Zhang 0023, Haofei Zhang, Ruqian Hao, Juanxiu Liu, Xiaoyang Wang 0005
Expert Syst. Appl.2
2025 Clip4Vis: Parameter-free fusion for multimodal video recognition
Qishi Zheng, Mengnan He, Jiuqin Duan, Gai Luo, Yimin Han, Qingyue Min, Ping Zhang 0023
Neurocomputing9
2023 Feature aggregation with transformer for RGB-T salient object detection
Ping Zhang 0023, Mengnan Xu, Pan Gao 0008, Jing Zhang 0043
Neurocomputing1
2023 Cross-Interaction Kernel Attention Network for Pansharpening
abstract
The aim of pansharpening is to fuse panchromatic (PAN) images and the corresponding low-resolution multispectral (LRMS) images to generate high-resolution multispectral (HRMS) images. In recent years, convolutional neural network (CNN)-based methods have obtained excellent performance in this field. In order to break through the limitations of standard convolution, the dynamic convolution is also proposed for pansharpening. However, the existing dynamic convolution ignores the context along channel dimension and does not consider the mutual enhancement between spatial and spectral information. In this letter, we propose a novel cross-interaction kernel attention network (CIKANet). The proposed model consists of two branches to extract spectral and spatial information. In particular, spatial and spectral kernel attention modules that can dynamically scale the parameters of convolutional kernels are applied. To strengthen the interactions between branches and obtain complementary information, we adopt a cross kernel attention structure. A series of experiments conducted on the WV3 and QB datasets suggest that CIKANet outperforms other state-of-the-art (SOTA) models visually and quantitatively.
Ping Zhang 0023, Yang Mei, Pan Gao 0008, Binxing Zhao
IEEE Geosci. Remote. Sens. Lett.1
2023 Enhanced Point Feature Network for Point Cloud Salient Object Detection
abstract
Point cloud salient object detection (SOD) aims to identify and segment the most prominent areas or targets in a 3D scene. Currently, research on point cloud SOD is still in its infancy, with most approaches neglecting the color information available in the point cloud. In this paper, we propose an enhanced point feature network (EPFNet) for point cloud SOD. Firstly, we extract RGB information from the point cloud with color and use it as input for the dual-stream network. Next, we introduce a multi-scale information enhancement (MIE) module to enhance common information and embed complementary information, acquiring RGB features at different scales and transforming them into point features. To gain access to global semantic information, we propose a point-image fusion (PIF) module, which aggregates the enhanced point features with the RGB global features and produces the final results. We conduct extensive experiments to validate the effectiveness of EPFNet and our approach outperforms 7 other state-of-the-art models on 4 metrics.
Pan Gao 0008, Siyi Peng, Chang Duan, Ping Zhang 0023
IEEE Signal Process. Lett.5
2023 Sparse Regularization-Based Spatial-Temporal Twist Tensor Model for Infrared Small Target Detection
abstract
Infrared (IR) small target detection under complex environments is an essential part of IR search and track systems. However, previously proposed IR small target detection algorithms cannot achieve complete suppression of complex and significant backgrounds. The spatial–temporal information of image sequences is not fully exploited. In this article, we present a sparse regularization-based twist tensor model for IR small target detection. First, the twist tensor model is built via perspective conversion based on the target’s local continuity in the spatial–temporal domain, which makes the original complicated background components more structured and increases the difference between the background and the target. Then, the structured sparsity-inducing norm is introduced to define the locality and continuity of the target. To further minimize the sparse background structures and global noise, the structured sparsity-inducing norm and the$l_{1}$norm are combined as the target’s parse constraint. Experimental results on real scenes reveal that the suggested method can process images with high detection accuracy and outstanding background suppression ability compared to various state-of-the-art methods.
Ping Zhang 0023, Lingyi Zhang
IEEE Trans. Geosci. Remote. Sens.2
2023 Infrared Small Target Detection Combining Deep Spatial-Temporal Prior With Traditional Priors
abstract
Infrared small target detection is a critical component of infrared search and track (IRST) systems. However, existing traditional methods only rely on traditional priors for detection, while existing deep learning methods depend heavily on training data with fine-grained annotations to obtain data-driven deep priors, which limits their effectiveness in infrared sequence detection. In this study, we propose a spatial-temporal tensor optimization model that combines the strengths of the dataset-free deep prior and traditional priors. On the one hand, we introduce the dataset-free deep spatial-temporal prior expressed by the untrained 3D Spatial-Temporal Prior Module (3DSTPM) to help capture the underlying spatial-temporal characteristics. On the other hand, we present the convolution-based TNN method and Nested Total Variation (NTV) regularization to respectively acquire the low-rank prior and local smoothness prior, aiding in target enhancement and background suppression. Furthermore, we can obtain accurate detection results by using an unsupervised solving algorithm based on the adaptive moment estimation (Adam) algorithm to learn parameters. Comprehensive experiments reveal that the proposed method performs better on real scenes than other state-of-the-art competitive detection approaches.
Pan Gao 0008, Sixiang Ji, Ping Zhang 0023
IEEE Trans. Geosci. Remote. Sens.5
2022 PMACNet: Parallel Multiscale Attention Constraint Network for Pan-Sharpening
abstract
Pan-sharpening, a task involving information fusion, entails merging panchromatic (PAN) images with high spatial resolution and low-resolution multispectral (LRMS) images in order to obtain high-resolution multispectral (HRMS) images. Due to deep learning’s excellent regression capabilities, it has recently become the dominating technique for this assignment. Meanwhile, the development of the transformer, a novel deep learning architecture for natural language processing, has provided researchers with new insights. In this letter, we seek to extend transformer’s excellent mechanisms to pixel-level fusion challenges. We designed a parallel convolutional neural network structure for learning both the regions of interest from the LRMS images and the residuals required for regression to HRMS images. Then, in our proposed pixelwise attention constraint (PAC) module, the residuals will be changed utilizing the learned region of interest. In addition, we presented a novel multireceptive-field attention block (MRFAB) to frame our network. Experiments on two datasets also show that our work is better than the mainstream algorithms at both indicators and visualization.
Yixun Liang, Ping Zhang 0023, Yang Mei, Tingqi Wang
IEEE Geosci. Remote. Sens. Lett.2
2021 Edge and Corner Awareness-Based Spatial-Temporal Tensor Model for Infrared Small-Target Detection
abstract
Infrared (IR) small-target detection has been a widely studied task in IR search and tracking systems. It remains a challenging problem, especially in heterogeneous scenarios, where it is very difficult to discriminate true targets from sparse residuals in the background. A novel edge and corner awareness-based spatial–temporal tensor (ECA-STT) model is presented in this article. First, we construct an STT based on a spatial–temporal correlation analysis of the IR video background. Then, we propose an indicator to highlight the target through adjustable importance measurements of the edge and corner. The tensor-based nonlocal total variation is also adopted to describe the edges in the background. The target–background separation problem is modeled as a tensor robust principal component analysis (TRPCA) problem with the tensor rank function replaced by the tensor truncated nuclear norm. The proposed model is solved by an effective optimization algorithm derived from the alternating direction method of multipliers (ADMM). Extensive experiments verify the superior abilities of the proposed model in target enhancement and background suppression.
Ping Zhang 0023, Lingyi Zhang, Xiaoyang Wang 0005, Fengcan Shen, Chun Fei
IEEE Trans. Geosci. Remote. Sens.1
2020 Stereoscopic video saliency detection based on spatiotemporal correlation and depth confidence optimization
Ping Zhang 0023, Xiaoyang Wang 0005, Chun Fei, Zhengkui Guo
Neurocomputing1
2018 Unsupervised Saliency Detection in 3-D-Video Based on Multiscale Segmentation and Refinement
abstract
In this letter, we propose an unsupervised salient object detection method in three-dimensional videos. Both temporal and depth information are efficiently considered, and multiscale architecture and graph-based refinement are built to improve accuracy and robustness. First, the input video frame is segmented into nonoverlapping superpixels by combining both appearance and depth information at the input. A multiscale architecture is also deployed after the segmentation with different segmentation parameters. Second, the initial saliency score of each segmented superpixel in each scale is calculated via global contrast, which is defined by appearance, depth, and motion cues from two consecutive frames. Third, the initial saliency in each scale is refined by smoothing over graphs built by three spatial–temporal feature priors—color, depth, and motion. Finally, the result is obtained by fusing three refined saliency maps in three scales. The experiments on two widely used datasets illustrate that our method outperforms state-of-the-art algorithms in terms of accuracy, robustness, and reliability.
Ping Zhang 0023, Pengyu Yan, Fengcan Shen
IEEE Signal Process. Lett.1
2017 Infrared dim target detection based on total variation regularization and principal component pursuit
Xiaoyang Wang 0005, Zhenming Peng, Dehui Kong, Ping Zhang 0023, Yanmin He
Image Vis. Comput.4
2017 Infrared Small Target Detection via Nonnegativity-Constrained Variational Mode Decomposition
abstract
Infrared small target detection is one of the key techniques in the infrared search and track system. Frequency differences among target, background, and noise are often important information for target detection. In this letter, a nonnegativity-constrained variational mode decomposition (NVMD) method is proposed. Unlike the traditional frequency-domain methods, the proposed method can adaptively decompose the input signal into several separated band-limited subsignals, with the nonnegativity constraint. First, a bandpass filter is used as a preprocessing step. Second, by exploring the frequency and nonnegativity properties of the small target, the NVMD model is constructed. The potential target subsignal can be obtained by solving the NVMD model. By performing threshold segmentation on the potential target subsignal, we can obtain the detection result of the infrared small target. Experiments on six real infrared image sequences demonstrate that the proposed method has a good performance in target enhancement and background suppression. Additionally, the proposed method shows strong robustness under various backgrounds.
Xiaoyang Wang 0005, Zhenming Peng, Ping Zhang 0023, Yanmin He
IEEE Geosci. Remote. Sens. Lett.3
2017 Multi-sensor image super-resolution with fuzzy cluster by using multi-scale and multi-view sparse coding for infrared image
Xiaomin Yang, Wei Wu 0002, Kai Liu 0012, Wei-long Chen, Ping Zhang 0023
Multim. Tools Appl.5