VLDB 2026 Research / reviewers in the wild / expert
Yan Pu
dblp:87/11071
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A non-local sparse unmixing based hyperspectral change detection with unsupervised deep clustering
Tianqi Gao, Maoguo Gong, Xiangming Jiang, Yue Zhao 0024, Hao Liu 0123, Yan Pu |
Knowl. Based Syst. | 6 |
| 2025 | Hierarchical Feature Alignment-based Progressive Addition Network for Multimodal Change Detection
Tongfei Liu, Yan Pu, Tao Lei 0003, Jianjian Xu, Maoguo Gong, Lifeng He, Asoke K. Nandi |
Pattern Recognit. | 2 |
| 2025 | Scale-Aware Pruning Framework for Remote Sensing Object Detection via Multifeature RepresentationabstractWith the rapid advancements in computer vision, high-resolution remote sensing imagery has become a crucial data source for object detection. Nevertheless, effectively utilizing limited computational resources and reducing the burden on satellite edge devices remains a significant challenge. To effectively reduce model complexity while maintaining its representational capacity, this article proposes a scale-aware pruning framework (SAPF) to enhance remote sensing object detection ability. First, this article classifies the convolutional layers in object detection models into two categories: layers with a single-scale feature representation and layers with a multiscale feature representation. For convolutional layers with single-scale features, we utilize singular value decomposition (SVD) to quantify feature importance and assess filter redundancy to enhance model efficiency. By removing less critical filters, this pruning criteria aims to reduce the model size and computational load without compromising performance. However, convolutional layers with multiscale features are crucial for optimizing feature extraction and balancing information capture across various scales. To address this, this article evaluates the similarity between convolutional layers with different scales to determine the contribution of various scale features in multiscale fusion. Surprisingly, the SAPF can reduce the FLOPs and parameters, as well as ensure the representational ability obviously when the YOLO v5s and Faster-RCNN are adopted to classify the NWPU VHR-10, RSOD, and SIMD datasets. This means we can save the training computation resources for the model. Additionally, SAPF can significantly improve the efficiency of the model in object detection to ensure its real-time performance. Zhuping Hu, Maoguo Gong, Yue Zhao 0024, Mingyang Zhang 0002, Yiheng Lu, Jianzhao Li, Yan Pu, Zhao Wang 0011 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | D3PM: Dual-Stream Denoising Diffusion Probabilistic Model for Change Detection in Multimodal Remote Sensing ImagesabstractDetecting land cover changes from multi-temporal and multi-modal remote sensing images acquired by different sensors at the same location is a complex yet highly valuable task. Recently, diffusion models, exemplified by the Denoising Diffusion Probabilistic Model (DDPM), have garnered significant attention for their remarkable performance and straightforward architecture. These models excel in image generation, distribution modeling, and feature extraction, making them highly promising for advancing Multimodal Change Detection (MCD). In this paper, we propose a Dual-stream Denoising Diffusion Probabilistic Model (D3PM) to address the challenges of MCD. Specifically, D3PM leverages DDPM to design two distinct processing streams, one for each image modality. The first stream employs an unconditional DDPM, whose denoising encoder-decoder network can achieve robust feature extraction. The second stream employs a conditional DDPM to facilitate modal translation, enabling the extracted features to align with the characteristics of the other modality, thereby improving cross-modality comparability. To further enhance performance, we constructed a CD task branch based on the decoder features of the two DDPMs across multiple denoising time steps. Additionally, we designed a collaborative learning optimization strategy with asynchronous time steps, fostering cross-task knowledge sharing and mutual enhancement while preserving the integrity of individual task learning. Experimental results on multiple public datasets demonstrate the effectiveness and superiority of the proposed D3PM, which achieves efficient modal transformation and alignment, mitigates modal heterogeneity interference, and significantly improves detection performance. Fenlong Jiang, Xinlong Huo, Mingyang Zhang 0002, Maoguo Gong, Yan Pu, Yu Zhou 0051, Wei Zhao 0019, Ziyu Guan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Collaborative Frequency-Aware Transformer for Unsupervised Multimodal Change Detection in Heterogeneous Remote Sensing ImagesabstractMultimodal change detection (MCD), as an emerging task, aims at recognizing change regions from bi-temporal remote sensing images (RSI) of different modalities. Inspired by the success of the self-attention mechanism in transformer, attempts have been made to solve MCD through the transformer variants. However, transformer-based network optimization requires high-quality training samples. In addition, due to the significant differences in the data distribution, semantic information, and feature representation of multimodal data, transformer-based methods have obvious deficiencies in local feature representation and spatial consistency, especially when dealing with heterogeneous images. To address the above challenges, we propose a collaborative frequency-aware transformer for MCD (CFAT-MCD). As an unsupervised framework, CFAT-MCD is capable of learning more fine-grained patterns of land cover change through a few pseudo-labels. The CFAT is designed to enhance spatial consistency and align the features on a multi-scale basis, which can effectively mitigate the effects of modal differences. In addition, we propose a window-based spatial-frequency collaborative representation (SFCR) module to introduce frequency information into the spatial domain and improve the discriminability of spatial features. Extensive experiments on public datasets and quantitative analyses have validated the superior detection performance of our approach and the effectiveness of each module. Yan Pu, Maoguo Gong, Tongfei Liu, Mingyang Zhang 0002, Jianzhao Li, Hanhong Zheng, Yue Zhao 0024 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Adversarial Feature Equilibrium Network for Multimodal Change Detection in Heterogeneous Remote Sensing ImagesabstractChange detection (CD) methods have been crucial in exploring geo-environmental science. With the advancement of remote sensing (RS) technology, multimodal images acquired from different platforms and sensors are widely used for CD tasks. As an emerging task, multimodal CD (MCD) aims to achieve more comprehensive and precise detection of land cover changes through complementary information in multimodal images. However, there are significant differences between modalities, particularly in heterogeneous images. How to deal with modal differences while effectively integrating change information remains a challenge in MCD. In this article, we propose a novel adversarial feature equilibrium network (AFENet), which establishes an additional adversarial optimization to solve the equilibrium problem between modal differences and land cover changes. Our AFENet aligns the features and reduces the modal gap through a multiscale adversarial domain adaptation (MADA) approach. Meanwhile, a divergence-aware contrastive module (DCM) is designed as a regularization term for adversarial optimization. DCM affects the sensitivity of feature extractors by constraining the mutual information between changed and unchanged pixels. In this case, AFENet can maintain the consistency of feature representation while maximizing the discriminability of change targets. The features extracted from AFENet will then be integrated by our multistream feature fusion (MFF) module and utilized to generate change maps. The effectiveness of our approach is demonstrated on two scene-level multimodal RS datasets. Compared with existing methods, our AFENet achieves state-of-the-art (SOTA) performance on both datasets and outperforms the second-best$F1$score by 4.64% and 1.1%, respectively. Yan Pu, Maoguo Gong, Tongfei Liu, Mingyang Zhang 0002, Tianqi Gao, Fenlong Jiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Gradient-Guided Multiscale Focal Attention Network for Remote Sensing Scene ClassificationabstractRemote sensing scene classification (RSSC) aims to understand and analyze the semantic information at the scene level with complex geographical properties. Despite the profound success of advanced deep models in automatically capturing hierarchical embedding representations and the gradual dominant trend in RSSC, it still remains a great challenge to precisely focus on targets at variable scales that are considered highly relevant to the corresponding scene and separated from the background. Motivated by this recognition, in this article, we present the gradient-guided multiscale focal attention network (GMFANet) for RSSC to adaptively localize the representative multiscale semantic representation for complex scenes. In particular, a lightweight parameterized hierarchical multiscale attention (HMA) mechanism is proposed, which constitutes the main aim of adaptively enhancing physical detail and high-level semantic information at different layers, rather than regarding each scale set with equivalent insight, while eliminating redundant information inherent in conventional attention mechanisms. Subsequently, a gradient-guided spatial focused attention (GSFA) module is specifically designed to accurately localize critical regions at multiple scales, with the dynamic combination of gradient-activated reference attention map and prediction attention map from supervised information-based learning. In addition, a curriculum-driven dynamic attention fusion (CDAF) strategy is tailored to fuse the spatial attention above from easy to hard for avoiding from poor local optimum and decreasing the early learning ambiguity. Our extensive comparative experiments and ablation analyses implemented on real-world public RSSC datasets indicate that our approach achieves the state-of-the-art performance exactly. The code is available athttps://github.com/bling2beyond/GMFANet. Yue Zhao 0024, Maoguo Gong, A. K. Qin 0001, Mingyang Zhang 0002, Zhuping Hu, Tianqi Gao, Yan Pu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | A 2-20-GHz 360° Variable Gain Phase Shifter MMIC With Reverse-Slope Phase CompensationabstractThis paper presents a 2–20-GHz highly accurate vector-sum variable gain phase shifter (VGPS) monolithic microwave integrated circuit (MMIC) for ultra-broadband phased-array applications. An active power splitter followed by a dual-path I/Q network is utilized for precise I/Q signals generation. A reverse-slope phase compensation technique is proposed to improve broadband I/Q phase balances. Two active single-to-differential convertors further divide the I/Q signals into four quadrature signals. The vector modulator consisting of variable gain amplifiers (VGAs), passive analog adder and differential-to-single convertor synthesizes the desired output phase with tunable gain by orthogonal modulation and summation of the four quadrature signals. The designed VGPS, fabricated in 0.25-$\mu \text{m}$GaAs pHEMT technology, achieves 360° of full-span phase shift with tunable gain and a one-decade instantaneous bandwidth. Measurement result shows that the peak gain and input referred 1-dB compression point of the proposed VGPS are–3.1-dB and +6.5-dBm, respectively. With the aid of off-chip frequency-independent logic circuits and digital-to-analog convertor (DAC), the VGPS exhibits 6-bit phase resolution with root-mean-square (RMS) phase error of less than 2.92°over 2–20-GHz. The VGPS also results in a 4-bit gain control performance with 7.5-dB gain tuning range (RMS gain error < 0.34-dB) and a 5-bit gain control performance with 15.5-dB gain tuning range (RMS gain error < 0.62-dB). The VGPS MMIC consumes 115-mA current from a 5-V voltage supply and occupies chip area of 4.5-mm2 (excluding pads). Yitong Xiong, Yan Pu, Zhinan Yu, Xiaozong Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |