VLDB 2026 Research / reviewers in the wild / expert
Yidong Peng
dblp:148/1417
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-3779-0360ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cardiac cavity segmentation review in the past decade: Methods and future perspectives
Weisheng Li 0001, Yucheng Shu, Yidong Peng, Bin Xiao 0002 |
Neurocomputing | 4 |
| 2025 | ABIBE: Adaptive Building Information-Based Extraction From Remote Sensing Imagery Using Vision-Language ModelsabstractBuilding extraction from remote sensing imagery is essential for urban planning, population monitoring, and emergency response. However, traditional methods often struggle with accurate building delineation due to complex architectural features and varying imaging conditions. In this paper, we propose ABIBE, an Adaptive Building Information-Based Extraction framework that leverages Vision-Language Models (VLMs) for high-precision building extraction from remote sensing imagery. Our approach introduces two key innovations: (1) a hierarchical feature transfer mechanism that selectively extracts and adapts visual representations from the JanusPro vision-language model, and (2) a dynamic multi-scale feature fusion technique that adaptively integrates these features with spatial information through cross-modal attention mechanisms. Extensive experiments on the WHU Building Dataset demonstrate that our method significantly outperforms state-of-the-art approaches, achieving 91.39% IoU and 95.50% F1 score. We also provide comprehensive performance analysis including computational complexity, runtime efficiency, and resource requirements to evaluate the practical applicability of the proposed framework. The ABIBE framework provides a promising solution for high-precision building extraction from remote sensing imagery, with particular advantages in areas with complex architectural structures and high building density. Yongtao Deng, Dajiang Lei, Yidong Peng, Weisheng Li 0001, Liping Zhang 0012 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | GAPL-SegNet: Geometry-Aware Prototype Learning for Few-Shot Building Segmentation in Remote Sensing ImageryabstractFew-shot building segmentation in remote sensing imagery remains challenging due to limited annotated data and complex geometric structures inherent in building footprints. Traditional prototype-based methods ignore rich geometric priors of buildings, leading to suboptimal feature representations and poor generalization across different geographic regions. We propose GAPL-SegNet, a novel architecture that integrates Geometry-Aware Prototype Learning with adaptive feature aggregation for superior few-shot performance. Our approach introduces several key innovations: (1) a geometric feature extractor that captures building-specific structural patterns including edges, corners, and rectangularity through dedicated detection branches; (2) a geometry-aware prototype learning framework that leverages spatial importance weighting for more discriminative prototype construction based on geometric significance; (3) a progressive adaptive training strategy with dynamic geometric loss weighting that ensures effective integration of geometric priors; and (4) systematic cross-dataset analysis quantifying domain adaptation challenges. Extensive experiments demonstrate strong performance across multiple datasets and scenarios. On the WHU Building dataset with 100 training samples, GAPL-SegNet achieves 85.75±0.43% IoU (95% CI: [85.22, 86.29]) based on five independent runs, with notable stability (CV < 0.5%), outperforming the best baseline by 3.67% IoU. Cross-dataset evaluations reveal significant challenges from resolution differences: WHU to Inria achieves 51.68% IoU (41.4% with resolution alignment) and Inria to WHU reaches 53.71% IoU (33.6% with resolution alignment). Multi-resolution analysis demonstrates that spatial resolution differences (0.3m to 1.2m) cause up to 46% IoU degradation, identifying resolution adaptation as the primary domain shift challenge that exceeds the scope of geometric priors alone. Computational efficiency analysis reveals that our method achieves a favorable performance-efficiency balance with 22.5M parameters and 5.3ms inference time. In building segmentation’s specific few-shot setting, geometric priors bring better boundary consistency and deployment efficiency compared to general-purpose foundation models. Yongtao Deng, Dajiang Lei, Liping Zhang 0012, Yidong Peng, Weisheng Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | An Analytical Thermal Anisotropy Model Considering Roof Effect and Multiple Scattering in the Urban Canopy Over Sloping TerrainabstractAs urbanization accelerates, more and taller buildings and less greenery are closely related to changes in the urban thermal environment (UTE). Knowledge of spatial and temporal variations of UTE is becoming increasingly concerning, and this can be measured with the land surface temperature (LST). Satellite observation of LST is an important tool for monitoring; the strong thermal anisotropy limits the use of satellite thermal infrared (TIR) data. Hitherto, the poor investigation was focused on the modeling and analysis of urban thermal anisotropy (UTA), especially in mountainous urban areas with multi-slope environments. These areas exhibit a distinctive “roof effect”, which is defined as the radiative transfer effect between the roof and the adjacent wall due to the slope that results in different heights between the roofs; multiple scattering has also been changed. Although an analytical thermal anisotropy model for the urban canopy over sloping terrain (AU3SM) has been proposed, its inability to effectively account for roof effects and multi-scattering mechanisms limits its daytime TIR observation applicability. To address these limitations, we developed an enhanced AU3SM that considers the roof effect and multiple scattering, which is labeled AU3SM-RS. The model was evaluated using measurements based on unmanned aerial vehicles (UAVs) in the mountainous city of Chongqing, China, with values of the root mean square error (RMSE) and coefficient of determination (R2) of 0.83 K and 0.93 in UTA. Comparison with a graphic processing unit-based solution for the faster 3-D radiative transfer model (GRay) further validates the model’s reliability with RMSE and R2values of 0.12 K and 0.96, respectively. Simulations in a certain scenario reveal that as the slope increases, the roof effect increases and the multiple scattering effect decreases in UTA and brightness temperature (BT), ignoring the roof effect and multiple scattering can result in maximum UTA biases of approximately 0.54, 0.48, and 0.72 K, BT biases approximately 1.02, 1.62, and 2.4 K at 5°, 15°, and 30° slopes, the biases due to neglecting the second scattering are very slight compared to the first scattering. Under certain conditions with a slope of 10°, the wider roof and narrower roadway, a related more dramatic roof effect; the narrower and deeper street canyon, a related more dramatic scattering effect. The proposed model is an efficient computational tool to assess UTA in mountainous areas quickly. Xinguang Sang, Xiaobo Luo, Zunjian Bian, Biao Cao, Panpan Zhu, Yidong Peng, Tengyuan Fan |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Whole Heart Segmentation Based on 3D Contour-Guided Multi-Head Attention Network From CT and MRI ImagesabstractHeart image segmentation is a critical task in medical image processing, which is crucial for the diagnosis and treatment planning of cardiovascular diseases. It helps doctors understand patients' cardiac anatomy and functional status more comprehensively and lays the foundation for personalized medicine and precision medicine research. Addressing the current challenges of rough surfaces on the entire heart, incomplete segmentation of heart substructures, and the lack of structured prediction of pulmonary arteries due to artifacts, scale diversity, uneven intensity, and boundary ambiguity in cardiac computed tomography (CT) and magnetic resonance imaging (MRI) images, we propose a whole heart segmentation algorithm based on 3D contour guided network. The proposed algorithm achieves robust whole heart segmentation results and has few network structure parameters. To enhance the consistency of features extracted by the codec, we propose a 3D codec information integration module to focus on task-related areas. In the final stage of information integration, features of different scales are combined. A 3D contour attention module enhances the perception of the heart's structure and shape. Contour prediction results from the initial stage, generating a low-resolution voxel of the entire heart with contour details. The second stage builds upon the initial phase of secondary learning to achieve multi-label segmentation results. The proposed algorithm achieved average Dice scores of 0.905 and 0.865 for the CT and MRI modalities, respectively, in 40 cases. Weisheng Li 0001, Yidong Peng, Yucheng Shu |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | ER-OCN: Toward efficient network routing in ocean city based on deep reinforcement learning
Shu Yang 0002, Yaofeng Liu, Laizhong Cui, Yidong Peng, Victor C. M. Leung |
Comput. Commun. | 4 |
| 2022 | Spatiotemporal Reflectance Fusion via Tensor Sparse RepresentationabstractTradeoffs between the spatial and temporal resolutions of current satellite instruments limit our ability to conduct high-quality and continuous monitoring of the earth’s surface dynamics. Spatiotemporal image fusion has become increasingly necessary to obtain remote sensing images with high spatiotemporal resolution. However, current learning-based methods concentrate on predicting images only from spatial similarity and neglect spectral correlations of remote sensing images, leading to significant spectral information loss. In this article, we develop a novel nonlocal tensor sparse representation-based semicoupled dictionary learning approach (SCDNTSR) for spatiotemporal fusion. In the SCDNTSR method, the spectral correlation and the spatial similarity of the nonlocal similar cubes are simultaneously exploited through the tensor–tensor product-based tensor sparse representation. Furthermore, the semicoupled mapping prior knowledge of sparse coefficients across the high- and low-spatial resolution (HSR\LSR) image spaces is exploited with the coupled dictionary to constrain the similarity of sparse coefficients to improve the prediction performance. In addition, to capture additional prior spatial information, the SCDNTSR provides a new method to determine the degradation relationship between the target HSR and LSR difference images with the help of the known HSR and LSR difference images. The proposed SCDNTSR method was tested on real datasets at both the Coleambally Irrigation Area study site and the Lower Gwydir Catchment study site. Results show that the proposed method outperforms five state-of-the-art methods, especially in maintaining the spectral information, proving the feasibility of integrating the degradation relationship, spatio-spectral-nonlocal correlation, and semicoupled mapping priors of the multisource data into the proposed model. Yidong Peng, Weisheng Li 0001, Xiaobo Luo, Jiao Du, Xiayan Zhang, Yi Gan, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | A Geographically and Temporally Weighted Regression Model for Spatial Downscaling of MODIS Land Surface Temperatures Over Urban Heterogeneous RegionsabstractThe fine spatial resolution (~100 m) land surface temperature (LST) is a key variable of great concern in various environmental studies over urban heterogeneous regions. An improvement in the spatial resolution of the coarse spatial resolution LST is an effective way to extend its potential uses in applications that have strict requests on both the spatial and temporal resolutions. However, previous statistical downscaling algorithms were proposed mainly by addressing the spatial variability in the LST while neglecting the temporal variability. In this paper, we propose a new algorithm based on a geographically and temporally weighted regression (GTWR) model for spatial downscaling of the Moderate Resolution Imaging Spectroradiometer LST data from 1000 to 100 m. The GTWR-based algorithm with temporally and geographically varying regression coefficients can capture both the spatial and temporal variabilities in the LST from the time series data at a coarse spatial resolution for effectively reconstructing the subpixel variability in the LST at fine spatial resolution. In addition, because of a better ability to explain the LST variability over urban heterogeneous regions, a normalized difference built-up index and a digital elevation model were selected as auxiliary variables. Taking Beijing and Lanzhou as examples, the performance of the GTWR-based algorithm was assessed by comparing the results with the TsHARP and GWR-based algorithms and the Landsat-8 LST. The results indicate that the GTWR-based algorithm outperforms the above-mentioned algorithms with lower mean root mean square error (1.62 °C) and mean absolute error (1.28 °C) and better agreement between the GTWR downscaled LST and the Landsat-8 LST. Yidong Peng, Weisheng Li 0001, Xiaobo Luo, Hua Li 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |