EDBT 2026 Demo / reviewers in the wild / expert
Jiepan Li
dblp:332/0731
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-3971-8194ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Complex Backgrounds: A Unified Difference-Aware Decoder for Binary SegmentationabstractBinary segmentation is used to distinguish objects of interest from background, and is an active area of convolutional encoder-decoder network research. The current decoders are designed for specific objects based on the common backbones as the encoders, but cannot deal with complex backgrounds. Inspired by the way human eyes detect objects, we propose a new unified dual-branch decoder paradigm, termed the difference-aware decoder, to better explore the differences between foreground and background and to separate objects of interest in optical images. This decoder operates in two stages, leveraging multi-level features from the encoder. In the first stage, coarse detection of foreground objects is achieved by directly utilizing high-level semantic features, mimicking the initial rough observation of human vision. In the second stage, the decoder refines segmentation by exploring differences in low-level features, guided by the coarse map from the first stage. To enhance this process, we introduce two key innovations. First, a difference-aware prototype generation strategy leverages the guide map to extract foreground and background prototypes from high-level features, and calculates the similarity between these prototypes and corresponding representations in low-level feature spaces. Second, an overlapped window cross-level semantic guidance mechanism integrates high-level semantic information into low-level features through channel grouping and multi-scale aligned window pairs, guided by the similarities computed in the first strategy. Together, these innovations significantly enhance the DAD’s ability to discern subtle differences, enabling precise foreground extraction and effectively addressing the challenges of complex and varied backgrounds. To verify the performance of the proposed difference-aware decoder, we choose three well known backbones including ResNet, Res2Net, PVT, and two binary segmentation tasks,i.e., salient object detection, and camouflaged object detection, for comparative experiments. The results demonstrate that the difference-aware decoder can achieve higher accuracy than the other state-of-the-art binary segmentation methods for these tasks. The source code will be available on https://github.com/Henryjiepanli/DAD. Jiepan Li, Wei He 0003, Fangxiao Lu, Hongyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Advancing Chinese Lip Reading through Contextual EnhancementabstractLip reading, also known as visual speech recognition, aims to extract spoken words from silent videos by analyzing lip movements. Most existing methods have been developed primarily for English-language videos and are not optimized for Chinese lip reading. Given that Chinese is rich in homophones, which demand nuanced contextual cues for accurate disambiguation, conventional lip reading techniques tend to experience a significant drop in performance when applied to Chinese. To address these challenges, we propose the Context-Enhanced Visual Speech Recognition Network (ConVSR), specifically designed to improve Chinese lip reading. ConVSR incorporates intermediate connectionist temporal classification (inter CTC) residual modules to provide enhanced CTC supervision, improving the model’s contextual comprehension. Additionally, a Bi-Transformer Decoder is employed to simultaneously capture both past and future contextual information, thereby strengthening the model’s semantic understanding. Our approach achieves a $74.40 \%$ character error rate (CER) on the CAS-VSR-MOV20 dataset in track 1 of the Mandarin Audio-Visual Speech Recognition Challenge (MAVSR) 2025, securing first place and underscoring the effectiveness of ConVSR in advancing Chinese lip reading accuracy. Ruoyao Xue, Jiepan Li, Zhehui Wu, Wei He 0003 |
FG | 2 |
| 2025 | A Comprehensive Deep-Learning Framework for Fine-Grained Farmland Mapping From High-Resolution ImagesabstractThe extraction of large-scale farmland is essential for optimizing agricultural production and advancing sustainable development. To meet the urgent need for efficient farmland extraction and overcome existing technical challenges, we have developed a comprehensive farmland mapping framework that integrates advanced data, methodology, and cartographic techniques. Regarding data, we present the fine-grained farmland dataset (FGFD), which compiles high-quality, meticulously annotated very high-resolution (VHR) satellite images and captures distinct regional characteristics across eastern, southern, western, northern, and central China. Building on the FGFD, we propose the dual-branch boundary-aware network (DBBANet), which employs ResNet-50 as the encoder to extract multilayer encoded features and introduces two parallel decoding branches: a spatial-aware branch and a boundary-aware branch. The dual-branch architecture leverages both unique semantic information relevant to farmland and detailed boundary information, facilitating a more comprehensive and accurate representation of farmland areas. By combining this dataset with our innovative methodology, we further propose a farmland mapping framework designed for large-scale applications. The proposed framework enables the direct generation of high-precision vector maps from VHR images, providing crucial technical support for farmland management, resource assessment, and agricultural planning. Extensive experiments conducted on the FGFD have established benchmarks for 13 segmentation methods, demonstrating the state-of-the-art (SOTA) performance of our approach. In practical large-scale applications, our mapping framework produces high-precision vector maps with clear boundaries, bridging the gap in fine-grained farmland mapping and paving the way for further research and applications in this field. The source code of the proposed DBBANet and FGFD is available at:https://github.com/Henryjiepanli/DBBANet. Jiepan Li, Yipan Wei, Tiangao Wei, Wei He 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Learning without Exact Guidance: Updating Large-Scale High-Resolution Land Cover Maps from Low-Resolution Historical LabelsabstractLarge-scale high-resolution (HR) land-cover mapping is a vital task to survey the Earth's surface and resolve many challenges facing humanity. However, it is still a nontrivial task hindered by complex ground details, various landforms, and the scarcity of accurate training labels over a wide-span geographic area. In this paper, we propose an efficient, weakly supervised framework (Paraformer) to guide large-scale HR land-cover mapping with easy-access historical land-cover data of low resolution (LR). Specifically, existing land-cover mapping approaches reveal the dominance of CNNs in preserving local ground details but still suffer from insufficient global modeling in various landforms. Therefore, we design a parallel CNN-Transformer feature extractor in Paraformer, consisting of a downsampling-free CNN branch and a Transformer branch, to jointly capture local and global contextual information. Besides, facing the spatial mismatch of training data, a pseudo-label-assisted training (PLAT) module is adopted to reasonably refine LR labels for weakly supervised semantic segmentation of HR images. Experiments on two large-scale datasets demonstrate the superiority of Paraformer over other state-of-the-art methods for automatically up-dating HR land-cover maps from LR historical labels. Zhuohong Li, Wei He 0003, Jiepan Li, Fangxiao Lu, Hongyan Zhang 0001 |
CVPR | 3 |
| 2024 | Overcoming the Uncertainty Challenges in Flood Rapid Mapping with SAR DataabstractThe escalating intensity and frequency of floods, exacerbated by global climate change, emphasize the urgent need to address the growing risks of floods. Rapid and precise flood detection is paramount for efficiently responding to emergencies and executing disaster relief measures, enabling swift reactions to flood disasters and minimizing the losses incurred thus. The 2024 IEEE GRSS Data Fusion Contest Track 1 is centered on leveraging multi-source remote sensing data, particularly synthetic aperture radar (SAR) data, to classify flood and non-flood areas. In this contest, we acknowledge the significance of managing uncertain predictions and present an efficient Uncertainty-Aware Fusion Network (UAFNet). Specifically, we build on the traditional encoder-decoder architecture, initially employing the pyramid visual transformer (PVT) as a feature extractor. Subsequently, we apply a typical decoding strategy, namely the feature pyramid network, to obtain a flood extraction map with relatively high uncertainty. Furthermore, leveraging the uncertain extraction map, we introduce an Uncertainty Rank Algorithm to quantify the uncertainty level of each pixel of the foreground and background. We seamlessly integrate this algorithm with our proposed Uncertainty-Aware Fusion Module, enabling level-by-level feature refinement and ultimately yielding a refined extraction map with minimal uncertainty. Employing the proposed UAFNet, we utilize diverse versions of PVT as encoders to train multiple UAFNets. Additionally, we enhance our approach with online testing augmentation and multi-model fusion strategy, aiming to enhance the final flood extraction accuracy. Our technical solution has exhibited outstanding performance, earning the first-place ranking in the 2024 IEEE GRSS Data Fusion Contest Track 1 and achieving an impressive F1 score of 82.985% on the official test set. Jiepan Li, Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2024 | Overcoming the Uncertainty Challenges in Flood Rapid Mapping with Multi-Source Optical DataabstractAs global climate change worsens, floods are becoming more severe and frequent, urgently demanding effective flood risk mitigation strategies. Timely and precise flood inundation mapping is crucial for emergency response and relief. The 2024 IEEE GRSS Data Fusion Contest Track 2 aims to pioneer innovative algorithms for accurate flood extraction using multi-source optical remote sensing (RS) data. However, data diversity introduces aleatoric uncertainty, especially with synthetic, non-real data. Meanwhile, the vast coverage of RS imagery and the small proportion of flood areas cause a significant class imbalance, leading to epistemic uncertainty. In this paper, we propose an Uncertainty-aware Detail-Preserving Network (UADPNet) for rapid flood mapping of multi-source optical data. Firstly, we design an Aleatoric Uncertainty Estimator to model aleatoric uncertainty in multi-source data. Secondly, we introduce a Multi-Scale Convolution Block to extract multi-scale information without downsampling. Thirdly, we utilize a multi-level supervised strategy to quantify epistemic uncertainty and highlight uncertain pixels via the Uncertainty-Aware Fusion Module. With UADPNet, we adopt a multi-model fusion and post-processing strategy to enhance the final flood extraction accuracy. Our outstanding experimental results in the official test set showcase the superiority of our method, which secured the first-place ranking in the 2024 IEEE GRSS Data Fusion Contest Track 2, boasting an impressive F1 score of 89.843%. Jiepan Li, Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2024 | Identifying Every Building's Function in Large-Scale Urban areas with Multi-Modality Remote-Sensing DataabstractBuildings, as fundamental man-made structures in urban environments, serve as crucial indicators for understanding various city function zones. Rapid urbanization has raised an urgent need for efficiently surveying building footprints and functions. In this study, we proposed a semi-supervised framework to identify every building’s function in large-scale urban areas with multi-modality remote-sensing data. In detail, optical images, building height, and nighttime-light data are collected to describe the morphological attributes of buildings. Then, the area of interest (AOI) and building masks from the volunteered geographic information (VGI) data are collected to form sparsely labeled samples. Furthermore, the multi-modality data and weak labels are utilized to train a segmentation model with a semi-supervised strategy. Finally, results are evaluated by 20,000 validation points and statistical survey reports from the government. The evaluations reveal that the produced function maps achieve an OA of 82% and Kappa of 71% among 1,616,796 buildings in Shanghai, China. This study has the potential to support large-scale urban management and sustainable urban development. All collected data and produced maps are open access at https://github.com/LiZhuoHong/BuildingMap. Zhuohong Li, Wei He 0003, Jiepan Li, Hongyan Zhang 0001 |
IGARSS | 3 |
| 2024 | C2F-SemiCD: A Coarse-to-Fine Semi-Supervised Change Detection Method Based on Consistency Regularization in High-Resolution Remote Sensing ImagesabstractA high-precision feature extraction model is crucial for change detection. In the past, many deep learning-based supervised change detection methods learned to recognize change feature patterns from a large number of labelled bi-temporal images, whereas labelling bi-temporal remote sensing images is very expensive and often time-consuming. Therefore, we propose a coarse-to-fine semi-supervised change detection method based on consistency regularization (C2F-SemiCD), which includes a coarse-to-fine change detection network with a multi-scale attention mechanism(C2FNet) and a semi-supervised update method. Among them, the C2FNet network "gradually" completes the extraction of change features from coarse-grained to fine-grained through multi-scale feature fusion, channel attention mechanism, spatial attention mechanism, global context module, feature refine module, initial aggregation module, and final aggregation module. The semi-supervised update method uses the mean teacher method. The parameters of the student model are updated to the parameters of the teacher Model by using the exponential moving average (EMA) method. Through extensive experiments on three datasets and meticulous ablation studies, including crossover experiments across datasets, we verify the significant effectiveness and efficiency of the proposed C2F-SemiCD method. The code will be open at: https://github.com/ChengxiHAN/C2F-SemiCD-and-C2FNet. Chengxi Han, Chen Wu 0003, Meiqi Hu, Jiepan Li, Hongruixuan Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | UANet: An Uncertainty-Aware Network for Building Extraction From Remote Sensing ImagesabstractBuilding extraction aims to segment building pixels from remote sensing images and plays an essential role in many applications, such as city planning and urban dynamic monitoring. Over the past few years, deep learning methods with encoder–decoder architectures have achieved remarkable performance due to their powerful feature representation capability. Nevertheless, due to the varying scales and styles of buildings, conventional deep learning models always suffer from uncertain predictions and cannot accurately distinguish the complete footprints of the building from the complex distribution of ground objects, leading to a large degree of omission and commission. In this paper, we realize the importance of uncertain prediction and propose a novel and straightforward Uncertainty-Aware Network (UANet) to alleviate this problem. Specifically, we first apply a general encoder–decoder network to obtain a building extraction map with relatively high uncertainty. Second, in order to aggregate the useful information in the highest-level features, we design a Prior Information Guide Module to guide the highest-level features in learning the prior information from the conventional extraction map. Third, based on the uncertain extraction map, we introduce an Uncertainty Rank Algorithm to measure the uncertainty level of each pixel belonging to the foreground and the background. We further combine this algorithm with the proposed Uncertainty-Aware Fusion Module to facilitate level-by-level feature refinement and obtain the final refined extraction map with low uncertainty. To verify the performance of our proposed UANet, we conduct extensive experiments on three public building datasets, including the WHU building dataset, the Massachusetts building dataset, and the Inria aerial image dataset. Results demonstrate that the proposed UANet outperforms other state-of-the-art algorithms by a large margin. The source code of the proposed UANet is available at https://github.com/Henryjiepanli/Uncertainty-aware-Network. Jiepan Li, Wei He 0003, Weinan Cao, Liangpei Zhang 0001, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |