EDBT 2026 Demo / reviewers in the wild / expert
Dilong Li
dblp:146/2332
· DBLP profile ↗
21ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-5826-5568ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SFE-CapsNet: Spatial Feature Enhanced Capsule Networks for Remote Sensing Object DetectionabstractRemote sensing imagery often involves complex backgrounds and multi-scale targets, while variations such as rotation and scaling significantly degrade the performance of existing object detection algorithms and hinder effective modeling of spatial relationships between objects. To address these challenges, we propose SFE-CapsNet. First, we fuse scalar features extracted by convolutional neural networks (CNNs) with vector features from capsule networks through structural reorganization to form virtual capsules, approximating the dynamic routing process via a single fully connected layer. This design preserves object pose and texture information while streamlining information flow. Second, we introduce a capsule attention module that generates attention masks to dynamically enhance target-relevant features and suppress background noise, strengthening multi-level feature representations. Integrated with a feature pyramid network (FPN) architecture, our approach achieves precise detection of targets at varying scales. Experimental results demonstrate that multi-level feature fusion and the capsule attention mechanism significantly improve detection accuracy and robustness, achieving 77.65% mean Average Precision (mAP) on the DOTA dataset and 97.63% mAP on HRSC2016, highlighting its effectiveness and efficiency in complex scenes. Ziyi Chen 0001, Wenhui Qiu, Huayou Wang, Dilong Li, Jin Gou, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2026 | Visible and Infrared Image Fusion Based on Adaptive Weighted Multimodal Features Extraction and Bidirectional Guidance StructureabstractEfficient fusion of infrared and visible images is of critical importance for real-time applications such as autonomous driving. While deep learning-based fusion methods have demonstrated significant improvements in fusion quality in recent years, current network architectures still exhibit unsatisfactory computational complexity and processing speed. To reduce computational complexity and improve fusion efficiency, mask-based methods or approaches driven by downstream tasks often prioritize key regions. However, such methods tend to overemphasize target objects, potentially overlooking contextually significant elements. To address this limitation and achieve more effective fusion, we propose BEFuse, a decoupled two-stage training strategy with an end-to-end inference framework. BEFuse extracts shallow features and gradient information from images in the first stage via a cross-modal image segmentation subnetwork. In the fusion stage, we use the Hadamard product to map features to an implicit quadratic feature space, combining feature similarity and gradient mask information, allowing automatic adjustment of loss weights and improving fusion accuracy. Experiments on four datasets (MSRS, TNO, RoadScene, and M3FD) show that BEFuse outperforms existing methods in both fusion quality and computational speed. Ziyi Chen 0001, Gaosheng Cai, Dilong Li, Jing Wang 0049, Jin Gou, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Leveraging Multi-View Images to Learn Domain-Invariant Discriminative Embeddings for Cross-View Geo-LocalizationabstractCross-view geo-localization (CVGL) aims to match images of the same location captured from different viewpoints, such as those captured by Unmanned Aerial Vehicles (UAVs) and satellite platforms. The task is particularly challenging due to significant variations in scale, viewpoint, and illumination. Most existing methods employ symmetric sampling strategy to construct drone–satellite image pairs for deep metric learning, but neglect the potential of incorporating multi-view drone images to enhance the viewpoint robustness of features. To address this, we propose leveraging multi-view images to learn Domain-Invariant Discriminative Embeddings (DIDE) for CVGL. DIDE introduces an Inter-view Feature Aggregation Module (IFAM), which dynamically integrates multi-view drone information into robust embeddings. These are used in contrastive learning with satellite embeddings within batches to learn view-invariant discriminative features, while representation learning further improves scene discrimination across batches. To reduce the domain gap, DIDE constructs and aligns drone and satellite prototypes for effective cross-domain feature alignment. Furthermore, we adopt a parameter-efficient transfer learning strategy that leverages the capabilities of pre-trained foundation models while fine-tuning only dual adapters, significantly reducing the trainable parameters. DIDE achieves the state-of-the-art on University-1652 and University-160k, competitive results on SUES-200, and demonstrates strong cross-dataset transferability, with fewer training parameters and lower computational cost. Ziyi Chen 0001, Dilong Li, Jin Gou, Cheng Wang 0003, Kyle Gao, Jonathan Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Serialization Based Point Cloud Oversegmentation
Chenghui Lu, Jianlong Kwan, Dilong Li, Ziyi Chen 0001, Haiyan Guan |
ICCV | 3 |
| 2025 | LGMamba: Large-Scale ALS Point Cloud Semantic Segmentation With Local and Global State-Space ModelabstractThe large scale and extensive coverage of point cloud data make large-scale airborne laser scanning (ALS) point cloud semantic segmentation a highly challenging task. Although transformers have shown impressive performance in large-scale point cloud semantic segmentation task, their quadratic complexity limits the processing capacity. To alleviate this issue, we propose Local and Global Mamba (LGMamba)—a novel state-space model (SSM)-based network for large-scale point cloud semantic segmentation. Specifically, we propose Local Mamba module to extract fine-grained local features by effectively capturing local dependencies. Then, we propose Global Mamba module to refine the learned local features by capturing the global long-distance dependencies of whole scenes. The validation of our method on the DALES datasets was conducted. Extensive experimental results demonstrate the effectiveness of LGMamba, with mean intersection over union (mIoU) of 82.3% and overall accuracy (OA) of 97.7% on DALES. Dilong Li, Chongkei Chang, Ziyi Chen 0001, Jixiang Du |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | SECBNet: Semantic Segmentation-Enhanced Color Balance Network for Optical Satellite ImagesabstractEarth observation satellites can capture optical images under different temporal, climatic conditions, and platforms exhibit substantial differences in color and brightness, leading to poor visual experiences when synthesizing large-area optical satellite images. The related issue of color balancing has attracted considerable attention from researchers, yet challenges such as a lack of research data and sensitivity to model parameters persist. To address these problems, this article publishes a publicly open dataset and presents a semantic segmentation-enhanced color balance network (SECBNet). First, to mitigate the scarcity of research data, we develop a publicly available remote sensing image color balance dataset, Zhu Hai color balance image (ZHCBI), to support related research activities. Second, to improve semantic consistency between the color-balanced images and the target images, we design a dual-branch U-Net architecture guided by segmentation results and propose a novel segmentation feature loss function. Finally, to address issues of seams and unnatural transitions between blocks in segmented processing, we introduce a postprocessing module based on weighted averaging. We conducted comparative experiments and analyses with existing mainstream color balancing algorithms on the ZHCBI dataset. The results demonstrate that our proposed method achieves state-of-the-art color balancing quality, with significant improvement in visual effects and a higher peak signal-to-noise ratio (PSNR) (23.64 dB) compared with other mainstream methods. Ziyi Chen 0001, Hanhuang Chen, Lujuan Gao, Dilong Li, Cheng Wang 0003, Linlin Xu, Somayeh Mollaee, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DBARCT: Road Extraction Based on Double-Branch Architecture and Random Block Coding TransformerabstractAlthough transformer models are main network architectures for the delineation of roads from remote sensing imagery, they have critical limitations due to their regular patch mechanism and inefficiency in local information learning. To address these limitations for enhanced road extraction, this letter presents a novel double-branch architecture and random block coding transformer (DBARCT), with the following contributions. First, to improve local spatial details’ learning, we integrate transformer with convolutional neural network (CNN) into a novel dual-branch encoder-decoder architecture, such that the resulting model is efficient at learning both the local edge information and the global context information that are highly complementary for accurate road extraction. Second, to additionally augment the learning of global contextual information, we integrate the regular patching approach in traditional transformer models with a new irregular patching approach, such that it can better capture the global spatial information correlations that might be ignored by the regular patching approach. Third, an array of tests was carried out to meticulously scrutinize the efficacy of the fundamental elements of the suggested model. The empirical findings reveal that the intersection over union (IoU) metric attained by the proposed methodology on the LRSNY dataset stands at 88.53%, thereby corroborating the efficacy and preeminence of our approach in tasks related to road extraction. Ziyi Chen 0001, Yucai Chen, Lujuan Gao, Dilong Li, Linlin Xu, Jonathan Li 0001, Cheng Wang 0003, Yewang Chen |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Local Enhanced Transformer Networks for Land Cover Classification With Airborne Multispectral LiDAR DataabstractTransformer networks have demonstrated remarkable performance in point cloud processing tasks. However, balancing local feature aggregation with long-range dependency modeling remains a challenging issue. In this work we present a local enhanced Transformer network (LETNet) for land cover classification with multispectral LiDAR data. Specifically, we first rethink position encoding in 3D Transformers and design a novel feature encoding module that embeds comprehensive geometric and semantic information, serving a similar purpose. Then, the proposed local enhanced Transformer module is used to capture the accurate global attention weights and refine the features. Finally, to effectively extract and integrate global features across various scales, an attention-based pooling module is introduced. This module extracts global features from each encoder and decoder layer and constructs a feature pyramid to fuse these multi-scale global features. Both quantitative assessments and comparative analyses demonstrate the competitive capability and advanced performance of the LETNet in land cover classification task. Dilong Li, Shenghong Zheng, Ziyi Chen 0001, Jonathan Li 0001, Jixiang Du |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | BrGAN: Blur Resist Generative Adversarial Network With Multiple Joint Dilated Residual Convolutions for Chlorophyll Color Image RestorationabstractThis paper presents a Blur Resist Generative Adversarial Network (GAN) (BrGAN) with multiple joint dilated residual convolutions for chlorophyll image restoration of the Geostationary Ocean Color Imager (GOCI). First, a publicly available dataset was built to support this study. Second, a multiple attention perception mechanism and a multiple joint dilated residual convolution module was proposed to cope with the challenge of large missing areas in GOCI chlorophyll images. Third, a patch GAN based discrimination module was proposed to avoid the restored areas with generating mosaic and shadows. Our experimental results demonstrate that the BrGAN can reach 37.06 in the peak signal-to-noise ratio (PSNR) and 0.0485 in the Learned Perceptual Image Patch Similarity (LPIPS), respectively. The comparative study shows that the BrGAN achieves the highest effectiveness and advancement among other seven state-of-the-art methods. Ziyi Chen 0001, Yuhua Luo, Yiping Chen 0002, Jing Wang 0049, Dilong Li, Kyle Gao, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | An Infrared Moving Small Object Detection Method Based on Trajectory Growth
Dilong Li, Shuixin Pan, Yueqiang Zhang, Linyu Huang, Hongxi Guo |
PRCV (4) | 1 |
| 2022 | JFT: A Robust Visual Tracker Based on Jitter Factor and Global Registration
Shuixing Pan, Dilong Li, Yueqiang Zhang, Linyu Huang, Hongxi Guo |
PRCV (4) | 3 |
| 2022 | RoadCapsFPN: Capsule Feature Pyramid Network for Road Extraction From VHR Optical Remote Sensing ImageryabstractRoad detection plays an important role in a wide range of applications. However, due to size variations, spectral diversities, occlusions, and complex scenarios, it is still challenging to accurately extract roads from very-high resolution (VHR) optical remote sensing images. This paper proposes a capsule feature pyramid network for extracting road networks from VHR optical images, termed as RoadCapsFPN. By designing a capsule feature pyramid network, the RoadCapsFPN extracts and integrates multiscale capsule features to recover a high-resolution and semantically strong road feature representation. Next, we also design a contextual feature module, including dense atrous convolution (DAC) and residual multi-kernel pooling (RMP) units, to further exploit rich contextual properties of the roads at a high-resolution perspective. Benefitting from the multiscale feature abstraction and context augmentation, our RoadCapsFPN shows impressing results in processing variedly-sized and diversely-spectral roads in complex environments. Two testing datasets, Google and Massichusate Roads Datasets, are used for evaluating the proposed RoadCapsFPN via four testing indicators -precision,recall, intersection-over-union (IoU), and$F_{1}$-score. Comparative studies also confirm the superior performance of the RoadCapsFPN in accurately extracting road networks. Haiyan Guan, Yongtao Yu, Dilong Li, Hanyun Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | CCapFPN: A Context-Augmented Capsule Feature Pyramid Network for Pavement Crack DetectionabstractPeriodically monitoring the pavement conditions is of great importance to many intelligent transportation activities. Timely and correctly identifying the distresses or anomalies on pavement surfaces can help to smooth traffic flows and avoid potential threats to pavement securities. In this paper, we develop a novel context-augmented capsule feature pyramid network (CCapFPN) to detect cracks from pavement images. The CCapFPN adopts vectorial capsules to represent high-level, intrinsic, and salient features of cracks. By designing a feature pyramid architecture, the CCapFPN can fuse different levels and different scales of capsule features to provide a high-resolution, semantically strong feature representation for accurate crack detection. To take advantage of the context properties, a context-augmented module is embedded into each stage of the CCapFPN to rapidly enlarge the receptive field. The CCapFPN performs effectively and efficiently in processing pavement images of diverse conditions and detecting cracks of different topologies. Quantitative evaluations show that an overall performance with a precision, a recall, and an F-score of 0.9200, 0.9149, and 0.9174, respectively, were achieved on the test datasets. Comparative studies with some existing deep learning and edge based crack detection methods also confirm the superior performance of the CCapFPN in crack detection tasks. Yongtao Yu, Haiyan Guan, Dilong Li, Shenghua Jin, Changhui Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Capsule Feature Pyramid Network for Building Footprint Extraction From High-Resolution Aerial ImageryabstractBuilding footprint extraction plays an important role in a wide range of applications. However, due to size and shape diversities, occlusions, and complex scenarios, it is still challenging to accurately extract building footprints from aerial images. This letter proposes a capsule feature pyramid network (CapFPN) for building footprint extraction from aerial images. Taking advantage of the properties of capsules and fusing different levels of capsule features, the CapFPN can extract high-resolution, intrinsic, and semantically strong features, which perform effectively in improving the pixel-wise building footprint extraction accuracy. With the use of signed distance maps as ground truths, the CapFPN can extract solid building regions free of tiny holes. Quantitative evaluations on an aerial image data set show that a precision, recall, intersection-over-union (IoU), and F-score of 0.928, 0.914, 0.853, and 0.921, respectively, are obtained. Comparative studies with six existing methods confirm the superior performance of the CapFPN in accurately extracting building footprints. Yongtao Yu, Yongfeng Ren, Haiyan Guan, Dilong Li, Changhui Yu, Shenghua Jin, Lanfang Wang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | A Cascaded Deep Convolutional Network for Vehicle Logo Recognition From Frontal and Rear Images of VehiclesabstractVehicle logo recognition provides an important supplement to vehicle make and model analysis. Some of the existing vehicle logo recognition methods depend on the detection of license plates to roughly locate vehicle logo regions using prior knowledge. The vehicle logo recognition performance is greatly affected by the license plate detection techniques. This paper presents a cascaded deep convolutional network for directly recognizing vehicle logos without depending on the existence of license plates. This is a two-stage processing framework composed of a region proposal network and a convolutional capsule network. First, potential region proposals that might contain vehicle logos are generated by the region proposal network. Then, the convolutional capsule network classifies these region proposals into the background and different types of vehicle logos. We have evaluated the proposed framework on a large test set towards vehicle logo recognition. Quantitative evaluations show that a detection rate, a recognition rate, and an overall performance of 0.987, 0.994, and 0.981, respectively, are achieved. Comparative studies with the Faster R-CNN and other three existing methods also confirm that the proposed method performs effectively and robustly in recognizing vehicle logos of various conditions. Yongtao Yu, Haiyan Guan, Dilong Li, Changhui Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | A Hybrid Capsule Network for Land Cover Classification Using Multispectral LiDAR DataabstractLand cover mapping is an effective way to quantify land resources and monitor their changes. It plays an important role in a wide range of applications. This letter proposes a hybrid capsule network for land cover classification using multispectral light detection and ranging (LiDAR) data. First, the multispectral LiDAR data were rasterized into a set of feature images to exploit the geometrical and spectral properties of different types of land covers. Then, a hybrid capsule network composed of an encoder network and a decoder network is trained to extract both high-level local and global entity-oriented capsule features for accurate land cover classification. Quantitative classification evaluations on two data sets show that the overall accuracy, average accuracy, and kappa coefficient of over 97.89%, 94.54%, and 0.9713, respectively, are obtained. Comparative studies with five existing methods confirm that the proposed method performs robustly and accurately in land cover classification using the multispectral LiDAR data. Yongtao Yu, Haiyan Guan, Dilong Li, Tiannan Gu, Lanfang Wang, Lingfei Ma, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | 3-D Feature Matching for Point Cloud Object ExtractionabstractEffective object extraction plays an important role in many point cloud-based applications. This letter proposes a 3-D feature matching framework for point cloud object extraction. To determine the optimal affine transformation parameters for each template feature point, a convex dissimilarity function and the locally affine-invariant geometric constraints are designed to construct the overall objective function. The 3-D feature matching framework is integrated into a point cloud object extraction workflow. Extraction results on six test data sets show that average completeness, correctness, quality, and F1-measure of 0.96, 0.97, 0.93, and 0.96, respectively, are obtained in extracting light poles, vehicles, and palm trees. Comparative studies also confirm that the proposed method performs effectively and robustly, and exhibits superior or compatible performance over the other compared methods. Yongtao Yu, Haiyan Guan, Dilong Li, Shenghua Jin, Taiyue Chen, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Road Manhole Cover Delineation Using Mobile Laser Scanning Point Cloud DataabstractPeriodical road manhole cover measurement is extremely important to ensure road safety and reduce traffic disasters. This letter proposes an effective method for delineating road manhole covers from mobile laser scanning point cloud data. To improve processing efficiency, first, road surface points are segmented and rasterized into georeferenced intensity images. Then, object-oriented patches are generated through superpixel segmentation and further fed to a convolutional capsule network classifier for manhole cover detection. Finally, manhole covers are accurately delineated through a marked point process of disks. Quantitative evaluations on three data sets show that an average completeness, correctness, quality, and F1-measure of 0.965, 0.961, 0.929, and 0.963, respectively, are obtained. Comparative studies with three existing methods confirm that the proposed method performs superiorly in delineating manhole covers of varying conditions and on complex road surface environments. Yongtao Yu, Haiyan Guan, Dilong Li, Cheng Wang 0003, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Vehicle Detection From High-Resolution Remote Sensing Imagery Using Convolutional Capsule NetworksabstractVehicle detection plays an important role in a variety of traffic-related applications. However, due to the scale and orientation variations and partial occlusions of vehicles, it is still challengeable to accurately detect vehicles from remote sensing images. This letter proposes a convolutional capsule network for detecting vehicles from high-resolution remote sensing images. First, a test image is segmented into superpixels to generate meaningful and nonredundant patches. Then, these patches are input to a convolutional capsule network to label them into vehicles or the background. Finally, nonmaximum suppression is adopted to eliminate repetitive detections. Quantitative evaluations on four test data sets show that average completeness, correctness, quality, and F1-measure of 0.93, 0.97, 0.90, and 0.95, respectively, are obtained. Comparative studies with three existing methods confirm that the proposed method effectively performs in detecting vehicles of various conditions. Yongtao Yu, Tiannan Gu, Haiyan Guan, Dilong Li, Shenghua Jin |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2015 | Evaluation of regional-scale snow albedo characteristics during winter season from 2003 to 2014abstractSnow is a very important component of the climate system. It can influence the energy budget of the atmosphere and hydrological system significantly. The main goal of this paper is to use remote sensing and geographical information system techniques to analyze the spatial and temporal variations in regional scale and to find the relations between meteorological parameters and the snow albedo for the future modeling in snow albedo study. The results revealed spatial and temporal variation throughout different months during the winter season. In addition, the Pearson correlation coefficient analysis showed partial correlation between snow albedo and meteorological variables, which can be used to model snow albedo in some hydrological studies. Jonathan Li 0001, Claude R. Duguay, Dilong Li |
IGARSS | 4 |
| 2014 | Semisupervised Classification for Hyperspectral Imagery With Transductive Multiple-Kernel LearningabstractThe classification of hyperspectral imagery is a challenging problem because few labeled pixels are available. In this letter, we propose a new semisupervised learning algorithm to combine both cluster and manifold assumptions to increase classification reliability and accuracy. The new method uses a concave-convex procedure and sequential minimization optimization technologies for transductive multiple-kernel learning (TMKL). Then, a one-against-all strategy is adopted to generalize the binary TMKL classifiers to solve the multiclass problem of remote sensing images. Experimental results on two real data sets indicate that the proposed method exhibits both high accuracy and good computational performance. Cheng Wang 0003, Dilong Li, Jonathan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |