VLDB 2026 Research / reviewers in the wild / expert
Chao Li 0066
dblp:66/190-66
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-1932-7698ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WildMDT-YOLO: A multi-scale wildlife detection model for complex environmentsabstractThe deterioration of the global ecological environment and increasing human activities pose severe threats to wildlife survival, making reliable detection methods crucial for wildlife protection and monitoring. However, existing detection methods often encounter inadequate feature extraction in complex settings such as tree occlusion, strong illumination, and low-light environments. To address this challenge, this paper proposes WildMDT-YOLO, an improved YOLOv8n-based wildlife detection method. The model introduces a Multi-scale Focusing Diffusion Network (MFDN) that enhances contextual information across scales through feature focusing and diffusion mechanisms. A novel detection head employs shared convolutions to reduce model parameters while task alignment and interactive feature extraction improve both classification and localization accuracy. The integration of Deformable Convolutional Networks v3 (DCNv3) and Mixed Local Channel Attention (MLCA) mechanism further enhances adaptability to wildlife species with complex shapes and varying scales. Experimental results show that WildMDT-YOLO achieves 92.6% mean average precision (mAP), a 3.1% increase over baseline YOLOv8n, while reducing parameters by 17.1%. Cross-dataset evaluation on Snapshot Serengeti demonstrates robust generalization capability, achieving 89.2% mAP and maintaining 83.6% small object detection performance despite significant domain shift between ecosystems. This model provides an effective tool for improving monitoring efficiency and accuracy in wildlife conservation. Chao Li 0066, Youbo Pang, Xianhang Liu, Minchao Sun |
Comput. Vis. Image Underst. | 1 |
| 2026 | MLaVQA: A multi-level attention method for remote sensing visual question answering with large language model
Weipeng Jing 0001, Wanlin Yang, Chao Li 0066, Mahmoud Emam |
Inf. Sci. | 3 |
| 2026 | Energy-Efficient Federated Learning With Dynamic Model Pruning for Industrial IoTabstractWith the advent of the Industry 4.0 era, Federated Learning (FL) provides robust data privacy protection for smart manufacturing and supply chain optimization, while facilitating collaborative intelligent optimization across enterprises and devices. However, the complex and overparameterized deep neural networks used in FL result in significant computational overhead for Industrial Internet of Things (IIoT) devices, leading to low energy efficiency and hindering the practical deployment of FL on IIoT devices. Moreover, the widespread data and device heterogeneity in the IIoT exacerbates the decrease in energy efficiency caused by inconsistent computational efficiency across nodes. This article proposes an energy-efficient dynamic model pruning method for FL, named EDPrune-FL, to address the aforementioned challenges. Compared to existing methods, this approach offers greater flexibility and efficiency by utilizing a dynamic pruning rate allocation mechanism. This mechanism updates the pruning rate for each participating client in every communication round, allowing the pruning upper bound to adapt to the varying importance of different learning stages in FL. EDPrune-FL ensures the global model’s performance while reducing the training energy consumption of clients in heterogeneous environments. To guarantee that dynamic pruning maintains the stability and effectiveness of the model in heterogeneous environments, we also demonstrated the convergence of EDPrune-FL and discussed the relationship between pruning rates and convergence, providing a qualitative analysis. Experimental results demonstrate that our method outperforms the state-of-the-art technique across four real-world datasets. With tests conducted on 100 clients, our approach reduces energy consumption by 10% while maintaining comparable accuracy. Guangsheng Chen, Fangyu Sun, Weitao Zou, Chao Li 0066, Yipeng Zhou, Moule Lin, Peng Liu 0023, Linkang Geng, Lei Fan 0007, Weipeng Jing 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Prototype-Aligned Federated Learning for Robust Object Extraction in Heterogeneous Remote SensingabstractFederated learning (FL) has emerged as a pivotal collaborative machine learning framework, enabling privacy-preserving analytics for smart city applications using distributed data from Internet of Things (IoT) devices. However, the inherent data heterogeneity that arises from diverse geographical and environmental factors poses significant challenges to the effectiveness of FL-based models. To address these challenges, this paper introduces a novel Prototype-Based FL framework for cross-domain object extraction in heterogeneous remote sensing images. The proposed framework employs multiple vectors to represent class prototypes for capturing the intricate intra-class variations and mitigating the adverse effects of non-identically distributed (non-IID) data across clients. Furthermore, we adopt a distance-based classification method to reduce classification errors. Additionally, we propose a Prototype-Anchored Metric Learning approach to minimize intra-class variance and enhance inter-class separability, which can facilitate the alignment of feature representations across heterogeneous datasets. The proposed method improves the coherence and stability of feature spaces in federated settings and enhances the global model’s generalization capabilities for complex urban monitoring tasks. Extensive experiments on three distinct remote sensing datasets(including infrastructure and disaster) demonstrate that the proposed method significantly outperforms state-of-the-art FL-based approaches in urban monitoring accuracy and robustness. The code is available at Guangsheng Chen, Ye Yuan 0011, Moule Lin, Lianchong Zhang, Chao Li 0066, Weitao Zou, Weipeng Jing 0001, Mahmoud Emam |
IEEE Internet Things J. | 6 |
| 2025 | Fine-grained forest net primary productivity monitoring: Software system integrating multisource data and smart optimizationabstractAbstract Net primary productivity (NPP) is essential for sustainable resource management and conservation, and it serves as a primary monitoring target in smart forestry systems. The predominant method for NPP inversion involves data collection through terrestrial and satellite sensing systems, followed by parameter estimation using models such as the Carnegie‐Ames‐Stanford Approach (CASA). While this method benefits from low costs and extensive monitoring capabilities, the data derived from multisource sensing systems display varied spatial scale characteristics, and the NPP inversion models cannot detect the impact of data heterogeneity on the outcomes sensitively, reducing the accuracy of fine‐grained NPP inversion. Therefore, this paper proposes a modular system for fine‐grained data processing and NPP inversion. Regarding data processing, a two‐stage spatial‐spectral fusion model based on non‐negative matrix factorization (NMF) is proposed to enhance the spatial resolution of remote sensing data. A spatial interpolation model based on stacking generalization with residual correction is introduced to get raster meteorological data compatible with remote sensing images. Furthermore, we optimize the CASA model with the kernel method to enhance model sensitivity and enrich the spatial details of the inversion results with high resolution. Through validation using real datasets, the proposed fusion and interpolation models have significant advantages over mainstream methods. Furthermore, the correlation coefficient () between the estimated NPP using our improved inversion model and the field‐measured NPP is 0.69, demonstrating the feasibility of this platform in detailed forest NPP monitoring tasks. Weitao Zou, Long Luo, Fangyu Sun, Chao Li 0066, Guangsheng Chen, Weipeng Jing 0001 |
Softw. Pract. Exp. | 4 |
| 2025 | Hypergraph BiFormer for Semantic Segmentation of High-Resolution Remote Sensing ImagesabstractWhile transformers are powerful neural network architectures for feature learning, current Transformer-based approaches for semantic segmentation of high-resolution remote sensing images (HRRSIs) struggle with the extraction of local semantic features. To address this issue, we incorporate a hypergraph into the Transformer. Hypergraph-based methods are proficient at discovering high-order correlations within limited-scale data, extracting pertinent representations to enhance the Transformer’s learning capabilities. We also propose dual pooling and feature aggregation modules (FAMs), inspired by the adaptive pooling’s potent local modeling capabilities, to additionally extract fine-grained features from HRRSIs. In particular, we conceive a hypergraph BiFormer (HGBT) based on these three proposed modules along with a BiFormer backbone. HGBT has the potential to learn general latent features as well as generate high-order representations of HRRSIs by modeling correlations of multiscale features and local topology within an entirely nonlinear space, leading to the aggregation of features in a compact and localized manner, enhancing the model’s ability to capture detailed variations within small areas. We validate our approach through extensive experiments on ISPRS Vaihingen and Potsdam datasets, where HGBT attains mean intersection over union (mIoU) of 83.71% and 87.88%, respectively. Both quantitative and qualitative assessments underscore the dominance of HGBT. Our code will be accessible at:https://github.com/ZhangIceNight/HGBFormer. Weipeng Jing 0001, Donglin Di, Chao Li 0066, Mahmoud Emam, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Gaussian-Based Swap Operator for Context-Aware Extraction of Building Boundary VectorsabstractAccurate extraction of building vector boundaries holds paramount importance within the domains of urban planning and Geographic Information Systems (GIS), providing indispensable support for urban construction endeavors and resource management initiatives. CNNs, while proficient in local feature extraction, often falter in capturing holistic, global image characteristics. Transformers excel in contextual feature comprehension but demand substantial computational resources and parameterization, impeding practical deployment. To address these challenges, this paper introduces an innovative computational operator known as G-Swap, which integrates Gaussian-distance-based feature correlation considerations, thereby significantly augmenting contextual comprehension within the computational framework. Additionally, a universal architecture for boundary vector extraction is proposed in this paper, comprising three primary components: 1) an Enhanced Backbone, integrating the G-Swap operator to enhance the backbone while bolstering model expressiveness; 2) a Decoder module, tasked with discriminating corner and edge features; and 3) a Two-branch Detection Head. Empirical experiments conducted on the Vectorizing World Building Dataset (VWB) underscore the model’s superior performance. Our G-Swap achieved F1 scores of 91.2% for vertices and 80.1% for edges, surpassing the previous state-of-the-art by 2.1% and 2.0% respectively. Moule Lin, Weipeng Jing 0001, Weitao Zou, Zhongwei Qiu, Chao Li 0066 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | RSVMamba for Tree Species Classification Using UAV RGB Remote Sensing ImagesabstractEffective forest tree species (TS) classification is critical for various application domains such as forest management, biodiversity conservation, and ecological research. However, existing studies on TS classification predominantly rely on high-cost and processing-intensive hyperspectral data, which limits practical applications on large scales. In this work, we focus on investigating the potential of cost-effective unmanned aerial vehicle (UAV) RGB images for TS classification in heterogeneous forests and propose a method that fully leverages the rich spatial, semantic, and visible spectral information of UAV RGB images. We propose an RSVMamba model, which incorporates improved visual state-space (VSS) blocks and an AutoDownsampling module to enhance accuracy and stability while paying particular attention to small objects in sparse spatial locations. The model achieves linear computational complexity while retaining the global receptive field, making it particularly suitable for processing high spatial-resolution images. Additionally, we collected UAV RGB images covering$40~\text {km}^{2}$of subtropical forest in southern China. A meticulous evaluation of this data shows that our method achieves an overall accuracy (OA) of 84.28% for eight TS, dead trees, and other broadleaves. We verify the superiority of our method through a series of comparative experiments on the collected and benchmark datasets. Our results affirm the usefulness of single-temporal UAV RGB images for TS classification in heterogeneous forest environments. Furthermore, the proposed method bridges the gap between data accessibility and precision in TS classification, broadening the boundaries of single-temporal UAV RGB images for practical forestry applications and providing a more cost-effective and time-flexible solution for this problem. Juntao Gu, Basim Azam, Moule Lin, Chao Li 0066, Weipeng Jing 0001, Naveed Akhtar |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | SRF: SpectrumRecombineFormer for Hyperspectral Image ClassificationabstractHyperspectral imaging is a valuable technique for accurately classifying materials because of the abundance of spectral information and high resolution it provides. However, the characteristics of Hyperspectral Imaging, such as high-dimensional features and information redundancy, pose significant challenges to data processing. Traditional dimensionality reduction methods often have information loss, high computational complexity, and easy to ignore the strong correlation between HSI bands when dealing with the HSI data. Although other methods can achieve satisfactory classification performance, they do not consider the dimensionality reduction of HSI, and they focus on the model performance, which limits further improvement in classification performance. This article proposes a transformer-based framework called “SpectrumRecombineFormer” (SRF), which is composed of two key modules, namely “Spatial–Spectral Recombination” (SSRC) and “Cross-Layer Fusion” (CF). The SSRC is capable of utilizing both adjacent and non-adjacent spectrums to generate the spatial-sequential perceptive representations, which alleviate the effect of the strong correlation between HSI bands. The CF can avoid the loss of information during the feed-forward procedure among layers. Extensive experiments on five existing datasets (widely adopted Indian Pines, Houston2013, Pavia University, Salinas, and KSC) demonstrate the capability of our proposed method to address the above-mentioned challenges. Both quantitative and qualitative experimental ablation studies, including visualization results, reveal that the proposed SRF method can successfully and efficiently classify HSIs and surpass the other state-of-the-art methods. For access to the source code, please visit https://github.com/kangpeilun/SRF-HSI-Classification-master . Weipeng Jing 0001, Peilun Kang, Donglin Di, Juntao Gu, Mahmoud Emam, Linda F. Mohaisen, Xun Yang 0001, Chao Li 0066 |
ACM Trans. Multim. Comput. Commun. Appl. | 9 |
| 2024 | ESNet: Perceptive Spatial-Spectral Fusion with Multi-stage Reconstruction for Pansharpening
Chao Li 0066, Juntao Gu, Moule Lin, Weipeng Jing 0001 |
ADMA (3) | 1 |
| 2024 | Adaptive Global-local Fusion Network Based Deep Unsupervised Hashing for Remote Sensing Image RetrievalabstractUnsupervised hashing methods have gained wide-spread popularity for remote sensing (RS) image retrieval due to their high efficiency. Existing methods heavily rely on similarity matrix generated by pre-trained models as supervised signals. However, such pre-trained models obtained from natural images fail to comprehensively extract features in RS images, yielding unreliable similarity relationships. Considering complex features in RS images, we propose a novel adaptive global-local fusion network based deep unsupervised hashing (AFDUH) method. The advantages of AFDUH lie in the use of large kernel convolution and bi-level routing attention mechanism for learning local and global features simultaneously. AFDUH fuses these features under different resolutions in an interactive fashion to deduce high-quality similarity matrix based on contrastive learning. Besides, we embed a sample selection strategy in AFDUH, which can filter out unreliable supervised signals to further improve retrieval accuracy. Extensive experiments on two RS datasets demonstrate that AFDUH outperforms the state-of-the-art baselines. Yipeng Zhou, Quan Z. Sheng, Chao Li 0066, Tongtong Lou, Weipeng Jing 0001 |
ICME | 4 |
| 2024 | BT-YOLO: Improved YOLOv5 Based on BiFormer Structure and Task-Specific Decoupled Head for Photovoltaic Infrared Defect Detection on UAV ScenariosabstractPrevious studies have demonstrated the importance of combining infrared defect detection methods with UAV inspection to promote the development of solar energy. The defect detection frequently encounters difficulties such as small objects, easy confusion between defects and environment, and uneven sample number, resulting in a low detection accuracy. To solve these problems, this paper proposed a model BT-YOLO to detect infrared photovoltaic images captured by UAV based on the YOLOv5 network. Firstly, BiFormer is a visual transformer structure embedded into the backbone network of the model, better preserving fine-grained details. Secondly, to achieve more accurate classification and finer localization, the feature encoding of classification and localization is decoded separately in the detection head. Finally, the regression loss is calculated using Wise-IoU instead of GIoU, thereby allowing the model to note the loss of ordinary-quality anchor boxes. The results demonstrate that the improved model improves the mAP performance by 5.1%. Weipeng Jing 0001, Baihong Guo, Peilun Kang, Mahmoud Emam, Chao Li 0066 |
MSN | 7 |
| 2024 | HGSNet: A hypergraph network for subtle lesions segmentation in medical imagingabstractAbstract Lesion segmentation is a fundamental task in medical image processing, often facing the challenge of subtle lesions. It is important to detect these lesions, even though they can be difficult to identify. Convolutional neural networks, an effective method in medical image processing, often ignore the relationship between lesions, leading to topological errors during training. To tackle topological errors, move is made from pixel‐level to hypergraph representations. Hypergraphs can model lesions as vertices connected by hyperedges, capturing the topology between lesions. This paper introduces a novel dynamic hypergraph learning strategy called DHLS. DHLS allows for the dynamic construction of hypergraphs contingent upon input vertex variations. A hypergraph global‐aware segmentation network, termed HGSNet, is further proposed. HGSNet can capture the key high‐order structure information, which is able to enhance global topology expression. Additionally, a composite loss function is introduced. The function emphasizes the global aspect and the boundary of segmentation regions. The experimental setup compared HGSNet with other advanced models on medical image datasets from various organs. The results demonstrate that HGSNet outperforms other models and achieves state‐of‐the‐art performance on three public datasets. Junze Wang, Chao Li 0066, Weipeng Jing 0001 |
IET Image Process. | 4 |
| 2024 | Optimized Vectorizing of Building Structures With Switch: High-Efficiency Convolutional Channel-Switch Hybridization StrategyabstractThe building planar graph reconstruction, a.k.a. footprint reconstruction, which lies in the domain of computer vision and geoinformatics, has been long afflicted with the challenge of redundant parameters in conventional convolutional models. Therefore, in this letter, we proposed an advanced and adaptive shift architecture, the “Switch” operator, which incorporates nonexponential growth parameters while retaining analogous functionalities to integrate local feature spatial information, resembling a high-dimensional convolution operation. The “Switch” operator, cross-channel operation, architecture implements the XOR operation to exchange adjacent or diagonal features alternately and then blends alternating channels through a$1 \times 1$convolution operation to consolidate information from different channels. The SwitchNN architecture, on the other hand, incorporates a group-based parameter-sharing mechanism inspired by the convolutional neural network (CNN) process, thereby significantly reducing the number of parameters. We validated our proposed approach through experiments on the SpaceNet corpus. Our method achieves 82.9% precision, 79.8% F1 score, 83.7% recall, and an MAE of 0.018 with the FLOPs of 24.83 G, outperforming existing state-of-the-art methods, such as Roof-Former and HEAT. These results demonstrate the effectiveness of this innovative architecture in building planar graph reconstruction from 2-D building images. Moule Lin, Weipeng Jing 0001, Chao Li 0066, András Jung |
IEEE Geosci. Remote. Sens. Lett. | 3 |