VLDB 2026 Research / reviewers in the wild / expert
Daxiang Li 0002
dblp:73/9638-2
· DBLP profile ↗
13ranked-venue papers
12as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical semantic alignment heterogeneous knowledge distillation model for smart agriculture crop leaf disease recognition
Daxiang Li 0002, Ying Liu 0026 |
Expert Syst. Appl. | 1 |
| 2025 | STRM-KD: Semantic topological relation matching knowledge distillation model for smart agriculture apple leaf disease recognition
Daxiang Li 0002, Ying Liu 0026 |
Expert Syst. Appl. | 1 |
| 2025 | Mamba-Wavelet Cross-Modal Fusion Network With Graph Pooling for Hyperspectral and LiDAR Data Joint ClassificationabstractRecently, with the rapid development of deep learning, the collaborative classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) image has become a research hotspot in remote sensing (RS) technology. However, existing methods either only consider complementary learning of spatial-domain information, or do not take into account the intrinsic dependencies between pixels and overlook the importance difference of pixels. In this letter, we propose a Mamba-Wavelet Cross-Modal Fusion Network with Graph Pooling (MW-CMFNet) for HSI and LiDAR joint classification. First, a Two-Branch Feature Extraction (TBFE) is used to extract spatial and spectral features. Then, in order to dig deeper into the complementary information of different modalities and fully fuse them under the guidance of frequency-domain information, a Mamba-Wavelet Cross-Modal Feature Fusion (MW-CMFF) Module is devised, it aims to utilize Mamba’s outstanding long-range modeling ability to learn complementary information in the spatial and frequency domains, Finally, the Graph Pooling module is designed to sense the intrinsic dependencies of neighbouring pixels and explore the importance difference of pixels, rather than assigning the same weight to different pixels. Experiments on the Houston2013 and Trento datasets show that the MW-CMFNet achieves higher classification accuracy compared to other state-of-the-art methods. Daxiang Li 0002, Bingying Li, Ying Liu 0026 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Mamba Cross-Modal Information Fusion Self-Distillation Model for Joint Classification of LiDAR and Hyperspectral DataabstractRecent studies have found that compared to single-modal data, the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) multimodal data can utilize their complementary information to further improve the accuracy of land-cover classification. However, due to the significant differences between multimodal data, the complementarity among them is difficult to be fully exploited and utilized, and the features after fusion are not refined and optimized, which limits the further improvement of land-cover classification accuracy. To alleviate these issues, a novel Mamba Cross-Modal Information Fusion Self-Distillation (Mb-CMIFSD) model is designed. Specifically, Mb-CMIFSD first uses conventional convolutional neural networks (CNN) to transform each patch into a token sequence. Second, a Mamba Cross Modal Information Fusion (MCMIF) module is developed to combine cross-modal attention with bidirectional Mamba mechanism, which can better explore the complementarity of multimodal remote sensing (RS) data and obtain more discriminative multimodal fusion features. Finally, a Prototype Constrained Self-Distillation (PCSD) module is designed to utilize the constructed prototype orthogonal regularization knowledge distillation function to further refine cross-modal fusion features, thereby enhancing the robustness and adaptability of feature extraction. The experimental results on three benchmark HSI and LiDAR datasets show that the designed Mb-CMIFSD model has higher classification accuracy compared to other state-of-the-art methods, and the ablation experiments also confirm the positive effect of the designed two key modules. Daxiang Li 0002, Bingying Li, Ying Liu 0026 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Micro-expression recognition based on a novel GCN-transformer cooperation model for IoT-eHealth
Daxiang Li 0002, Nannan Qiao, Ying Liu 0026 |
Expert Syst. Appl. | 1 |
| 2024 | Image recognition based on lightweight convolutional neural network: Recent advancesabstractImage recognition is an important task in computer vision with broad applications. In recent years, with the advent of deep learning, lightweight convolutional neural network (CNN) has brought new opportunities for image recognition, which allows high-performance recognition algorithms to run on resource-constrained devices with strong representation and generalization capabilities. This paper first presents an overview of several classical lightweight CNN models. Then, a comprehensive review is provided on recent image recognition techniques using lightweight CNN. According to the strategies applied to optimize image recognition performance, existing methods are classified into three categories: (1) model compression, (2) optimization of lightweight network, and (3) combining Transformer with lightweight network. In addition, some representative methods are tested on three commonly used datasets for performance comparison. Finally, technical challenges and future research trends in this field are discussed. Ying Liu 0026, Jiahao Xue, Daxiang Li 0002, Weidong Zhang 0005, Tuan Kiang Chiew, Zhijie Xu |
Image Vis. Comput. | 3 |
| 2024 | HFSI-TF: Hierarchical Full-Scale Interactive Transformer Model for Object Detection in Remote Sensing ImageabstractTransformer-based object detection models usually adopt an encoding-decoding architecture that mainly combines self-attention (SA) and multilayer perceptron (MLP). Although this architecture does not require nonmaximum suppression (NMS) and can really achieve end-to-end object detection, it also suffers from the disadvantage of insufficient multiscale object perception in the image, which leads to low accuracy in detecting small objects. Focusing on these issues, a new full-scale bidirectional interactive attention (FSBDIA) mechanism is constructed, thereby a novel hierarchical full-scale interactive transformer (HFSI-TF) model is designed for object detection in remote sensing image (RSI). First, in order to enhance the multiscale perception ability of the model, the FSBDIA mechanism is designed under the guidance of full-scale information. Then, based on FSBDIA, a hierarchical HFSI-TF encoder is constructed to interactively fuse multilayer feature maps layer by layer, thereby obtaining multiscale encoded features of RSI. Finally, a mixed cross attention (MCA) mechanism is also constructed, and an iterative decoding architecture is designed based on it to improve the accuracy of small object detection. Comparative experiments based on two benchmark datasets (i.e., DIOR and HRSC2016) show that the designed HFSI-TF model can effectively improve the accuracy of object detection in RSI, and the model we designed has superior performance compared to other state-of-the-art methods. Daxiang Li 0002, Bingying Li, Ying Liu 0026 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | PSCLI-TF: Position-Sensitive Cross-Layer Interactive Transformer Model for Remote Sensing Image Scene ClassificationabstractIn the scene classification task of remote sensing image (RSI), in order to fully perceive multi-scale local objects in the image and explore their interdependencies to mine the scene semantics of RSI, this letter designs a novel Position-Sensitive Cross-Layer Interactive Transformer (PSCLI-TF) model to improve the accuracy of RSI scene classification. Firstly, ResNet50 is utilized as the backbone to extract the multi-layer feature maps of RSI. Then, in order to enhance the model’s position sensitivity to local objects in RSI, a new Position-Sensitive Cross-Layer Interactive Attention (PSCLIA) mechanism is designed, and based on it a novel PSCLI-TF encoder is constructed to perform layer-by-layer interactive fusion on the multi-layer feature maps to obtain the multi-granularity Cross-Layer Fusion (CLF) feature of RSI. Finally, a prototype-based self-supervised loss function is constructed to alleviate the semantic gap problem of "large intra-class variance and small inter-class variance" in RSI scene classification. Comparative experimental results based on three datasets (i.e., AID, NWPU and UCM) indicate that the classification performance of the designed PSCLI-TF model is highly competitive compared to other state-of-the-art methods. Daxiang Li 0002, Runyuan Liu, Ying Liu 0026 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Shoeprint Image Retrieval Based on Dual Knowledge Distillation for Public Security Internet of ThingsabstractIn order to implement rapid retrieval of large-scale crime scene investigation shoeprint image (SPI) in the intelligent mobile terminal of Public Security Internet of Things (PSIoT), a novel dual knowledge distillation (DKD) network is designed by fusing spatial attention (SA) distillation and feature distillation to solve its problems of limited computing and storage capacity. First, ResNet50 was modified by adding a new SA module and hash layer as the teacher model, and a light convolutional neural network (CNN) with the same SA and hash layer is designed as the student model. Then, the attention distillation loss function is constructed to distill the SA knowledge in the teacher module to the student module to improve the ability of the convolutional layer at the front end of the student network to capture the underlying visual features of the SPI. Finally, the feature distillation loss function is constructed to distill the semantic knowledge in the teacher module to the student module to improve the ability of the student module to express high-level semantics of the SPI. we compare our method on two data sets of SPID and FID-300 with other state-of-the-art methods in the SPI retrieval domain. The experimental results show that our method can improve the SPI retrieval baseline by a large margin and better than other methods. Daxiang Li 0002, Yang Li 0161, Ying Liu 0026 |
IEEE Internet Things J. | 1 |
| 2022 | Remote Sensing Image Scene Classification Model Based on Dual Knowledge DistillationabstractIn the application of remote sensing image (RSI) scene classification, in order to solve the contradiction between the accuracy of Convolutional Neural Network (CNN) and the large amount of model parameters, a novel dual knowledge distillation (DKD) model combining dual attention (DA) and spatial structure (SS) is designed. First, new DA and SS modules are constructed and introduced into ResNet101 and light-weight CNN designed as teacher and student networks respectively. Then, in order to improve its local feature extraction and high-level semantic representation abilities for RSI by transmission the DA and SS knowledge in the teacher network to the student network, we design the corresponding DA and SS distillation losses. The comparative experimental results based on AID and NWPU-45 datasets show that when the training ratio is 20%, the accuracy of the student network after DKD is improved by 7.57% and 7.28% respectively, and in the case of fewer parameters, DKD has higher accuracy than most other methods. Daxiang Li 0002, Yixuan Nan, Ying Liu 0026 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Pornographic images recognition based on spatial pyramid partition and multi-instance ensemble learning
Daxiang Li 0002, Jing Wang 0033, Tingge Zhu |
Knowl. Based Syst. | 1 |
| 2014 | Pyramid Match Kernel and Classifier Ensemble-Based MIL Algorithm for Pornographic Images FilteringabstractIn this paper, a novel multi-instance learning (MIL) algorithm based on pyramid match kernel (PMK) and classifier ensemble is proposed for recognizing pornographic scene from image database. First, an improved JSEG image segmentation technique is deployed for dividing every image into several regions, and regards the whole image as a "bag", the low-level visual features (i.e. color and texture) of each segmented region as "instance". As a result, the pornographic images filtering problem can be transferred into a typical MIL problem. Second, similarity between the multi-instance bags is measured by PMK method, which allows MIL problem to be solved directly by the support vector machine (SVM). Finally, many base classifiers based on PMK with different levels are constructed, and the performance weighting rule is used to dynamically determine the weights of them, so the strategy of classifier ensemble is used to improve the filtering accuracy. In a real condition image set that the ratio of normal image to pornographic image is 9:1, experimental results show that the proposed algorithm, named PMKCE-MIL, is robust, and its performance is superior to other algorithms. Daxiang Li 0002, Jing Wang 0033, Ying Liu 0026 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2014 | Multiple kernel-based multi-instance learning algorithm for image classification
Daxiang Li 0002, Jing Wang 0033, Ying Liu 0026, Dianwei Wang |
J. Vis. Commun. Image Represent. | 1 |