Daixun Li

dblp:361/0252 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
19since 2021 · last 2026
0009-0006-8689-6929ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MMFormer: Multi-Modality semi-Supervised vision transformer in remote sensing imagery classification
Daixun Li, Weiying Xie, Leyuan Fang, Yunke Wang, Mingxiang Cao, Jitao Ma, Yunsong Li 0001, Chang Xu 0002
Neural Networks1
2026 FA-Mamba: frequency attention driven Mamba for multimodal remote sensing classification
Danian Yang, Daixun Li, Jitao Ma, Yibing Lu, Yunsong Li 0001, Leyuan Fang, Weiying Xie
Neural Networks2
2026 LoME: LoRA-Driven Multimodal Extractor for RGB-X Vision Tasks
abstract
RGB-X multimodal vision tasks present a highly promising approach to enhancing model performance in complex visual conditions. Existing multimodal frameworks are based on either the symmetric parallel network of feature fusion or the shared network of input fusion. However, parallel networks suffer from uncontrollable parameters and imbalanced optimization across modal branches, while shared networks often lead to a lack of diversity in gradient optimization. To address these challenges, we propose the LoRA-driven Multimodal Extractor (LoME), following a comprehensive analysis of existing multimodal frameworks. The low-rank properties of modal adapters for LoME ensure controllable growth in model parameters as the number of modalities increases. The dynamic parameter fusion between adapters and the shared feature extractor decouples gradient optimization directions, effectively mitigating imbalances caused by multimodal data biases while preserving complementary features. Moreover, we employ a training strategy based on dynamic rank allocation to reduce computational overhead and enhance modal diversity expression. We validate the effectiveness and generalizability of LoME across three multimodal vision tasks. LoME achieves superior performance compared to previous state-of-the-art methods on multiple datasets. For example, on the DroneVehicle dataset, our method achieves a 10.4% improvement in accuracy compared to the SOTA method, while the parameter overhead is reduced to 23% of the previous network (44.63M). The code has been open-sourced at https://github.com/zyszxhy/LoME.
Weiying Xie, Tianlin Hui, Daixun Li, Jie Lei 0001, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Aligning and Prompting Anything for Zero-Shot Generalized Anomaly Detection
abstract
Zero-shot generalized anomaly detection (ZGAD) plays a critical role in industrial automation and health screening. Recent studies have shown that ZGAD methods built on visual-language models (VLMs) like CLIP have excellent cross-domain detection performance. Different from other computer vision tasks, ZGAD needs to jointly optimize both image-level anomaly classification and pixel-level anomaly segmentation tasks for determining whether an image contains anomalies and detecting anomalous parts of an image, respectively, this leads to different granularity of the tasks. However, existing methods ignore this problem, processing these two tasks with one set of broad text prompts used to describe the whole image. This limits CLIP to align textual features with pixel-level visual features and impairs anomaly segmentation performance. Therefore, for precise visual-text alignment, in this paper we propose a novel fine-grained text prompts generation strategy. We then apply the broad text prompts and the generated fine-grained text prompts for visual-textual alignment in classification and segmentation tasks, respectively, accurately capturing normal and anomalous instances in images. We also introduce the Text Prompt Shunt (TPS) model, which performs joint learning by reconstruction the complementary and dependency relationships between the two tasks to enhance anomaly detection performance. This enables our method to focus on fine-grained segmentation of anomalous targets while ensuring accurate anomaly classification, and achieve pixel-level comprehensible CLIP for the first time in the ZGAD task. Extensive experiments on 13 real-world anomaly detection datasets demonstrate that TPS achieves superior ZGAD performance across highly diverse datasets from industrial and medical domains.
Jitao Ma, Weiying Xie, Hangyu Ye, Daixun Li, Leyuan Fang
AAAI4
2025 FedCS: Coreset Selection for Federated Learning
abstract
Federated Learning (FL) is an emerging direction in distributed machine learning that enables jointly training a model without sharing the data. However, as the size of datasets grows exponentially, computational costs of FL increase. In this paper, we propose the first Coreset Selection criterion for Federated Learning (FedCS) by exploring the Distance Contrast (DC) in feature space. Our FedCS is inspired by the discovery that DC can indicate the intrinsic properties inherent to samples regardless of the networks. Based on the observation, we develop a method that is mathematically formulated to prune samples with high DC. The principle behind our pruning is that high DC samples either contain less information or represent rare extreme cases, thus removal of them can enhance the aggregation performance. Besides, we experimentally show that samples with low DC usually contain substantial information and reflect the common features of samples within their classes, such that they are suitable for constructing coreset. With only two time of linear-logarithmic complexity operation, FedCS leads to significant improvements over the methods using whole dataset in terms of computational costs, with similar accuracies. For example, on the CIFAR-10 dataset with Dirichlet coefficient α = 0.1, FedCS achieves 58.88% accuracy using only 44% of the entire dataset, whereas other methods require twice the data volume as FedCS for same performance.
Chenhe Hao, Weiying Xie, Daixun Li, Hangyu Ye, Leyuan Fang, Yunsong Li 0001
CVPR3
2025 Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and Memory
Daixun Li, Mingxiang Cao, Donglai Liu, Weiying Xie, Tianlin Hui, Lunkai Lin, Yunsong Li 0001
ICCV1
2025 FusionSAM: Visual Multi-Modal Learning with Segment Anything Model
abstract
Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance during training. While the Segment Anything Model (SAM) allows precise control during fine-tuning through its flexible prompting encoder, its potential remains largely unexplored in the context of multimodal segmentation for natural images. In this paper, we introduce SAM into multimodal image segmentation for the first time, proposing a novel framework that combines Latent Space Token Generation (LSTG) and Fusion Mask Prompting (FMP) modules. This approach transforms the training methodology for multimodal segmentation from a traditional black-box approach to a controllable, prompt-based mechanism. Specifically, we obtain latent space features for both modalities through vector quantization and embed them into a cross-attention-based inter-domain fusion module to establish long-range dependencies between modalities. We then use these comprehensive fusion features as prompts to guide precise pixel-level segmentation. Extensive experiments on multiple public datasets demonstrate that our method significantly outperforms SAM and SAM2 in multimodal autonomous driving scenarios, achieving an average improvement of 4.1% over the state-of-the-art method in segmentation mIoU, and the performance is also optimized in other multi-modal visual scenes.
Daixun Li, Weiying Xie, Mingxiang Cao, Yunke Wang, Leyuan Fang, Yunsong Li 0001, Chang Xu 0002
KDD (2)1
2025 Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal Fusion
abstract
Vision-Language-Action (VLA) systems are crucial for autonomous decision-making in embodied intelligence. While current systems have advanced the instruction-following capabilities, their limited spatial perception often leads to suboptimal performance for mobile manipulation tasks in unstructured environments. To address this challenge, we propose Uni-Sight, an end-to-end VLA system for robust mobile manipulation. Uni-Sight unifies decision-making, perception, and control through joint training, enabling synchronized cross-component optimization. Within the system, we introduce Latent Feature Aligner (LFA) that ensures accurate target localization by aligning multi-view data. Specifically, we develop Domain Transfer Policy (DTP), a hierarchical policy constrained by LiDAR-guided spatial priors, which ensures 3D spatial understanding with limited visual coverage. Extensive experiments on 20 real-world mobile manipulation tasks demonstrate the high task success rate and robust execution performance of Uni-Sight. Our Uni-Sight achieves a 3.04× the success rate of existing methods, and exhibits superior generalization in both long-horizon and zero-shot scenes. Code and dataset are publicly available at https://github.com/trantor2nd/Uni-Sight.
Daixun Li, Sibo He, Jiayun Tian, Weiying Xie, Mingxiang Cao, Donglai Liu, Tianlin Hui, Yunsong Li 0001
ACM Multimedia1
2025 Semi-Mamba: Mamba-Driven Semi-Supervised Multimodal Remote Sensing Feature Classification
abstract
Mamba architecture achieves the same performance as attention mechanisms with linear complexity, leading to significant progress in remote sensing land cover classification. However, existing Mamba methods rarely leverage the representational complementarity and consistency between different modalities, resulting in challenges such as incomplete fusion. To address these issues, we propose Semi-Mamba, a novel semi-supervised framework specifically designed for high-dimensional multi-modal data fusion. We introduce the Mamba Cross-Modality Fusion Module, which enables cross-modal learning of temporal features through state-space model interactions and smooth integration of input matrices, enhancing the fusion of richer feature representations. Additionally, to tackle the inherent difficulty of acquiring pixel-level annotations in remote sensing datasets, we introduce a multi-modal semi-supervised mechanism. This mechanism utilizes cross-modal supervision between different modalities to maximize data utilization and improve learning efficiency. It effectively enables joint training on both labeled and unlabeled data without relying on pseudo-labels. We integrate these innovations into a unified end-to-end framework. Compared to state-of-the-art CNN and Transformerbased architectures, our framework shows a significant improvement of over 3.12%, setting a new benchmark for semi-supervised multi-modal data fusion. The code has open sourced at https://github.com/LDXDU/Semi_Mamba_RS.
Yunsong Li 0001, Daixun Li, Weiying Xie, Jitao Ma, Sibo He, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.2
2025 Dual-Depth Unified Joint Optimization: Adaptive Curvature-Based Compression
abstract
Model compression methods such as pruning and quantization have been proposed to facilitate the deployment of convolutional neural networks (CNNs) on resource-constrained devices. Existing methods aim to combine the two for simultaneous improvement in compression ratio and runtime efficiency. However, most of the joint methods adopt linear tandem structures. Due to the lack of a unified framework, different optimization directions result in suboptimal solutions, especially when the compression ratio is extremely high. In this paper, we propose a novel adaptive curvature-based compression (ACC) method, which achieves a dual-depth unified joint optimization of pruning and quantization. In the first depth, we unify the pruning and quantization criteria using mean curvature, which leverages the discrete nature of image data and the continuum theory of differential geometry. In the second depth, we replace the traditional training process in the joint pruning-quantization method with curvature-aware knowledge distillation (CKD), unifying the two-stage approach into a simple but powerful parallel step. Our method is effective and interpretable by utilizing inherent properties to promote the understanding of information distribution and the importance of feature maps. Extensive experiments on multiple advanced benchmarks and diverse downstream task datasets have validated the superiority and generalizability of our ACC. Notably, we can achieve a 1.05% Top-1 accuracy improvement over the baseline under an extreme compression ratio of 454.55×, outperforming existing state-of-the-art (SOTA) methods.
Yunsong Li 0001, Xin Zhang 0092, Weiying Xie, Daixun Li, Hangyu Ye, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.5
2025 MIFNet: Multi-Scale Interaction Fusion Network for Remote Sensing Image Change Detection
abstract
Change Detection (CD) is a crucial and challenging task in remote sensing observations. Despite the remarkable progress driven by deep learning in remote sensing change detection, several challenges remain regarding global information representation and efficient interaction. The traditional Siamese network structure, which extracts features from bitemporal images using a weight-sharing network and generates a change map, but often neglects phase interaction information between images. Additionally, multi-scale feature fusion methods frequently use FPN-like structures, leading to lossy cross-layer information transmission and hindering the effective utilization of features. To address these issues, we propose a multi-scale interaction fusion network (MIFNet) that fuses bitemporal features at an early stage, using deep supervision techniques to guide early fusion features in obtaining abundant semantic representation of changes, also we construct a dual complementary attention module (DCA) to capture temporal information. Furthermore, we introduce a collection-allocation fusion mechanism, which is different from previous layer-by-layer fusion methods since it collects global information and embeds features at different levels to achieve effective cross-layer information transmission and promote global semantic feature representation. Extensive experiments demonstrate that our method achieves competitive results on the LEVIR-CD+ dataset, outperforming other advanced methods on both the LEVIR-CD and SYSU-CD datasets, with F1 improved by 0.96% and 0.61%, respectively, compared to the most advanced models.
Weiying Xie, Wenjie Shao, Daixun Li, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.3
2024 MDFL: Multi-Domain Diffusion-Driven Feature Learning
abstract
High-dimensional images, known for their rich semantic information, are widely applied in remote sensing and other fields. The spatial information in these images reflects the object's texture features, while the spectral information reveals the potential spectral representations across different bands. Currently, the understanding of high-dimensional images remains limited to a single-domain perspective with performance degradation. Motivated by the masking texture effect observed in the human visual system, we present a multi-domain diffusion-driven feature learning network (MDFL) , a scheme to redefine the effective information domain that the model really focuses on. This method employs diffusion-based posterior sampling to explicitly consider joint information interactions between the high-dimensional manifold structures in the spectral, spatial, and frequency domains, thereby eliminating the influence of masking texture effects in visual models. Additionally, we introduce a feature reuse mechanism to gather deep and raw features of high-dimensional data. We demonstrate that MDFL significantly improves the feature extraction performance of high-dimensional data, thereby providing a powerful aid for revealing the intrinsic patterns and structures of such data. The experimental results on three multi-modal remote sensing datasets show that MDFL reaches an average overall accuracy of 98.25%, outperforming various state-of-the-art baseline schemes. Code available at https://github.com/LDXDU/MDFL-AAAI-24.
Daixun Li, Weiying Xie, Yunsong Li 0001
AAAI1
2024 FedSLS: Exploring Federated Aggregation in Saliency Latent Space
abstract
Federated Learning (FL) is an emerging direction in distributed machine learning that enables jointly training a global model without sharing data with server. However, data heterogeneity biases the parameter aggregation at the server, leading to slower convergence and poorer accuracy of the global model. To cope with this, most of the existing works involve enforcing regularization in local optimization or improving the model aggregation scheme at the server. Though effective, they lack a deep understanding of cross-client features. In this paper, we propose a saliency latent space feature aggregation method (FedSLS) across federated clients. By Guided BackPropagation (GBP), we transform deep models into powerful and flexible visual fidelity encoders, applicable to general state inputs across different image domains, and achieve powerful aggregation in the form of saliency latent features. Notably, since GBP is label-insensitive, it is sufficient to capture saliency features only once on each client. Experimental results demonstrate that FedSLS leads to significant improvements over the state-of-the-arts in terms of accuracies, especially in highly heterogeneous settings. For example, on CIFAR-10 dataset, FedSLS achieves 63.43% accuracy within the strongly heterogeneous environment α=0.05, which is 6% to 23% higher than other baselines.
Hengyi Wang, Weiying Xie, Jitao Ma, Daixun Li, Yunsong Li 0001
ACM Multimedia4
2024 E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection
abstract
Multimodal image fusion and object detection are crucial for autonomous driving. While current methods have advanced the fusion of texture details and semantic information, their complex training processes hinder broader applications. Addressing this challenge, we introduce E2E-MFD, a novel end-to-end algorithm for multimodal fusion detection. E2E-MFD streamlines the process, achieving high performance with a single training phase. It employs synchronous joint optimization across components to avoid suboptimal solutions associated to individual tasks. Furthermore, it implements a comprehensive optimization strategy in the gradient matrix for shared parameters, ensuring convergence to an optimal fusion detection configuration. Our extensive testing on multiple public datasets reveals E2E-MFD's superior capabilities, showcasing not only visually appealing image fusion but also impressive detection outcomes, such as a 3.9\% and 2.0\% $\text{mAP}_{50}$ increase on horizontal object detection dataset M3FD and oriented object detection dataset DroneVehicle, respectively, compared to state-of-the-art approaches.
Mingxiang Cao, Weiying Xie, Jie Lei 0001, Daixun Li, Wenbo Huang 0001, Yunsong Li 0001
NeurIPS5
2024 FedDiff: Diffusion Model Driven Federated Learning for Multi-Modal and Multi-Clients
abstract
With the rapid development of imaging sensor technology in the field of remote sensing, multi-modal remote sensing data fusion has emerged as a crucial research direction for land cover classification tasks. While diffusion models have made great progress in generative models and image classification tasks, existing models primarily focus on single-modality and single-client control, that is, the diffusion process is driven by a single modal in a single computing node. To facilitate the secure fusion of heterogeneous data from clients, it is necessary to enable distributed multi-modal control, such as merging the hyperspectral data of organization A and the LiDAR data of organization B privately on each base station client. In this study, we propose a multi-modal collaborative diffusion federated learning framework called FedDiff. Our framework establishes a dual-branch diffusion model feature extraction setup, where the two modal data are inputted into separate branches of the encoder. Our key insight is that diffusion models driven by different modalities are inherently complementary in terms of potential denoising steps on which bilateral connections can be built. Considering the challenge of private and efficient communication between multiple clients, we embed the diffusion model into the federated learning communication structure, and introduce a lightweight communication module. Qualitative and quantitative experiments validate the superiority of our framework in terms of image quality and conditional consistency. To the best of our knowledge, this is the first instance of deploying a diffusion model into a federated learning framework, achieving optimal both privacy protection and performance for heterogeneous data. Our FedDiff surpasses existing methods in terms of performance on three multi-modal datasets, achieving a classification average accuracy of 96.77% while reducing the communication cost.
Daixun Li, Weiying Xie, Yibing Lu, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.1
2024 Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification
abstract
In multimodal land cover classification (MLCC), a common challenge is the redundancy in data distribution, where task-irrelevant information from multiple modalities can hinder the effective integration of their unique features. To tackle this, we introduce the Multimodal Informative Vit (MIVit), a system with an innovative information aggregate-distributing mechanism. This approach redefines redundancy levels and integrates performance-aware elements into the fused representation, facilitating the learning of semantics in both forward and backward directions. MIVit stands out by significantly reducing redundancy in the empirical distribution of each modality’s separate and fused features. It employs oriented attention fusion (OAF) for extracting shallow local shape features across modalities in horizontal and vertical dimensions, and a Transformer feature extractor for extracting deep global features through long-range attention. We also propose an information aggregation constraint (IAC) based on mutual information, designed to remove redundant information and preserve complementary information within embedded features. Additionally, the information distribution flow (IDF) in MIVit enhances performance-awareness by distributing global classification information across different modalities’ feature maps. This architecture also addresses missing modality challenges with lightweight independent modality classifiers, reducing the computational load typically associated with Transformers. Our results show that MIVit’s bidirectional aggregate-distributing mechanism between modalities is highly effective, achieving an average overall accuracy of 95.56% across three multimodal datasets. This performance surpasses current state-of-the-art methods in MLCC. The code for MIVit is accessible at https://github.com/icey-zhang/MIViT.
Jie Lei 0001, Weiying Xie, Geng Yang 0001, Daixun Li, Yunsong Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 FedFusion: Manifold-Driven Federated Learning for Multi-Satellite and Multi-Modality Fusion
abstract
Multi-Satellite, multi-modality in-orbit fusion is a challenging task as it explores the fusion representation of complex high-dimensional data under limited computational resources. Deep neural networks can reveal the underlying distribution of multimodal remote sensing data, but the in-orbit fusion of multimodal data is more difficult because of the limitations of different sensor imaging characteristics, especially when the multimodal data follow nonindependent identically distribution (Non-IID) distributions. To address this problem while maintaining classification performance, this article proposes a manifold-driven multi-modality fusion framework, FedFusion, which randomly samples local data on each client to jointly estimate the prominent manifold structure of shallow features of each client and explicitly compresses the feature matrices into a low-rank subspace through cascading and additive approaches, which is used as the feature input of the subsequent classifier. Considering the physical space limitations of the satellite constellation, we developed a multimodal federated learning (FL) module designed specifically for manifold data in a deep latent space. This module achieves iterative updating of the subnetwork parameters of each client through global weighted averaging, constructing a framework that can represent compact representations of each client. The proposed framework surpasses existing methods in terms of performance on three multimodal datasets, achieving a classification average accuracy of 94.35% while compressing communication costs by a factor of 4. Furthermore, extensive numerical evaluations of real-world satellite images were conducted on the orbiting edge computing architecture based on Jetson TX2 industrial modules, which demonstrated that FedFusion significantly reduced training time by 48.4 min (15.18%) while optimizing accuracy. The codes will be available at:https://github.com/LDXDU/FedFusion.
Daixun Li, Weiying Xie, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.1
2024 Ebbinghaus-Curve Guided Low-Rank Component-Induced Attention for Multisource Remote Sensing Classification
abstract
The integration of multisource remote sensing (RS) data is crucial in land use and land cover (LULC) studies, offering numerous applications. Using diverse data sources enhances the accuracy of land cover classification. However, due to differences in imaging mechanisms, existing methods face challenges in capturing complex local and global relationships. Moreover, current multimodal fusion approaches often fail to efficiently preserve heterogeneous data, leading to the overfusion of redundant features. To address these challenges, we propose Ebbinghaus-curve guided multisource RS classification network (ECNet). This framework maximizes the benefits of convolutional operators for local feature representation and leverages Transformer architecture for learning long-distance dependencies. Inspired by the forgetting strategy of human brain neurons, we extend the concept of information loss to address feature preservation issues in neural networks. ECNet effectively preserves essential features while eliminating redundancy in high-dimensional manifold structures, thus mitigating overfitting caused by redundant multimodal features. Extensive experiments on four publicly available datasets demonstrate the competitiveness of ECNet in classification tasks. Notably, on the Houston2013 dataset, ECNet achieves an impressive overall accuracy (OA) of 96.63%, surpassing various state-of-the-art baseline approaches. The code is available athttps://github.com/lyb10087/ECNetfor the sake of reproducibility.
Weiying Xie, Yibing Lu, Daixun Li, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 RS-DGC: Exploring Neighborhood Statistics for Dynamic Gradient Compression on Remote Sensing Image Interpretation
abstract
Distributed deep learning has recently been attracting more attention in remote sensing (RS) applications due to the challenges posed by the increased amount of open data that are produced daily by Earth observation programs. However, the high communication costs of sending model updates among multiple nodes are a significant bottleneck for scalable distributed learning. Gradient sparsification has been validated as an effective gradient compression (GC) technique for reducing communication costs and thus accelerating the training speed. Existing state-of-the-art gradient sparsification methods are mostly based on the “larger-absolute-more-important” criterion, ignoring the importance of small gradients, which is generally observed to affect the performance. Inspired by informative representation of manifold structures from neighborhood information, we propose a simple yet effective dynamic gradient compression scheme leveraging neighborhood statistics indicator for RS image interpretation, termed RS-DGC. We first enhance the interdependence between gradients by introducing the gradient neighborhood to reduce the effect of random noise. The key component of RS-DGC is a Neighborhood Statistical Indicator (NSI), which can quantify the importance of gradients within a specified neighborhood on each node to sparsify the local gradients before gradient transmission in each iteration. Further, a layer-wise dynamic compression scheme is proposed to track the importance changes of each layer in real time. Extensive downstream tasks validate the superiority of our method in terms of intelligent interpretation of RS images. For example, we achieve an accuracy improvement of 0.51% with more than 50× communication compression on the NWPU-RESISC45 dataset using VGG-19 network. To the best of our knowledge, this is the first gradient compression method designed for RS images and downstream tasks, achieving a successful trade-off between high compression ratio and performance.
Weiying Xie, Jitao Ma, Daixun Li, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.4