Jitao Ma

dblp:334/0636 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MMFormer: Multi-Modality semi-Supervised vision transformer in remote sensing imagery classification
Daixun Li, Weiying Xie, Leyuan Fang, Yunke Wang, Mingxiang Cao, Jitao Ma, Yunsong Li 0001, Chang Xu 0002
Neural Networks7
2026 FA-Mamba: frequency attention driven Mamba for multimodal remote sensing classification
Danian Yang, Daixun Li, Jitao Ma, Yibing Lu, Yunsong Li 0001, Leyuan Fang, Weiying Xie
Neural Networks3
2026 Cross-Modal Visual Perception Consistency: A Language-Enhanced Approach for Heterogeneous Change Detection
abstract
Heterogeneous remote sensing image change detection (HRSICD) seeks to identify surface changes by comparing images captured at different times. However, CD faces significant challenges due to heterogeneity arising from varying sensor types and imaging conditions. Recently, powerful vision-language models like CLIP have emerged, with strong semantic decoding abilities. Opening new possibilities for using linguistic information as an auxiliary in visual tasks, potentially driving breakthroughs in HCD. Capitalizing on this prospect, we investigate graph learning with vision-language features and introduce LEVPC, the first language-enhanced visual perception consistency framework for HCD. First, we create a mutual information-guided graph aggregation module. Specifically, it builds modality-invariant structured relationships among visual nodes by using language features as connecting bridges, providing a consistent foundation for comparing changes. To reduce modeling bias from heterogeneity, language is used as an anchor to aggregate features, ensuring a unified expression of visual representations. In summary, language guides the generation and aggregation of multiple subgraphs from visual inputs, ultimately building robust representations of structural relationships within a shared semantic space. Moreover, a change semantic compensation module is introduced, which analyses the change intensity between bi-temporal data from a vision-language perspective. And then adds change-related semantic descriptions for salient change regions, enhancing the expressiveness of visual change features. Experiments on multiple datasets validate the superior performance of LEVPC in HCD, achieving an average increase of 2.6% in Kappa. The code will be publicly available at https://github.com/sylXIDIAN/LEVPC.
Siyao Li, Weiying Xie, Jitao Ma, Leyuan Fang, Yunsong Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 MultiGS: Multi-Dimensional Information-Aware Gradient Sparsification
abstract
Gradient sparsification (GS) is an effective method for reducing communication overhead in distributed training. For the first time, we introduce the concept of Multi-dimensional information into GS and propose a new gradient sparsification method named Multi-dimensional information-aware Gradient sparsification (MultiGS), which achieves high compression ratio with negligible accuracy loss and is applicable to mainstream network architectures. MultiGS reconstructs the layer-wise gradient by combining the high-frequency components of the local gradient and the low-frequency components of the sparsified global gradient that effectively addresses the issue of stale gradients and alleviates model bifurcation. Through the convergence proof of MultiGS for smooth non-convex problems and comparison with momentum SGD in convergence speed, we show that such new perspective approach is theoretically reasonable and practically effective. As validated with several mainstream model families (i.e., ResNets, VGGNet, LSTM, Vision Transformer, and Large Language Models), our MultiGS shows better accuracy over previous GS methods. Moreover, empirical results show that when a sufficient number of training nodes are available, MultiGS accelerates the distributed training by more than 3×, which is better than existing sparsification method.
Jitao Ma, Donglai Liu, Weiying Xie, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.1
2026 BSDM: Background Suppression Diffusion Model for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) is widely used in Earth observation and deep space exploration. A major challenge for HAD is the complex background of the input hyperspectral images (HSIs), resulting in anomalies confused in the background. On the other hand, most existing HAD methods require training a separate model for each HSI, resulting in poor generalization in practical applications. This paper starts the first attempt to study a new and generalizable background learning problem without labeled samples. We present a novel solution BSDM (background suppression diffusion model) for HAD, which can simultaneously learn latent background distributions and generalize to different datasets for suppressing complex background. It is featured in three aspects: (1) For the complex background of HSIs, we design pseudo-background noise and learn the potential background distribution in it with a diffusion model (DM). (2) For the generalizability problem, we apply a statistical offset module so that the BSDM adapts to datasets of different domains without labeling samples. (3) For achieving background suppression, we innovatively improve the inference process of DM by feeding the original HSIs into the denoising network, which removes the background as noise. Our work paves a new background suppression way for HAD that can improve HAD performance without the prerequisite of manually labeled data. Assessments and generalization experiments of four HAD methods on several real HSI datasets demonstrate the above three unique properties of the proposed method. Our project is available at https://github.com/majitao-xd/BSDM-HAD.
Jitao Ma, Weiying Xie, Xueshuang Xiang, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Aligning and Prompting Anything for Zero-Shot Generalized Anomaly Detection
abstract
Zero-shot generalized anomaly detection (ZGAD) plays a critical role in industrial automation and health screening. Recent studies have shown that ZGAD methods built on visual-language models (VLMs) like CLIP have excellent cross-domain detection performance. Different from other computer vision tasks, ZGAD needs to jointly optimize both image-level anomaly classification and pixel-level anomaly segmentation tasks for determining whether an image contains anomalies and detecting anomalous parts of an image, respectively, this leads to different granularity of the tasks. However, existing methods ignore this problem, processing these two tasks with one set of broad text prompts used to describe the whole image. This limits CLIP to align textual features with pixel-level visual features and impairs anomaly segmentation performance. Therefore, for precise visual-text alignment, in this paper we propose a novel fine-grained text prompts generation strategy. We then apply the broad text prompts and the generated fine-grained text prompts for visual-textual alignment in classification and segmentation tasks, respectively, accurately capturing normal and anomalous instances in images. We also introduce the Text Prompt Shunt (TPS) model, which performs joint learning by reconstruction the complementary and dependency relationships between the two tasks to enhance anomaly detection performance. This enables our method to focus on fine-grained segmentation of anomalous targets while ensuring accurate anomaly classification, and achieve pixel-level comprehensible CLIP for the first time in the ZGAD task. Extensive experiments on 13 real-world anomaly detection datasets demonstrate that TPS achieves superior ZGAD performance across highly diverse datasets from industrial and medical domains.
Jitao Ma, Weiying Xie, Hangyu Ye, Daixun Li, Leyuan Fang
AAAI1
2025 Allowing Oscillation Quantization: Overcoming Solution Space Limitation in Low Bit-Width Quantization
Weiying Xie, Zihan Meng, Jitao Ma, Wenjin Guo, Leyuan Fang, Yunsong Li 0001
ICCV3
2025 TF-ATM: Training-Free Adaptive Token Merging
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Tianlin Hui, Jitao Ma, Leyuan Fang
ACM Multimedia6
2025 Simultaneous Suppression of Random Noise and Ground Roll With a Fuzzy Logic-Guided Deep Network
abstract
Seismic data are frequently contaminated by incoherent random noise and coherent surface waves; both severely degrade subsurface data quality. While random noise is statistically irregular and uncorrelated, surface waves exhibit structured regularity in both temporal and spatial domains. Their coherent low-frequency propagation makes them fundamentally distinct from random noise, posing significant challenges to simultaneous suppression. This letter presents a Deep Adaptive Signal Denoising network (DASDNet) designed for random and ground roll simultaneous suppression using weakly supervised training on synthetic clean-noisy pairs. The architecture incorporates short-time Fourier transform (STFT) and discrete wavelet transform (DWT) for multiscale time-frequency decomposition. A fuzzy logic-based encoder is employed to mitigate feature uncertainty, complemented by a time-frequency regularization module that preserves structural coherence. A unified loss function across time, frequency, and spatial domains guides robust training. DASDNet demonstrates remarkable generalization capability to both synthetic and field-recorded datasets, achieving robust performance without reliance on labeled field data or manual post-hoc parameter tuning. Experimental results show that DASDNet improves signal-to-noise ratio (SNR) by 6.86 dB, reduces mean squared error (MSE) by 4433, and increases structural similarity index (SSIM) by 0.22, outperforming other methods. Qualitative analysis further demonstrates its advantage in preserving primary fidelity while suppressing noises.
Xuebin Zuo, Jitao Ma, Zhen Liao, Xiaohong Chen 0003
IEEE Geosci. Remote. Sens. Lett.2
2025 Exploring hyperspectral anomaly detection with human vision: A small target aware detector
Jitao Ma, Weiying Xie, Yunsong Li 0001
Neural Networks1
2025 Semi-Mamba: Mamba-Driven Semi-Supervised Multimodal Remote Sensing Feature Classification
abstract
Mamba architecture achieves the same performance as attention mechanisms with linear complexity, leading to significant progress in remote sensing land cover classification. However, existing Mamba methods rarely leverage the representational complementarity and consistency between different modalities, resulting in challenges such as incomplete fusion. To address these issues, we propose Semi-Mamba, a novel semi-supervised framework specifically designed for high-dimensional multi-modal data fusion. We introduce the Mamba Cross-Modality Fusion Module, which enables cross-modal learning of temporal features through state-space model interactions and smooth integration of input matrices, enhancing the fusion of richer feature representations. Additionally, to tackle the inherent difficulty of acquiring pixel-level annotations in remote sensing datasets, we introduce a multi-modal semi-supervised mechanism. This mechanism utilizes cross-modal supervision between different modalities to maximize data utilization and improve learning efficiency. It effectively enables joint training on both labeled and unlabeled data without relying on pseudo-labels. We integrate these innovations into a unified end-to-end framework. Compared to state-of-the-art CNN and Transformerbased architectures, our framework shows a significant improvement of over 3.12%, setting a new benchmark for semi-supervised multi-modal data fusion. The code has open sourced at https://github.com/LDXDU/Semi_Mamba_RS.
Yunsong Li 0001, Daixun Li, Weiying Xie, Jitao Ma, Sibo He, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Distributed Deep Learning With Gradient Compression for Big Remote Sensing Image Interpretation
abstract
Fast and reliable interpretation of high-dimensional hyperspectral images (HSIs) can provide great support to remote sensing-based Earth observations. Targets of interest in HSI can be detected using deep neural networks (DNNs) for background learning on an acquired image where the occurrence probability of background samples is much greater than that of targets, accounting for more than 95% of the whole scene. However, there is an increasing gap between theory and feasible application, because of the contradiction between massive hyperspectral data and resource-limited Internet of Things (IoT)/edge device hardware like satellite. To facilitate the deployment of hyperspectral target detection (HTD) in an edge computing environment, we introduce distributed background learning-a decentralized deep learning approach to meet the computing requirements of exploding high-dimensional data and larger DNNs. To address the communication bottleneck caused by gradient exchange during distributed learning, the proposed gradient compression solution, named gradient compression via centroid (GCC), uniquely compresses the most replaceable gradients with redundant information, thereby reducing communication overhead while maintaining accuracy. To illustrate the feasibility of the proposed method, we test it over two very large hyperspectral datasets with a total size of about 3.2 gigabytes (GBs) on a distributed system based on Ring All-reduce. We show that HTD based on distributed background learning outperforms those developed on a single node in terms of speed. Besides, the GCC compresses 50% gradients with only 0.01% loss of target detection accuracy to greatly reduce the communication overhead, surpassing existing gradient compression methods. It is expected that this framework will accelerate the introduction of distributed training on IoT/edge devices.
Weiying Xie, Jitao Ma, Tianen Lu, Yunsong Li 0001, Jie Lei 0001, Leyuan Fang, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 JointSQ: Joint Sparsification-Quantization for Distributed Learning
abstract
Gradient sparsification and quantization offer a promising prospect to alleviate the communication overhead problem in distributed learning. However, direct combination of the two results in suboptimal solutions, due to the fact that sparsification and quantization haven't been learned together. In this paper, we propose Joint Sparsification-Quantization (JointSQ) inspired by the discovery that sparsification can be treated as 0-bit quantization, regardless of architectures. Specifically, we mathematically formu-late JointSQ as a mixed-precision quantization problem, expanding the solution space. It can be solved by the designed MCKP-Greedy algorithm. Theoretical analysis demon-strates the minimal compression noise of JointSQ, and ex-tensive experiments on various network architectures, including CNN, RNN, and Transformer, also validate this point. Under the introduction of computation overhead consistent with or even lower than previous methods, JointSQ achieves a compression ratio of 1000× on different models while maintaining near-lossless accuracy and brings 1.4× to 2.9× speedup over existing methods.
Weiying Xie, Jitao Ma, Yunsong Li 0001, Jie Lei 0001, Donglai Liu, Leyuan Fang
CVPR3
2024 FedSLS: Exploring Federated Aggregation in Saliency Latent Space
abstract
Federated Learning (FL) is an emerging direction in distributed machine learning that enables jointly training a global model without sharing data with server. However, data heterogeneity biases the parameter aggregation at the server, leading to slower convergence and poorer accuracy of the global model. To cope with this, most of the existing works involve enforcing regularization in local optimization or improving the model aggregation scheme at the server. Though effective, they lack a deep understanding of cross-client features. In this paper, we propose a saliency latent space feature aggregation method (FedSLS) across federated clients. By Guided BackPropagation (GBP), we transform deep models into powerful and flexible visual fidelity encoders, applicable to general state inputs across different image domains, and achieve powerful aggregation in the form of saliency latent features. Notably, since GBP is label-insensitive, it is sufficient to capture saliency features only once on each client. Experimental results demonstrate that FedSLS leads to significant improvements over the state-of-the-arts in terms of accuracies, especially in highly heterogeneous settings. For example, on CIFAR-10 dataset, FedSLS achieves 63.43% accuracy within the strongly heterogeneous environment α=0.05, which is 6% to 23% higher than other baselines.
Hengyi Wang, Weiying Xie, Jitao Ma, Daixun Li, Yunsong Li 0001
ACM Multimedia3
2024 Adaptive Pruning of Channel Spatial Dependability in Convolutional Neural Networks
abstract
Deep Convolutional Neural Networks (CNNs) have demonstrated excellent performance in various multimedia application scenarios. However, complex models often require significant computational resources and energy costs. Therefore, CNN compression is crucial for addressing deployment challenges of multimedia application on resource constrained edge devices. However, existing CNN channel pruning strategies primarily focus on the "weights" or "activations" of the model, overlooking its "interpretability" information. In this paper, we explore CNN pruning strategies from the perspective of model interpretability. We model the correspondence between channel feature maps and interpretable visual perception based on class saliency maps, aiming to assess the contribution of each channel to the desired output. Additionally, we utilize Discrete Wavelet Transform (DWT) to capture the global features and structure of class saliency maps. Based on this, we propose a Channel Spatial Dependability (CSD) metric, evaluating the importance and contribution of channels in a bidirectional manner to guide model pruning. And we dynamically adjust the pruning rate of each layer based on performance changes, in order to achieve more accurate and efficient adaptive pruning. Our method achieves significant results across a range of different networks and datasets. For instance, we achieved a 51.3% pruning on the ResNet-56 model while maintaining an accuracy of 94.16%, outperforming feature-map or other State-of-the-Art (SOTA).
Weiying Xie, Mei Yuan, Jitao Ma, Yunsong Li 0001
ACM Multimedia3
2024 A graphic structure based branch-and-bound algorithm for complex quadratic optimization and applications to magnitude least-square problem
Cheng Lu 0007, Jitao Ma, Zhibin Deng, Wenxun Xing
J. Glob. Optim.2
2024 RS-DGC: Exploring Neighborhood Statistics for Dynamic Gradient Compression on Remote Sensing Image Interpretation
abstract
Distributed deep learning has recently been attracting more attention in remote sensing (RS) applications due to the challenges posed by the increased amount of open data that are produced daily by Earth observation programs. However, the high communication costs of sending model updates among multiple nodes are a significant bottleneck for scalable distributed learning. Gradient sparsification has been validated as an effective gradient compression (GC) technique for reducing communication costs and thus accelerating the training speed. Existing state-of-the-art gradient sparsification methods are mostly based on the “larger-absolute-more-important” criterion, ignoring the importance of small gradients, which is generally observed to affect the performance. Inspired by informative representation of manifold structures from neighborhood information, we propose a simple yet effective dynamic gradient compression scheme leveraging neighborhood statistics indicator for RS image interpretation, termed RS-DGC. We first enhance the interdependence between gradients by introducing the gradient neighborhood to reduce the effect of random noise. The key component of RS-DGC is a Neighborhood Statistical Indicator (NSI), which can quantify the importance of gradients within a specified neighborhood on each node to sparsify the local gradients before gradient transmission in each iteration. Further, a layer-wise dynamic compression scheme is proposed to track the importance changes of each layer in real time. Extensive downstream tasks validate the superiority of our method in terms of intelligent interpretation of RS images. For example, we achieve an accuracy improvement of 0.51% with more than 50× communication compression on the NWPU-RESISC45 dataset using VGG-19 network. To the best of our knowledge, this is the first gradient compression method designed for RS images and downstream tasks, achieving a successful trade-off between high compression ratio and performance.
Weiying Xie, Jitao Ma, Daixun Li, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.3