Xueqiong Li

dblp:198/4468 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-2364-4947ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Leveraging VLMs for MUDA: Category-specific prompt with multi-modal interactive LoRA
Xihuai He, Xueqiong Li, Wanrong Huang, Hengzhu Liu, Huibin Tan
Neural Networks3
2025 HyperSDT: HyperNetwork Slide Decision Tree for Interpretable Tabular Learning
abstract
Recently, substantial progress has been achieved in leveraging deep learning models for tabular data learning. However, despite significant advancements, the predominant focus of these endeavors has been on augmenting the performance of contemporary deep learning models. Consequently, the interpretability of such models is frequently overlooked or rendered secondary, thereby posing a challenge in comprehending their underlying decision-making processes. In this work, we propose a novel HyperNetwork Slide Decision Tree (HyperSDT) approach to achieve interpretable deep learning for tabular data while maintaining a comparable accuracy to state-of-the-art methods. HyperSDT provides a comprehensive interpretable framework with interpretability by using Silde Decision Tree and Decision Transformer together. Our experimental results demonstrate that our framework is competitive with prior baselines under various tabular learning benchmarks while providing better interpretability. The code can be achieved via https://github.com/hunan-create/HyperSDT.
Xueqiong Li, Zhenhua Liang, Shaowu Yang, Ji Wang 0001
ICASSP2
2025 Multi-layer Network Disintegration via Deep Reinforcement Learning
abstract
Multi-layer networks (MLN) effectively model interactions across layers, and the network disintegration (ND) problem yields significant importance in the analysis of MLN. Unfortunately, previous advances in ND for single-layer networks exhibits inefficiency and lack of scalability when extended to MLN, as MLN involve complex inter-layer dependencies and interactions that are absent in single-layer networks. To bridge this gap, we propose a pioneer framework named Multi-layer Network Disintegration via Deep Reinforcement Learning (MNDRL), by re-formulating the disintegration process into a preprocessing-encoding-decoding pipeline. To be specific, MNDRL consists of two Graph Neural Network (GNN) models to effectively capture intra-layer and inter-layer representations separately. Empowered with DRL, our MNDRL achieves near-optimal results in the ND task with the NP-hard complexity by autonomously learning efficient disintegration strategies. Extensive experiments demonstrate that our MNDRL outperforms the results of baseline disintegration methods by over 30% on real-world datasets and by over 10% on synthetic datasets with varying node counts.
Zhenhua Liang, Xueqiong Li, Shaowu Yang, Hengzhu Liu
ICASSP2
2025 UniIVFT: Towards a Unified Framework for Infrared-Visible Fusion and Translation
abstract
Infrared-visible image fusion (IVF) and infrared-to-visible image translation (I2V) are two closely related tasks in multimodal image processing, both aimed at combining or transforming infrared and visible modalities to enhance image information content. Existing methods typically focus on either fusion or translation, often requiring redundant construction of similar components for each task, which limits the effective utilization of cross-modal interactions and feature encoding capabilities. Furthermore, these approaches are often hindered by their reliance on complex feature extract models, limiting their overall effectiveness and adaptability. In this paper, we introduce the Unified Multimodal Infrared-Visible Image Fusion and Translation (UniIVFT) framework, which integrates both fusion and translation tasks within a single architecture. We employ a vision transformer (ViT) encoder-decoder structure augmented with task-specific tokens and introduce a contrastive loss to effectively align infrared and visible image features before multimodal encoding. This alignment enhances the encoder’s ability to capture cross-modal interactions. In UniIVFT, both IVF and I2V tasks share a unified encoder architecture and use task-specific tokens to control model outputs, reducing redundant model construction and training. Extensive experiments demonstrate that UniIVFT achieves performance on par with that of SOTAs across multiple tasks while maintaining a lightweight architecture with fewer model parameters.
Xueqiong Li, Shaowu Yang, Huibin Tan, Yuhua Tang
ICASSP2
2025 AGFT-Tracker: Adaptive Game-Based PEFT for Object Tracking with PLMs
abstract
The rise of pre-trained large models (PLMs) has sparked interest in vision tasks like object tracking. However, as PLMs scale, fully fine-tuning all parameters becomes impractical, highlighting the need for parameter-efficient fine-tuning (PEFT). While adapter tuning, which adds tunable parameters to Multi-Head Attention (MHA) or Feed-Forward Networks (FFN), is common, critical parameters like Layer Normalization (LN), vital for stability and convergence, are often overlooked. Furthermore, traditional fine-tuning strategies fail to differentiate module importance, limiting performance improvements. To solve these issues, we propose a new PEFT method for unlocking large model potential in object tracking: Adaptive Game-Based Fine-tuning Tracker (AGFT-Tracker). AGFT-Tracker combines adapter tuning with direct LN fine-tuning and adaptively allocates parameter budgets based on tracking attention losses. Important sensitive modules use higher-rank LoRA and frozen LN, while stable modules undergo lower-rank LoRA and LN adjustments. This approach improves effectiveness and efficiency, achieving state-of-the-art results on challenging benchmarks.
Mingyu Cao, Xihuai He, Xueqiong Li, Kedi Zhang, Yuhua Tang, Wanrong Huang, Huibin Tan
ICME3
2025 Adaptive Distribution-Aware Modeling for Transformer Tracking
abstract
Adapting to changes in data distribution is a major challenge in visual object tracking. In Transformer-based tracking, Layer Normalization (LN) is often applied uniformly to both template and search features, limiting feature diversity. Additionally, models tend to converge to trivial solutions, and tracking samples are sensitive to distribution shifts, affecting robustness. To address these issues, we propose the Adaptive Distribution-Aware Transformer Tracker (ADAT), incorporating three key components: the Target-Aware Module (TAM), the Region-Aware Module (RAM), and the Self-Feedback-Aware Module (SFAM). TAM normalizes template and search features separately, preserving flexibility and enhancing target learning. RAM refines target perception by distinguishing between near and far target regions. SFAM filters out noisy samples and fine-tunes normalization parameters through self-feedback. While TAM and RAM regulate feature-level distribution, SFAM adjusts at the sample level. Extensive experiments show that ADAT outperforms existing methods, achieving superior performance on challenging benchmarks.
Mingyu Cao, Huibin Tan, Xueqiong Li, Wanrong Huang, Kedi Zhang, Yuhua Tang, Shaowu Yang
ICME3
2025 AIM-VR: All-in-One Video Restoration via Dual-Path Mamba with Frequency Adaptive Fusion
abstract
Real-world video-based vision systems frequently suffer concurrent degradations caused by unpredictable weather conditions such as rain, haze, and snow, severely affecting the visual quality of both human observers and the downstream computer vision tasks. In this paper, we propose All In Mamba Video Restoration (AIM-VR) model, a novel multi-degradation video restoration framework based on the Selective State Space Model. The proposed AIM-VR model effectively handles adverse weather conditions through three key innovations: a Dual-Path Mamba Modeling (DPMM) backbone with complementary Space-Time Sequence Mamba Block (SSMB) and Hilbert Sequence Mamba block (HSMB) for efficient temporal modeling, a Frequency Adaptive Fusion block (FAFB) for degradation-specific feature modulation, and a Universal Multi-degradation Contrastive Learning (UMCL) strategy for robust pattern discovery in multi-degradation scenarios. Experimental results demonstrate that the proposed AIM-VR achieves superior performance in terms of both restoration quality (0.82 dB PSNR over the state-of-the-art methods) and computational efficiency across multiple weather scenarios. Code is available at https://github.com/StephenLockhart/AIM-VR.
Zhizhou Lu, Junjie Huang 0001, Xueqiong Li, Baili Xiao
ICME5
2025 Multi-Resolution Infrared-Visible Image Fusion using Multi-Scale Residual Quantization
abstract
Infrared-visible image fusion (IVF) is an essential task in multimodal image processing that integrates infrared and visible modalities to enhance the overall image information content. However, existing methods often suffer from limited precision and efficiency. Furthermore, they fail to address practical requirements such as multi-resolution fusion and mutual translation. In this paper, we propose the Multi-Scale Residual Quantized Infrared-Visible Image Fusion (M-RQIVF) framework to efficiently generate high-quality fusion images. M-RQIVF trains multi-scale residual quantized infrared and visible autoencoders that convert images into multi-scale discrete token maps. This approach approximates the residuals from the features on a scale-by-scale basis, allowing for coarse-to-fine fused image generation that aligns well with human visual perception. Furthermore, by leveraging these discrete token maps, we train Visual Auto-Regressive (VAR) transformers using next-scale prediction. The VAR transformer ensures that features of corresponding sizes can be generated, even when the input infrared and visible images have different resolutions, facilitating fine-grained fusion. Additionally, the autoregressive structure enables image translation to be treated as a conditional generation task, thereby enabling mutual translation between infrared and visible images. Extensive experiments demonstrate that M-RQIVF outperforms the SOTAs while maintaining a much faster inference speed.
Huibin Tan, Wanrong Huang, Yuhua Tang, Xueqiong Li
ICME6
2025 DFDUN: Deep Infrared and Visible Image Fusion with Diffusion Prior Unfolding Network
abstract
Infrared and Visible Image Fusion (IVF) intends to aggregate information from infrared and visible modalities, generating comprehensive images. While deep-unfolding-based methods and diffusion-based methods show satisfactory performances, the former suffers from insufficiently deterministic priors, and latter is limited by the inaccurate prior generation or opaque working mechanisms. In this paper, we propose Deep infrared and visible image Fusion with Diffusion prior Unfolding Network (DFDUN), aiming for effective fusion with the generative diffusion prior in a transparent mechanism. DFDUN starts with a model-based fusion optimization formulation, which is unfolded into a Denoising Diffusion Module (DDM) for generating informative diffusion prior, and a Data Consistent Module (DCM) that transparently and effectively aggregates complementary information from diffusion prior and source modalities. Moreover, DFDUN employs a hypernetwork for adaptive dictionary parameter generation in DCM, enhancing fusion flexibility. Experimental results indicate that DFDUN outperforms existing methods, providing superior IVF performance with efficient and transparent fusion. Code is available at https://github.com/XiongMaoyi2001/DFDUN.
Maoyi Xiong, Tianrui Liu 0001, Xueqiong Li, Yuhua Tang
ICME5
2024 Diversifying Cross-Domain Few-Shot Learning via Multimodal Image Editing
abstract
Standing out as one of the most widely used tools in Cross-Domain Few-Shot Learning (CDFSL), data augmentation forms the bedrock of numerous recent advancements. However, the current augmentations in CDFSL are limited in their ability to modify high-level semantic attributes, resulting in a lack of diversity along key semantic dimensions. One of the most promising tools to edit images with key semantic attributes, e.g. backgrounds, is image-to-image generation via large multimodal models (LMMs). Given the promising image editing results of recent LMMs, we delve into leveraging LMMs to augment data diversity for CDFSL. We propose a novel method named, Multimodal Few-shot Image Editing (MFIE), which uses LMMs to automatically translate class-specific images into class-agnostic natural language descriptions for various key semantic attributes in target domains and editing origin images based on class-agnostic natural language descriptions. To filter out corrupted data that disturbs the class-specific information, we apply semantic filtering using image-language similarity. Experiments on Meta-Datset show that MFIE surpasses SOTA CDFSL algorithms.
Wenjing Yang 0002, Long Lan, Mingyang Geng, Haotian Wang 0001, Haoang Chi, Xueqiong Li, Ji Wang 0001
ICASSP7
2024 Radar Recognition in the Wild: Enhancing Radar Emitter Recognition through Auto-Correlation Model-Agnostic Meta Learning
abstract
In Electronic Support Measure (ESM) systems, the recognition of radar emitters stands as a pivotal yet intricate task. The complex electromagnetic environments, however, often hinders the collection of clean radar signal data, and results in data with different noise levels. Consequently, formulating a robust recognition model with limited data becomes a big challenge, further compounded by the demand for generalizability across scenarios with different noise levels. While Model-Agnostic Meta Learning (MAML) has proven its effectiveness in solving few-shot learning problems in computer vision, its application in radar signal processing has remained unexplored deeply. This paper pioneers the incorporation of MAML and autocorrelation into radar signal processing. To fit MAML to radar signals, we introduce a novel loss function, termed AC-Loss, designed to facilitate learning effective signal representation by retaining the periodicity of the radar pulses which is the key feature for recognizing different Pulse Repetition Intervals (PRIs). This proposed Autocorrelation Model-Agnostic Meta Learning (AC-MAML) enhances its recognition capabilities while using only a sparse number of signal samples in both source and target domains. Empirical results show the superiority of AC-MAML, achieving an impressive average recognition accuracy of 90.4% across seven diverse target domain scenarios.
Yixian Luo, Shaowu Yang, Huibin Tan, Ruochun Jin, Hengzhu Liu, Xueqiong Li
ICASSP7
2024 Modality Re-Balance for Visual Question Answering: A Causal Framework
abstract
Visual Question Answering (VQA) models often prioritize language cues over visual knowledge, leading to the "language prior" phenomenon. To address this, researchers have proposed methods to balance language and image information during training and inference. However, these approaches often struggle to capture important linguistic components due to the excessive exclusion of language information. Inspired by causal inference, we introduce a novel approach called the SyMmetrically Balanced Causal framework (SMBC) that rebalances visual and textual information in VQA tasks. This framework allows for an equal contribution of knowledge from both modalities to inference results. Experimental evaluation shows that SMBC: 1) applies to prevalent VQA models, including those with data augmentation, and 2) consistently improves performance on established benchmarks.
Xinpeng Lv, Wanrong Huang, Haotian Wang 0001, Ruochun Jin, Xueqiong Li, Shuman Li, Yongquan Feng, Yuhua Tang
ICASSP5
2024 CRNet: Cross-Reconstruction Network for Inconsistent Point Cloud Registration
abstract
Deep learning methods have made significant advancements in point cloud registration, achieving excellent performance on consistent point clouds. However, these methods face challenges when dealing with point clouds exhibiting inconsistent spatial distributions. To address this issue, we propose the Cross-Reconstruction Network (CRNet), a novel approach designed to register two inconsistent point clouds by reconstructing them into a consistent shape. CRNet utilizes a cross-learning framework that facilitates feature interaction between input point clouds at both global and point-wise levels. This interaction network enables the bidirectional generation of corresponding points to reconstruct consistent point clouds for transformation estimation. Furthermore, the transformation parameters can be refined by a regression network to achieve more accurate registration. The experimental results valuated on benchmark datasets demonstrate that CRNet outperforms state-of-the-art methods in inconsistent scenarios.
Yunzhe Xiao, Xueqiong Li, Shaowu Yang, Wenjing Yang 0002, Yong Dou
ICME2
2024 TIG-CL: Teacher-Guided Individual- and Group-Aware Contrastive Learning for Unsupervised Person Reidentification in Internet of Things
abstract
Unsupervised person reidentification (Re-ID) has attracted widespread due to its potential in Internet of Things applications, such as intelligent visual surveillance, it refers to retrieving the same individual across different camera views without using labeled data. To tackle the problem, a prevalent technique adopted by existing methods involves generating pseudo labels through clustering algorithms. However, this approach can result in merging individuals with different identities into the same group (i.e., cluster) during the training process. As a result, the resulting group centers may obscure the inherent characteristics of individual identities, thereby hindering the model from learning discriminative representations. To address the issue, we present a teacher-guided individual- and group-aware contrastive learning framework. Specifically, we propose a departure from the traditional approach of relying solely on contrastive learning between individual features and their corresponding group centers. Instead, we also exploit the relationship among individuals to construct contrast pairs and facilitate the learning of more discriminative features. This strategy enables the model to learn more about the individual characteristics that distinguish different persons, thus enhancing its ability to reidentify individuals accurately. Moreover, our method introduces a novel hybrid distillation module that enables simultaneous probability distillation at the group level and relationship distillation at the individual level. Guided by the teacher model, this module leads to improved feature representations of the student model. Extensive experimental results verify the effectiveness of our approach on four popular Re-ID data sets. The code will be made publicly available.
Xiao Teng, Xueqiong Li, Xinwang Liu 0002, Long Lan
IEEE Internet Things J.3
2024 Joint Spatial-Spectral Optimization for the High-Magnification Fusion of Hyperspectral and Multispectral Images
abstract
The fusion of hyperspectral and multispectral images is an important strategy for enhancing the spatial resolution of hyperspectral images. With the rapid advancement of multispectral imaging technology, the disparity in spatial resolution between multispectral and hyperspectral images is increasing. In certain scenarios, termed high-magnification, this difference can exceed$32\times $. Previous methods do not perform well under high-magnification fusion, and naturally, a challenge arises in achieving effective high-magnification super-resolution fusion. In light of the above analysis, this article introduces a novel algorithm for high-magnification super-resolution fusion of hyperspectral and multispectral images based on the joint optimization of spatial and spectral information. Specifically, our algorithm consists of three stages: 1) a fast preliminary fusion stage based on the Moore-Penrose inverse and singular value correlation priors for the rapid acquisition of preliminary solutions; 2) a joint spatial-spectral optimization stage where a coupled optimization framework is constructed to achieve integrated optimization of spatial and spectral information; and 3) an error backpropagation optimization stage where an effective error optimization term is introduced to further refine the fusion performance. We conducted extensive experiments on widely employed publicly available simulated datasets and real datasets. The experimental results unequivocally indicate that our proposed methodology consistently exhibits superior fusion performance compared with state-of-the-art methods, even under the condition of${\geq }60\times $super-resolution.
Yibing Zhan, Zhengbin Pang, Tong Zhou 0008, Xueqiong Li, Long Lan, Yuanxi Peng
IEEE Trans. Geosci. Remote. Sens.5
2024 Highly Efficient Active Learning With Tracklet-Aware Co-Cooperative Annotators for Person Re-Identification
abstract
Supervised person re-identification (ReID) has attracted widespread attentions in the computer vision community due to its great potential in real-world applications. However, the demand of human annotation heavily limits the application as it is costly to annotate identical pedestrians appearing from different cameras. Thus, how to reduce the annotation cost while preserving the performance remains challenging and has been studied extensively. In this article, we propose a tracklet-aware co-cooperative annotators' framework to reduce the demand of human annotation. Specifically, we partition the training samples into different clusters and associate adjacent images in each cluster to produce the robust tracklet which decreases the annotation requirements significantly. Besides, to further reduce the cost, we introduce a powerful teacher model in our framework to implement the active learning strategy and select the most informative tracklets for human annotator, the teacher model itself, in our setting, also acts as an annotator to label the relatively certain tracklets. Thus, our final model could be well-trained with both confident pseudo-labels and human-given annotations. Extensive experiments on three popular person ReID datasets demonstrate that our approach could achieve competitive performance compared with state-of-the-art methods in both active learning and unsupervised learning (USL) settings.
Xiao Teng, Long Lan, Xueqiong Li, Yuhua Tang
IEEE Trans. Neural Networks Learn. Syst.4
2017 pHMM-tree: phylogeny of profile hidden Markov models
abstract
Protein families are often represented by profile hidden Markov models (pHMMs). Homology between two distant protein families can be determined by comparing the pHMMs. Here we explored the idea of building a phylogeny of protein families using the distance matrix of their pHMMs. We developed a new software and web server (pHMM-tree) to allow four major types of inputs: (i) multiple pHMM files, (ii) multiple aligned protein sequence files, (iii) mixture of pHMM and aligned sequence files and (iv) unaligned protein sequences in a single file. The output will be a pHMM phylogeny of different protein families delineating their relationships. We have applied pHMM-tree to build phylogenies for CAZyme (carbohydrate active enzyme) classes and Pfam clans, which attested its usefulness in the phylogenetic representation of the evolutionary relationship among distant protein families. Availability and Implementation: This software is implemented in C/C ++ and is available at http://cys.bios.niu.edu/pHMM-Tree/source/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Luyang Huo, Han Zhang 0017, Xueting Huo, Yasong Yang, Xueqiong Li
Bioinform.5