Shuxiang Song 0001

dblp:162/4519-1 · also Shu-Xiang Song 0001 · DBLP profile ↗
← Back
53ranked-venue papers
0as first author
47since 2021 · last 2026
0000-0003-0280-2640ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 27 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
abstract
Music emotion recognition is a key task in symbolic music understanding (SMER). Recent approaches have shown promising results by fine-tuning large-scale pre-trained models (e.g., MIDIBERT, a benchmark in symbolic music understanding) to map musical semantics to emotional labels. While these models effectively capture distributional musical semantics, they often overlook tonal structures, particularly musical modes, which play a critical role in emotional perception according to music psychology. In this paper, we investigate the representational capacity of MIDIBERT and identify its limitations in capturing mode-emotion associations. To address this issue, we propose a Mode-Guided Enhancement (MoGE) strategy that incorporates psychological insights on mode into the model. Specifically, we first conduct a mode augmentation analysis, which reveals that MIDIBERT fails to effectively encode emotion-mode correlations. Motivated by this observation, we further identify the MIDIBERT layer that shows the weakest emotion relevance and introduce a Mode-guided Feature-wise linear modulation injection (MoFi) framework to inject explicit mode features, thereby enhancing the model's capability in emotional representation and inference. Extensive experiments on the EMOPIA and VGMIDI datasets demonstrate that our mode injection strategy significantly improves SMER performance, achieving accuracies of 75.2% and 59.1%, respectively. These results validate the effectiveness of mode-guided modeling in symbolic music emotion recognition.
Haiying Xia, Yumei Tan, Shuxiang Song 0001
AAAI4
2026 Motion-Aware Object Tracking via Motion and Geometry-Aware Cues
abstract
Understanding motion is essential for visual object tracking, especially in complex and dynamic scenarios. Yet, many existing methods rely on simplistic strategies such as template updates or temporal feature propagation, often overlooking the deeper modeling of motion information. To mitigate this limitation, we introduce a motion-aware spatio-temporal framework that enhances motion perception by explicitly matching motion patterns and modeling inter-frame motion relationships. Central to our design is a motion pattern dictionary, which encodes a diverse set of representative motion cues as learnable features. During tracking, features from the search region interact with the dictionary to retrieve the most relevant motion patterns, allowing the model to adapt to the current motion state. A dedicated decoder further incorporates temporal correlations to refine motion awareness. To complement motion modeling, we embed geometric cues into the search region features, which strengthens spatial perception, reduces ambiguity under occlusion, and improves foreground-background separation. Extensive evaluations on seven challenging benchmarks demonstrate the effectiveness of our design. In particular, MoDTrack_384 surpasses recent SOTA trackers on LaSOT by 1.2% in AUC, highlighting the benefits of motion pattern modeling and geometry-guided enhancement in mitigating tracking drift.
Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Yufei Tan, Haiying Xia, Shuxiang Song 0001
AAAI7
2026 AVSCNet: A dual-branch network for synchronization detection and content consistency learning in audio-video forgery detection
Guangwei Zhu, Haiying Xia, Shuxiang Song 0001
Neurocomputing4
2026 Dynamic cross-instance context mining for multimodal sentiment analysis
Haiying Xia, Youyong Cheng, Yumei Tan, Shuxiang Song 0001
Inf. Process. Manag.4
2026 HEL-Net: Heterogeneous Ensemble Learning for comprehensive diabetic retinopathy multi-lesion segmentation via Mamba-UNet
Lingyu Wu, Haiying Xia, Shuxiang Song 0001
Image Vis. Comput.3
2026 Global-local co-regularization network for facial action unit detection
Yumei Tan, Haiying Xia, Shuxiang Song 0001
J. Vis. Commun. Image Represent.3
2026 GLA: Globally-aware graph and dilated local attention for enhanced 3D object detection
Hai-Sheng Li 0001, Haizhen Liu, Shuxiang Song 0001
Pattern Recognit. Lett.3
2026 Unified Multi-Modal Tracking via Proxy Visual Prompts
abstract
Recent unified multi-modal tracking frameworks often encounter high computational overhead due to complex fusion operations. In this paper, we propose PoATrack, a proxy-based bidirectional fusion framework designed to facilitate efficient multi-modal tracking via interactive proxy prompts. Specifically, the framework treats features from all modalities equally and introduces an adaptive feature enhancement module to improve spatial representations within the search region. To further reduce the fusion cost, a lightweight proxy prompt fusion module is developed following a two-stage strategy: (1) representing modality-specific features through proxy visual prompts, and (2) dynamically learning hierarchical cross-modal relationships via proxy-guided fusion for enriched contextual modeling. Extensive experiments on six public benchmarks (RGB-T, RGB-E, and RGB-D) demonstrate that PoATrack achieves competitive accuracy while operating 1.8× faster than SDSTrack [17] under identical hardware settings.
Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiruo Zhu, Shuxiang Song 0001
IEEE Trans Autom. Sci. Eng.6
2026 Mamba-Driven Diffusion Model for Salient Object Detection in Optical Remote Sensing Images
abstract
Existing Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) methods mainly rely on a semantic segmentation paradigm, which relies on pixel-wise probabilities, leading to overconfident mispredictions. In contrast, the random sampling process of the diffusion model allows multiple possible predictions to be drawn from the mask distribution, effectively alleviating this problem. However, existing diffusion models mainly use Transformers as conditional feature extraction networks. Although they are good at global modeling, they have limited ability to handle long-range dependencies due to computational complexity. To overcome these challenges, we introduce MambaDif, an innovative diffusion model architecture based on Mamba. Specifically, we regard ORSI-SOD as a conditional mask generation task leveraging the diffusion model and achieving target distribution matching by adding noise to the mask and iteratively denoising it to match the target distribution. Then, we adopt Mamba to extract global features, efficiently process long sequences, and capture global contextual information with linear complexity. In addition, we introduce the global-local feature collaborative completion module (GLM), which combines the ability of convolutional layers to extract local features with the advantage of Mamba in capturing long-range dependencies, thereby achieving excellent denoising performance. Extensive experiments show that MambaDif outperforms SOTA methods in eight evaluation metrics on two standard datasets (EORSSD and ORSSD). We also report the generalization performance of the model on the challenging ORSI-4199 to evaluate its robustness.
Bineng Zhong 0001, Qihua Liang, Yufei Tan, Haiying Xia, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 Efficient Shortest Path-Driven Tabu Search Reconfiguration Algorithm for Multiprocessor Array
abstract
Fault-tolerant reconfiguration is essential for improving the reliability and efficiency of multiprocessor arrays. However, existing reconfiguration approaches primarily focus on algorithmic optimization, often neglecting architectural imbalance and communication bottlenecks arising from asymmetric redundancy placement. To overcome these limitations, this paper proposes an efficient framework that combines a novel double-sided redundant architecture with two-stage optimization algorithms to enhance interconnect efficiency under permanent processing element (PE) faults. The proposed$R_{l}+M+R_{r}$architecture symmetrically distributes redundant columns on both sides of the array while maintaining the same total number of spare PEs as single-sided redundancy, thereby mitigating unbalanced compensation paths and long communication distances without increasing hardware cost. Based on this structure, a shortest-path fault compensation algorithm is developed to minimize inter-PE communication distances, and a wide Tabu-based optimization algorithm is introduced as a second-stage refinement for global exploration and multi-directional reconfiguration. Experimental evaluations demonstrate that the proposed framework achieves superior communication performance and scalability compared with state-of-the-art methods under both single-sided and double-sided redundant architectures. It achieves lower latency, higher throughput, and reduced energy consumption while maintaining identical redundant resources and a smaller interconnect area footprint, confirming the efficiency and practicality of the proposed fault-tolerant topology reconfiguration strategy.
Hao Ding 0007, Yupeng Chi, Junyan Qian, Shuxiang Song 0001
IEEE Trans. Dependable Secur. Comput.5
2025 MambaLCT: Boosting Tracking via Long-term Context State Space Model
abstract
Effectively constructing context information with long-term dependencies from video sequences is crucial for object tracking. However, the context length constructed by existing work is limited, only considering object information from adjacent frames or video clips, leading to insufficient utilization of contextual information. To address this issue, we propose MambaLCT, which constructs and utilizes target variation cues from the first frame to the current frame for robust tracking. First, a novel unidirectional Context Mamba module is designed to scan frame features along the temporal dimension, gathering target change cues throughout the entire sequence. Specifically, target-related information in frame features is compressed into a hidden state space through a selective scanning mechanism. The target information across the entire video is continuously aggregated into target variation cues. Next, we inject the target change cues into the attention mechanism, providing temporal information for modeling the relationship between the template and search frames. The advantage of MambaLCT is its ability to continuously extend the length of the context, capturing complete target change cues, which enhances the stability and robustness of the tracker. Extensive experiments show that long-term context information enhances the model's ability to perceive targets in complex scenarios. MambaLCT achieves new SOTA performance on six benchmarks while maintaining real-time runing speeds.
Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Guorong Li, Zhiyi Mo, Shuxiang Song 0001
AAAI6
2025 Robust Tracking via Mamba-based Context-aware Token Learning
abstract
How to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming learning that combining temporal and appearance information by input more and more images (or features). Consequently, these methods not only increase the model's computational source and learning burden but also introduce much useless and potentially interfering information. To alleviate the above issues, we propose a simple yet robust tracker that separates temporal information learning from appearance modeling and extracts temporal relations from a set of representative tokens rather than several images (or features). Specifically, we introduce one track token for each frame to collect the target's appearance information in the backbone. Then, we design a mamba-based Temporal Module for track tokens to be aware of context by interacting with other track tokens within a sliding window. This module consists of a mamba layer with autoregressive characteristic and a cross-attention layer with strong global perception ability, ensuring sufficient interaction for track tokens to perceive the appearance changes and movement trends of the target. Finally, track tokens serve as a guidance to adjust the appearance feature for the final prediction in the head. Experiments show our method is effective and achieves competitive performance on multiple benchmarks at a real-time speed.
Jinxia Xie, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Zhiyi Mo, Shuxiang Song 0001
AAAI6
2025 Less Is More: Token Context-Aware Learning for Object Tracking
abstract
Recently, several studies have shown that utilizing contextual information to perceive target states is crucial for object tracking. They typically capture context by incorporating multiple video frames. However, these naive frame-context methods fail to consider the importance of each patch within a reference frame, making them susceptible to noise and redundant tokens, which deteriorates tracking performance. To address this challenge, we propose a new token context-aware tracking pipeline named LMTrack, designed to automatically learn high-quality reference tokens for efficient visual tracking. Embracing the principle of Less is More, the core idea of LMTrack is to analyze the importance distribution of all reference tokens, where important tokens are collected, continually attended to, and updated. Specifically, a novel Token Context Memory module is designed to dynamically collect high-quality spatio-temporal information of a target in an autoregressive manner, eliminating redundant background tokens from the reference frames. Furthermore, an effective Unidirectional Token Attention mechanism is designed to establish dependencies between reference tokens and search frame, enabling robust cross-frame association and target localization. Extensive experiments demonstrate the superiority of our tracker, achieving state-of-the-art results on tracking benchmarks such as GOT-10K, TrackingNet, and LaSOT.
Chenlong Xu, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Guorong Li, Shuxiang Song 0001
AAAI6
2025 Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
abstract
The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale and diversity of existing tracking datasets. In this work, we present a novel Self-Supervised Tracking framework, named SSTrack, designed to eliminate the need of box annotations. Specifically, a decoupled spatio-temporal consistency training framework is proposed to learn rich target information across timestamps through global spatial localization and local temporal association. This allows for the simulation of appearance and motion variations of instances in real-world scenarios. Furthermore, an instance contrastive loss is designed to learn instance-level correspondences from a multi-view perspective, offering robust instance supervision without additional labels. This new design paradigm enables SSTrack to effectively learn generic tracking representations in a self-supervised manner, while reducing reliance on extensive box annotations. Extensive experiments on nine benchmark datasets demonstrate that SSTrack surpasses SOTA self-supervised tracking methods, achieving an improvement of more than 25.3%, 20.4%, and 14.8% in AUC (AO) score on the GOT10K, LaSOT, TrackingNet datasets, respectively.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Shuxiang Song 0001
AAAI5
2025 Dynamic Updates for Language Adaptation in Visual-Language Tracking
abstract
The consistency between the semantic information provided by the multi-modal reference and the tracked object is crucial for visual-language (VL) tracking. However, existing VL tracking frameworks rely on static multi-modal references to locate dynamic objects, which can lead to semantic discrepancies and reduce the robustness of the tracker. To address this issue, we propose a novel vision-language tracking framework, named DUTrack, which captures the latest state of the target by dynamically updating multimodal references to maintain consistency. Specifically, we introduce a Dynamic Language Update Module, which leverages a large language model to generate dynamic language descriptions for the object based on visual features and object category information. Then, we design a Dynamic Template Capture Module, which captures the regions in the image that highly match the dynamic language descriptions. Furthermore, to ensure the efficiency of description generation, we design an update strategy that assesses changes in target displacement, scale, and other factors to decide on updates. Finally, the dynamic template and language descriptions that record the latest state of the target are used to update the multi-modal references, providing more accurate reference information for subsequent inference and enhancing the robustness of the tracker. DUTrack achieves new state-of-the-art performance on five mainstream vision-language and two vision-only tracking benchmarks, including LaSOT, LaSOText, TNL2K, OTB99-Lang, MGIT, GOT-10K, and UAV123. Code and models are available at https://github.com/GXNU-ZhongLab/DUTrack.
Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Jian Nong, Shuxiang Song 0001
CVPR6
2025 Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking
abstract
Vision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV) tracking which extremely emphasizes efficiency. In this study, we discover that many layers within lightweight ViT-based trackers tend to learn relatively redundant and repetitive target representations. Based on this observation, we propose a similarity-guided layer adaptation approach to optimize the structure of ViTs. Our approach dynamically disables a large number of representation-similar layers and selectively retains only a single optimal layer among them, aiming to achieve a better accuracy-speed trade-off. By incorporating this approach into existing ViTs, we tailor previously complete ViT architectures into an efficient similarity-guided layer-adaptive framework, namely SGLATrack, for real-time UAV tracking. Extensive experiments on six tracking benchmarks verify the effectiveness of the proposed approach, and show that our SGLATrack achieves a state-of-the-art real-time speed while maintaining competitive tracking precision. Codes and models are available at https://github.com/GXNU-ZhongLab/SGLATrack.
Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Ning Li 0044, Yuanliang Xue, Shuxiang Song 0001
CVPR7
2025 Explicit Context Reasoning with Supervision for Visual Tracking
abstract
Contextual reasoning with constraints is crucial for enhancing temporal consistency in cross-frame modeling for visual tracking. However, mainstream tracking algorithms typically associate context by merely stacking historical information without explicitly supervising the association process, making it difficult to effectively model the target's evolving dynamics. To alleviate this problem, we propose RSTrack, which explicitly models and supervises context reasoning via three core mechanisms. 1) Context Reasoning Mechanism : Constructs a target state reasoning pipeline, converting unconstrained contextual associations into a temporal reasoning process that predicts the current representation based on historical target states, thereby enhancing temporal consistency. 2) Forward Supervision Strategy : Utilizes true target features as anchors to constrain the reasoning pipeline, guiding the predicted output toward the true target distribution and suppressing drift in the context reasoning process. 3) Efficient State Modeling : Employs a compression-reconstruction mechanism to extract the core features of the target, removing redundant information across frames and preventing ineffective contextual associations. These three mechanisms collaborate to effectively alleviate the issue of contextual association divergence in traditional temporal modeling. Experimental results show that RSTrack achieves state-of-the-art performance on multiple benchmark datasets while maintaining real-time running speeds. Our code is available at https://github.com/GXNU-ZhongLab/RSTrack.
Fansheng Zeng, Bineng Zhong 0001, Haiying Xia, Yufei Tan, Xiantao Hu, Liangtao Shi, Shuxiang Song 0001
ACM Multimedia7
2025 TEMSA:Text enhanced modal representation learning for multimodal sentiment analysis
Shuxiang Song 0001, Yumei Tan, Haiying Xia
Comput. Vis. Image Underst.2
2025 Context transformer with multiscale fusion for robust facial emotion recognition
Yanling Gan, Luhui Xu, Shuxiang Song 0001, Xiaomei Tao
Pattern Recognit.3
2025 SIEVL-Track: Exploring Semantic Information Enhancement for Visual-Language Object Tracking
abstract
With the assistance of language descriptions, Visual-Language (VL) object tracking can obtain more accurate semantic information compared to traditional Visual-Only object tracking. However, the ability of current VL trackers to obtain target semantic information has not been fully developed due to limitations such as wasted modeling capabilities and insufficient utilization of historical temporal information. On the one hand, the modeling output from Transformer shallow encoders often does not directly participate in the prediction of tracking results, resulting in a certain degree of model capability waste. On the other hand, the semantic information of historical tracking results has also not been fully utilized in the tracking process, resulting in a certain degree of lack of semantic assistance capability. Therefore, we propose a novel hierarchical multi-stage VL tracker called SIEVL-Track to enhance target semantic information. Specifically, we first design a multi-stage visual language tracking framework for modeling multi-scale semantic information in Visual-Language tracking pipeline. Secondly, we propose a selective deep and shallow semantic information fusion module (S-DSFM) that explicitly integrates shallow output features into deep output features, so to reduce the waste of modeling capabilities and obtain more high-frequency semantic information related to the target. Finally, we design a temporal cue modeling module based on linguistic classification and multi-frame historical information(MHLS-TCM), with the aim of more comprehensive utilization of historical temporal semantic information. Benefit from the above designs, our VL tracker can obtain stronger target semantic information. Competitive performance from extensive experimental results on five popular vision-language tracking benchmarks, including LaSOT, OTB99-Lang, WebUAV-3M, LaSOText and TNL2K, have demonstrated the superiority and effectiveness of our SIEVL-Track.
Ning Li 0044, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Jian Nong, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Mamba Adapter: Efficient Multi-Modal Fusion for Vision-Language Tracking
abstract
Utilizing the high-level semantic information of language to compensate for the limitations of vision information is a highly regarded approach in single-object tracking. However, most existing vision-language (VL) trackers employ full-parameter fine-tuning, which can easily lead to catastrophic forgetting. Therefore, they fail to fully exploit the prior knowledge of pre-trained models from upstream tasks, resulting in unsatisfactory tracking performance. To alleviate the above problem, we propose a simple yet effective Vision-Language Tracking pipeline based on Mamba Adapter, named MAVLT, which adopts the idea of parameter-efficient fine-tuning (PEFT) to realize the interaction between vision-language modalities. This novel approach offers the following advantages: (1)The knowledge of the upstream pre-trained model is efficiently inherited by freezing its parameters. This ensures that the VL tracking framework only learns the modules for vision and language interaction, with a focus on the fusion between modalities. (2)The modal interaction between language and vision encoders is flexibly bridged in each encoder layer via proposed mamba adapter, enabling efficient interaction of visual and language information at multiple levels. Extensive experiments on five popular vision-language tracking benchmarks validate the effectiveness of the proposed MAVLT. Particularly, the MAVLT achieves 73.4% AUC score on the LaSOT benchmarks with only 0.18%(0.32M) of the total parameters updates. Code and models are available at https://github.com/GXNU-ZhongLab/MAVLT.
Liangtao Shi, Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Unifying Motion and Appearance Cues for Visual Tracking via Shared Queries
abstract
The rich motion and appearance cues between consecutive frames are crucial for robust visual tracking. However, most existing tracking methods are still limited in designing different components to separately employ corresponding cues and even ignore one of them. This makes them difficult to maintain effective interaction between different cues, thus hindering the models from fostering a comprehensive understanding of the target objects. To address these issues, we propose a unified spatio-temporal cues learning framework (named USCLTrack) that comprehensively mines the variation patterns of targets between consecutive frames in complex video streams. Specifically, USCLTrack firstly aggregates motion and appearance cues into shared queries to provide the bridge of interaction between both cues. Then, it directly generates object locations on the condition of these shared queries in an autoregressive manner, unifying different cues to guide future inferences. To effectively learn multiple spatio-temporal cues aggregated in the shared queries, we develop a spatio-temporal attention mechanism. This mechanism integrates motion cues with appearance cues according to the time steps for ensuring temporal consistency. Moreover, it concurrently captures motion trends and appearance changes to facilitate the understanding of the target objects. Extensive experiments on eight popular tracking benchmarks validate the effectiveness of the proposed USCLTrack.
Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Haiying Xia, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Robust Multi-Stage Tracking via Multi-Scale and Multi-Level Representation Learning
abstract
How to learn multi-scale and multi-level representations is crucial for robust tracking. However, most current one-stream structure based trackers with visual transformers (dubbed ViTs) cannot effectively capture multi-scale representations due to the structure of their adopted ViTs is non-hierarchical. Meanwhile, they often only use the output features from the final layer for predicting results (i.e., ignoring the utilization of low-level features from the shallow layers) which may result in a certain degree of lacking multi-level representation learning ability. To address these issues, we propose a robust multi-stage tracker that effectively combines the advantages of both hierarchical and one-stream structured ViT as a tracking backbone to improve the multi-scale and multi-level representation learning abilities. Specifically, first of all, we design a hierarchical tracker with a three-stage backbone. In the first two stages of our tracker, we utilize a dual-branch structure to obtain multi-scale features of the template and search region separately. Especially, We design the local scale awareness modules based on simple MLP layers to capture multi-scale features. These modules remove complex operations such as convolutions or shifted window attentions, thus avoiding the performance degradation caused by traditional hierarchical ViTs. In the third stage (i.e. the main stage), we construct a global encoder based on the one-stream ViT to achieve efficient feature extraction and feature interaction for our tracker. Then, we design a multi-level feature integration module in the main stage to explicitly utilize the representation information learned from the shallow layers and fuse them with the features of the final layer to obtain multi-level representation information. Lastly, benefit from the these designs, our tracker can effectively capture more multi-scale and multi-level representations for robust tracking. Comprehensive experiments on GOT-10k, LaSOT, LaSOT$_{ext}$, TNL2K, UAV123, TrackingNet and VOT2020 benchmarks validate the effectiveness and robustness of our method.
Ning Li 0044, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Multim.5
2025 Uncertainty-Guided Diffusion Model for Camouflaged Object Detection
abstract
Recently, diffusion models have significantly improved the performance of Camouflaged Object Detection (COD) by adding noise to a mask and iteratively denoising it to match the target distributions. Due to the direct extraction of features from noisy masks and the lack of conditional constraints on a prediction area, the diffusion model may deviate from a correct prediction range and produces mispredictions in regions with high uncertainty. To address this issue, we propose an uncertainty-guided diffusion model (UGDNet) for COD, which explicitly quantifies uncertainty and integrates it as an anchor condition into the diffusion models to provide an initialization of the diffusion regions. The core idea is first to utilize a probability representation and transformer to explicitly model uncertainty, aiming to identify areas where a model may generate overconfident mispredictions. Then, we use the uncertainty as an anchor condition to provide a reference prediction range for the diffusion model, guiding each step of the diffusion process. Furthermore, we use uncertainty to guide feature aggregation, prompting the model to pay extra attention to the semantic features of regions with high uncertainty to refine the segmentation results further. The experimental results indicate that our proposed UGDNet achieves higher accuracy than existing state-of-the-art models on five COD benchmarks, including COD10K, NC4K, CAMO, CHAMELEON, and CDS2K.
Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Shuxiang Song 0001
IEEE Trans. Multim.6
2025 Locating Target Regions for Image Retrieval in an Unsupervised Manner
abstract
Image retrieval performance can be improved by training a convolutional neural network (CNN) model with annotated data to facilitate accurate localization of target regions. However, obtaining sufficiently annotated data is expensive and impractical in real settings. It is challenging to achieve accurate localization of target regions in an unsupervised manner. To address this problem, we propose a new unsupervised image retrieval method named unsupervised target region localization (UTRL) descriptors. It can precisely locate target regions without supervisory information or learning. Our method contains three highlights: 1) we propose a novel zero-label transfer learning method to address the problem of co-localization in target regions. This enhances the potential localization ability of pretrained CNN models through a zero-label data-driven approach; 2) we propose a multiscale attention accumulation method to accurately extract distinguishable target features. It distinguishes the importance of features by using local Gaussian weights; and 3) we propose a simple yet effective method to reduce vector dimensionality, named twice-PCA-whitening (TPW), which reduces the performance degradation caused by feature compression. Notably, TPW is a robust and general method that can be widely applied to image retrieval tasks to improve retrieval performance. This work also facilitates the development of image retrieval based on short vector features. Extensive experiments on six popular benchmark datasets demonstrate that our method achieves about 7% greater mean average precision (mAP) compared to existing state-of-the-art unsupervised methods.
Bo-Jian Zhang, Guanghai Liu 0001, Shuxiang Song 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Robust consistency learning for facial expression recognition under label noise
Yumei Tan, Haiying Xia, Shuxiang Song 0001
Vis. Comput.3
2024 Autoregressive Queries for Adaptive Tracking with Spatio-Temporal Transformers
abstract
The rich spatio-temporal information is crucial to capture the complicated target appearance variations in visual tracking. However, most top-performing tracking algorithms rely on many hand-crafted components for spatio-temporal information aggregation. Consequently, the spatio-temporal information is far away from being fully explored. To alleviate this issue, we propose an adaptive tracker with spatio-temporal transformers (named AQA-Track), which adopts simple autoregressive queries to effectively learn spatio-temporal information without many hand-designed components. Firstly, we introduce a set of learnable and autoregressive queries to capture the instantaneous target appearance changes in a sliding window fashion. Then, we design a novel attention mechanism for the interaction of existing queries to generate a new query in current frame. Finally, based on the initial target template and learnt autoregressive queries, a spatio-temporal information fusion module (STM) is designed for spatiotemporal formation aggregation to locate a target object. Benefiting from the STM, we can effectively combine the static appearance and instantaneous changes to guide robust tracking. Extensive experiments show that our method significantly improves the tracker's performance on six popular tracking benchmarks: LaSOT, LaSOText, TrackingNet, GOT-10k, TNL2K, and UAV123. Code and models will be https://github.com/orgs/GXNU-ZhongLab.
Jinxia Xie, Bineng Zhong 0001, Zhiyi Mo, Shengping Zhang, Liangtao Shi, Shuxiang Song 0001, Rongrong Ji
CVPR6
2024 Diffusion Mask-Driven Visual-language Tracking
Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IJCAI5
2024 DTIL-Net: Dual-Task Interactive Learning Network for Automated Grading of Diabetic Retinopathy and Macular Edema
Yumei Tan, Shuxiang Song 0001, Haiying Xia
PRCV (14)3
2024 Joint pyramidal perceptual attention and hierarchical consistency constraint for gaze estimation
Haiying Xia, Zhuolin Gong, Yumei Tan, Shuxiang Song 0001
Comput. Vis. Image Underst.4
2024 Image retrieval using compact deep semantic correlation descriptors
Bo-Jian Zhang, Guanghai Liu 0001, Shuxiang Song 0001
Inf. Process. Manag.4
2024 Dual-consistency constraints network for noisy facial expression recognition
Haiying Xia, Chunhai Su, Shuxiang Song 0001, Yumei Tan
Image Vis. Comput.3
2024 Learning informative and discriminative semantic features for robust facial expression recognition
Yumei Tan, Haiying Xia, Shuxiang Song 0001
J. Vis. Commun. Image Represent.3
2024 Hard semantic mask strategy for automatic facial action unit recognition with teacher-student model
Zichen Liang, Haiying Xia, Yumei Tan, Shuxiang Song 0001
Multim. Syst.4
2024 Top-Down Cross-Modal Guidance for Robust RGB-T Tracking
abstract
Most RGB-T trackers heavily rely on bottom-up attention and thus overlook top-down cross-modal guidance for learning target features. Consequently, the discriminative power of the learnt target features is weak. To address this issue, we propose a novel RGB-T tracker (called TGTrack) that designs a Top-down Cross-modal Guidance mechanism to learn target features in two stages. In the first stage, our TGTrack effectively generates top-down cross-modal guidance signals with multi-modal encoders-decoders and prior vectors. In the second stage, these signals are transmitted and integrated to improve the discriminative power of our target features by the attention layers of the cross-modal encoders. Moreover, we introduce an Attention-Driven Spatio-Temporal Updater for updating discriminative target features. Through cross-frame attention guidance, it can effectively eliminates irrelevant features within the search region. As a result, our TGTrack can effectively avoid the complex multi-modal fusion modules and thus achieve robust RGB-T tracking. Extensive experiments on three popular RGB-T tracking benchmarks (i.e., LasHeR, RGBT234, and RGBT210) demonstrate that our TGTrack achieves new state-of-the-art performances.
Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Robust Tracking via Combing Top-Down and Bottom-Up Attention
abstract
Transformer attention plays an important role in current top-performing trackers. However, it is bottom-up, driven by stimulus and lacks intrinsic prior guidance. This bottom-up attention mechanism leads to an emphasis on all objects in the input images, rather than the task related objects. As a result, the performance of the bottom-up attention based trackers is deteriorated in complicated scenes. To address this issue, we propose a robust tracker that combines bottom-up attention with top-down attention to comply with the existing ViT framework, named TBTrack. TBTrack can not only utilize the existing bottom-up attention mechanisms to model the long-range relationship of input tokens, but also utilize a newly added top-down attention mechanism to pay more attention to task related object and further eliminate interference from similar objects and backgrounds. Specifically, we firstly design a top-down prior generation module using an adaptive learning parameter combined with the template inputs to obtain top-down task guided signals. Then, we inject the prior signals into a bottom-up attention module to obtain a top-down and bottom-up attention combination block (TB-Block). Finally, we stack these TB-Blocks to construct our tracker (TBTrack) with top-down prior guidance capability, which focuses more on the task related object. Through extensive experiments, our TBTrack achieves impressive performance on multiple tracking benchmarks, including GOT-10k, LaSOT, LaSOText, TNL2K, TrackingNet, UAV123 and so on. The code and trained models will be publicly available.
Ning Li 0044, Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 One-Stream Stepwise Decreasing for Vision-Language Tracking
abstract
Based on the fixed language descriptions in the initial frames, a vision-language tracker typically adopts a two-stream model structure to align vision and language features at the feature fusion stages. However, this paradigm may degrade the tracking performance due to inaccurate language descriptions and lacks further modal interaction. To address these issues, we propose a one-stream vision-language model called One-stream Stepwise Decreasing for Vision-Language Tracking (OSDT). Specifically, we first encode the language description using a language encoder. The obtained language features are then combined with visual images and entered jointly into a visual encoder, in which the encoder’s self-attention mechanism is utilized to facilitate more interactions between language and visual features. Moreover, to mitigate the problems caused by inaccurate language descriptions, we design a stepwise decreasing multi-modal interaction framework, in which a Feature Filter Module (FFM) is introduced to select language features that are more relevant to visual information to provide semantic guidance for visual feature extraction. Furthermore, without additional feature fusion modules, our one-stream model framework can efficiently utilize the proposed feature filtering module for feature selection. Consequently, our tracker can achieve fast tracking speed in the vision-language tracking domain compared to existing state-of-the-art methods. We extensively evaluate our tracker on three benchmarks, i.e. TNL2K, LaSOT, and OTB99, demonstrating competing performance compared to state-of-the-art vision-language tracking methods.
Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Ning Li 0044, Shuxiang Song 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Feature fusion of multi-granularity and multi-scale for facial expression recognition
Haiying Xia, Lidan Lu, Shuxiang Song 0001
Vis. Comput.3
2023 ST-VQA: shrinkage transformer with accurate alignment for visual question answering
Haiying Xia, Richeng Lan, Hai-Sheng Li 0001, Shuxiang Song 0001
Appl. Intell.4
2023 RT-Net: Region-Enhanced Attention Transformer Network for Polyp Segmentation
Yilin Qin, Haiying Xia, Shuxiang Song 0001
Neural Process. Lett.3
2023 Quantum Bilinear Interpolation Algorithms Based on Geometric Centers
abstract
Bilinear interpolation is widely used in classical signal and image processing. Quantum algorithms have been designed for efficiently realizing bilinear interpolation. However, these quantum algorithms have limitations in circuit width and garbage outputs, which block the quantum algorithms applied to noisy intermediate-scale quantum devices. In addition, the existing quantum bilinear interpolation algorithms cannot keep the consistency between the geometric centers of the original and target images. To save the above questions, we propose quantum bilinear interpolation algorithms based on geometric centers using fault-tolerant implementations of quantum arithmetic operators. Proposed algorithms include the scaling-up and scaling-down for signals (grayscale images) and signals with three channels (color images). Simulation results demonstrate that the proposed bilinear interpolation algorithms obtain the same results as their classical counterparts with an exponential speedup. Performance analysis reveals that the proposed bilinear interpolation algorithms keep the consistency of geometric centers and significantly reduce circuit width and garbage outputs compared to the existing works.
Hai-Sheng Li 0001, Jinhui Quan, Shuxiang Song 0001, Li Qing
ACM Trans. Quantum Comput.3
2022 HT-Net: hierarchical context-attention transformer network for medical ct image segmentation
Haiying Xia, Yumei Tan, Hai-Sheng Li 0001, Shuxiang Song 0001
Appl. Intell.5
2022 MC-Net: multi-scale context-attention network for medical CT image segmentation
Haiying Xia, Hai-Sheng Li 0001, Shuxiang Song 0001
Appl. Intell.4
2022 MFC-Net: Multi-scale fusion coding network for Image Deblurring
Haiying Xia, Yumei Tan, Shuxiang Song 0001
Appl. Intell.5
2022 HRNet: A hierarchical recurrent convolution neural network for retinal vessel segmentation
Haiying Xia, Lingyu Wu, Hai-Sheng Li 0001, Shuxiang Song 0001
Multim. Tools Appl.5
2022 Multilevel 2-D Quantum Wavelet Transforms
abstract
Wavelet transform is being widely used in classical image processing. One-dimension quantum wavelet transforms (QWTs) have been proposed. Generalizations of the 1-D QWT into multilevel and multidimension have been investigated but restricted to the quantum wavelet packet transform (QWPTs), which is the direct product of 1-D QWPTs, and there is no transform between the packets in different dimensions. A 2-D QWT is vital for image processing. We construct the multilevel 2-D QWT's general theory. Explicitly, we built multilevel 2-D Haar QWT and the multilevel Daubechies D4 QWT, respectively. We have given the complete quantum circuits for these wavelet transforms, using both noniterative and iterative methods. Compared to the 1-D QWT and wavelet packet transform, the multilevel 2-D QWT involves the entanglement between components in different degrees. Complexity analysis reveals that the proposed transforms offer exponential speedup over their classical counterparts. Also, the proposed wavelet transforms are used to realize quantum image compression. Simulation results demonstrate that the proposed wavelet transforms are significant and obtain the same results as their classical counterparts with an exponential speedup.
Hai-Sheng Li 0001, Huiling Peng, Shuxiang Song 0001, Gui-Lu Long 0001
IEEE Trans. Cybern.4
2021 A multi-scale segmentation-to-classification network for tiny microaneurysm detection in fundus images
Haiying Xia, Shuxiang Song 0001, Hai-Sheng Li 0001
Knowl. Based Syst.3
2020 Combination of multi-scale and residual learning in deep CNN for image denoising
abstract
To better restore a clean image from a noise observation under high noise levels, the authors propose an image denoising network based on the combination of multi‐scale and residual learning. Instead of using filters with different large sizes in traditional multi‐scale schemes, they arrange multi‐layer convolutions with the filters of the same size to speed up the model. Some dilated convolutions of different rates are combined with the common convolutions to enrich the extracted features in multi‐layer convolutions. Furthermore, they cascade the multi‐layer convolutions with residual blocks to improve the performance of image denoising. Their extensive evaluations on several challenging datasets demonstrate that the proposed model outperforms the state‐of‐art methods under all different noise levels in terms of peak signal‐to‐noise ratio, and the visual effects achieved by the proposed model are also better than the competing methods.
Haiying Xia, Fuyu Zhu, Hai-Sheng Li 0001, Shuxiang Song 0001, Xiangwei Mou
IET Image Process.4
2020 Md-Net: Multi-scale Dilated Convolution Network for CT Images Segmentation
Haiying Xia, Weifan Sun, Shuxiang Song 0001, Xiangwei Mou
Neural Process. Lett.3
2019 Content-Based Image Retrieval Using Color Volume Histograms
abstract
Human visual perception has a close relationship with the HSV color space, which can be represented as a cylinder. The question of how visual features are extracted using such an attribute is important. In this paper, a new feature descriptor; namely, a color volume histogram, is proposed for image representation and content-based image retrieval. It converts a color image from RGB color space to HSV color space and then uniformly quantizes it into 72 bins of color cues and 32 bins of edge cues. Finally, color volumes are used to represent the image content. The proposed algorithm is extensively tested on two Corel datasets containing 15[Formula: see text]000 natural images. These image retrieval experiments show that the color volume histogram has the power to describe color, texture, shape and spatial features and performs significantly better than the local binary pattern histogram and multi-texton histogram approaches.
Ji-Zhao Hua, Guanghai Liu 0001, Shuxiang Song 0001
Int. J. Pattern Recognit. Artif. Intell.3
2019 Quantum multi-level wavelet transforms
Hai-Sheng Li 0001, Haiying Xia, Shuxiang Song 0001
Inf. Sci.4
2019 Quantum vision representations and multi-dimensional quantum transforms
Hai-Sheng Li 0001, Shuxiang Song 0001, Huiling Peng, Haiying Xia
Inf. Sci.2
2018 Fast Single Image De-raining via a Weighted Residual Network
Ruibin Zhuge, Haiying Xia, Hai-Sheng Li 0001, Shuxiang Song 0001
ICONIP (6)4