Xiangyue Zhang

dblp:250/1932 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints
abstract
Reliable co-speech motion generation requires precise motion representation and consistent structural priors across all joints. Existing generative methods typically operate on local joint rotations, which are defined hierarchically based on the skeleton structure. This leads to cumulative errors during generation, manifesting as unstable and implausible motions at end-effectors. In this work, we propose GlobalDiff, a diffusion-based framework that operates directly in the space of global joint rotations for the first time, fundamentally decoupling each joint’s prediction from upstream dependencies and alleviating hierarchical error accumulation. To compensate for the absence of structural priors in global rotation space, we introduce a multi-level constraint scheme. Specifically, a joint structure constraint introduces virtual anchor points around each joint to better capture fine-grained orientation. A skeleton structure constraint enforces angular consistency across bones to maintain structural integrity. A temporal structure constraint utilizes a multi-scale variational encoder to align the generated motion with ground-truth temporal patterns. These constraints jointly regularize the global diffusion process and reinforce structural awareness. Extensive evaluations on standard co-speech benchmarks show that GlobalDiff generates smooth and accurate motions, improving the performance by 46.0% compared to the current SOTA under multiple speaker identities.
Xiangyue Zhang, Jianqiang Ren
AAAI1
2026 A fault-tolerant target tracking localization algorithm based on extended dimension cubature kalman filter and variational bayesian
Zedong Liang, Jinliang Ding, Hongli Xu 0003, Xiangyue Zhang, Qi Liu 0027
Signal Process.7
2025 SemTalk: Holistic Co-Speech Motion Generation with Frame-Level Semantic Emphasis
abstract
Co-speech gesture generation must carefully integrate common rhythmic motion with rare yet essential semantic gestures. In this work, we propose SemTalk for holistic co-speech gesture generation with frame-level semantic emphasis. Our key insight is to separately learn base motions and sparse motions, and then adaptively fuse them. In particular, coarse2fine cross-attention module and rhythmic consistency learning are explored to establish rhythm-related base motion, ensuring a coherent foundation that synchronizes gestures with the speech rhythm. Subsequently, semantic emphasis learning is designed to generate semantic-aware sparse motion, focusing on frame-level semantic cues. Finally, to integrate sparse motion into the base motion and generate semantic-emphasized co-speech gestures, we further leverage a learned semantic score for adaptive synthesis. Qualitative and quantitative comparisons on two public datasets demonstrate that our method outperforms the state-of-the-art, delivering high-quality co-speech motion with enhanced semantic richness over a stable base motion.
Xiangyue Zhang, Jianfang Li 0001, Ziqiang Dang, Jianqiang Ren, Liefeng Bo, Zhigang Tu 0001
ICCV1
2025 EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
Xiangyue Zhang, Jianfang Li 0001, Jianqiang Ren, Liefeng Bo, Zhigang Tu 0001
ACM Multimedia1
2025 FATL: Frozen-feature augmentation transfer learning for few-shot long-tailed sonar image classification
Zhongyu Bai, Hongli Xu 0003, Qichuan Ding, Xiangyue Zhang
Neurocomputing4
2025 MFEGPT: A PID controller-inspired multimodal feature enhancement for MLLMs
Hexiao Li, Zixi Jia, Xiangyue Zhang, Chuanjiang Leng, Heyao Li, Hongbin Gao, Jiqiang Liu
Neurocomputing3
2025 Conditional variational underwater image enhancement with kernel decomposition and adaptive hybrid normalization
Haopeng Zhang 0018, Hongli Xu 0003, Hao Liu 0008, Xiaosheng Yu 0001, Xiangyue Zhang, Chengdong Wu 0001
Neurocomputing5
2025 Robust 2D Skeleton Action Recognition via Decoupling and Distilling 3D Latent Features
abstract
Human skeletons provide a compact representation for action recognition. Compared to 3D skeletons, 2D skeletons lack view-independence and depth, making them less robust for motion analysis. However, 3D skeleton data requires specialized hardware, limiting its practicality, especially in outdoor or dynamic settings. In contrast, 2D skeletons can be extracted from standard RGB videos, making them more accessible. To address this, we propose 2D³-SkelAct, a 2D skeleton-based action recognition model. It maps 2D inputs to a 3D latent space, where pose and view features are decoupled. Additionally, 2D³-SkelAct distills motion cues from 3D models, enhancing motion detail capture while keeping the benefits of 2D data. Specifically, the pipeline of our 2D3-SkelAct consists of two steps:pose-view decouplingandpose-view distilling. First, we use a spatio-temporal transformer to decouple 2D skeleton sequences into latent pose and view features, enhancing the model’s ability to learn motion dynamics. Next, these decoupled features are separately integrated into the 2D skeleton model through two cross-attention modules, allowing it to extract discriminative motion features while mitigating uncertainties in 3D viewpoint and depth. Additionally, we distill motion cues from 3D models to compensate for the limitations of 2D skeletons. Remarkably, our model can seamless integrate with various skeleton feature extractors. We validate the proposed 2D3-SkelAct through extensive experiments, demonstrating its adaptability across different model architectures as where consistent improvement achieving. When combined with advanced skeleton feature extractors, 2D3-SkelAct achieves state-of-the-art performance in 2D skeleton-based action recognition.
Xiangyue Zhang, Yifan Jia 0007, Zhigang Tu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 CDF-UIE: Leveraging Cross-Domain Fusion for Underwater Image Enhancement
abstract
Underwater image enhancement (UIE) aims to restore image quality by mitigating inherent degradations in underwater imaging systems. While existing learning-based methods show promise, they face limitations in separating and processing frequency components, effectively fusing domain information, and balancing the enhancement of structures and details. To resolve these limitations, we propose cross-domain fusion (CDF)-UIE, a novel network that leverages and fuses cross-domain information for mitigating the degradation in underwater images. CDF-UIE first performs domain decoupling of input features using the proposed spatial-frequency decoupling (SFD) block. Then, we design an innovative CDF block, which effectively bridges the spatial- and frequency-domain features through the cross-domain attention mechanism. To produce stable and detailed enhanced outputs, we exploit the coarse and fine-scale information in the image reconstruction stage. In addition, we introduce a multiscale objective function that incorporates pixel-level, structural, and perceptual constraints to guide the enhancement process. We conduct extensive experiments on six diverse real-world underwater image datasets. Comprehensive experiments and real-world application tests demonstrate that CDF-UIE significantly outperforms existing methods, offering promising future applications in various underwater scenarios. The source code is available athttps://github.com/hpzhan66/CDF-UIE.
Haopeng Zhang 0018, Hongli Xu 0003, Xiaosheng Yu 0001, Xiangyue Zhang, Xiujing Gao, Chengdong Wu 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 A Gradient Vector Self-Learning Network for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) is challenging due to low target-background contrast and small target size, leading to missed detections and false alarms (Fas). To address these problems, a novel gradient vector self-learning network (GVSLNet) is proposed. First, the self-learning gradient vector (SLGV) module is designed based on the unique high correlation of infrared small target gradient. Traditional gradient vector field operators cannot update and learn the deep features. SLGV overcomes these limitations by using CNN to adaptively learn and calculate the gradient vector of infrared images, improving the ability to distinguish between targets and backgrounds under complex environments. Then, edge features are encoded into the global-local attention fusion (GLAF) module, which is based on Transformer and dilated convolution. Infrared small target images often display nonlocal self-similarity, where background signals tend to share similar structures. The GLAF module leverages this characteristic to further enhance target intensity while effectively suppressing background noise. The proposed GVSLNet can dynamically calculate gradients based on scene information, improving the detection ability of small targets in complex environments. Experimental results prove that the proposed GVSLNet outperforms state-of-the-art methods on public datasets while maintaining high inference speed.
Xiangyue Zhang, Xinhao Zheng, Chengdong Wu 0001, Jingyu Ru
IEEE Geosci. Remote. Sens. Lett.1
2024 GCSANet: Arbitrary Style Transfer With Global Context Self-Attentional Network
abstract
Arbitrary style transfer is attracting increasing attention in the computer vision community due to its application flexibility. Existing approaches directly fuse deep style features with deep content features or adaptively normalize content features for global statistical matching. Although effective, it is prone to suffer from local unnatural outputs and artifacts owing to the lack of exploring the global contextual semantic distribution of style image features. In this article, a novel global context self-attentional network (GCSANet) is proposed to efficiently generate high-quality stylized results based on the global semantic spatial distributions of style images. First, a context modeling module is proposed to aggregate the depth features of style images into global context features. Then, channel-wise interdependencies are captured with the feature transform module. Finally, the style features are appropriately aggregated to each location of the content image. In addition, novel external contrastive losses are proposed to balance the distribution of content and style features to ensure the reasonableness of the texture patterns in the stylized images. The ablation studies validate the effectiveness of the proposed components. Various quantitative and qualitative experiments demonstrate the superiority of our method for real-time arbitrary image/video style transfer.
Zhongyu Bai, Hongli Xu 0003, Xiangyue Zhang, Qichuan Ding
IEEE Trans. Multim.3
2023 MFSANet: Zero-Shot Side-Scan Sonar Image Recognition Based on Style Transfer
abstract
Side-scan sonar (SSS) is attracting increasing attention in ocean exploration for its utility and stability on autonomous underwater vehicles (AUVs). Existing SSS image recognition methods mainly employ deep neural networks (DNNs) for various tasks. However, the effectiveness of DNN-based approaches is limited in situations with zero samples. In this letter, a multifeature fusion self-attention network (MFSANet) is proposed to generate SSS images of novel categories, transforming this problem into a conventional supervised learning problem. Specifically, optical-acoustic image pairs are used as inputs to the network to synthesize pseudo-SSS images. First, the shallow and deep features of the input images are extracted by different layers of the encoder. Then, the long-range dependencies of the acoustic images are efficiently modeled with the proposed simplified self-attention module (SSAM). Finally, the acoustic features are appropriately aggregated to each position of the optical features to efficiently generate pseudo-SSS samples for training the classification network. In addition, a novel contrastive loss is proposed to optimize the cross-modal feature space distribution. Experimental results demonstrate that our method can efficiently generate high-quality pseudo-SSS samples, which improves the accuracy of zero-shot SSS image recognition.
Hongli Xu 0003, Zhongyu Bai, Xiangyue Zhang, Qichuan Ding
IEEE Geosci. Remote. Sens. Lett.3
2023 Infrared Small Target Detection Based on Gradient Correlation Filtering and Contrast Measurement
abstract
Infrared small target detection under complex backgrounds, especially in dense cloud and changeable clutter scenes, has always been a challenging research task. In order to improve the detection ability of small targets under complex backgrounds, an infrared small target detection method based on gradient correlation filtering and gradient contrast measurement (GCF-CM) is proposed in this article. The infrared gradient vector field (IGVF) of the original image is first constructed through the facet model. Then, considering the unique gradient characteristics of small targets, a gradient correlation filtering (GCF) method is proposed to filter small targets and background clutters. Meanwhile, a gradient contrast measurement (GCM) method is designed to further enhance the intensity of the small target. Finally, after fusing the two response maps, an adaptive threshold is adopted to extract small targets. Experimental results demonstrate that the proposed method can improve the intensity of the small target and suppress clutter sufficiently. In comparison with other excellent methods, the proposed method exhibits a robust detection performance.
Xiangyue Zhang, Jingyu Ru, Chengdong Wu 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Skeleton-based similar action recognition through integrating the salient image feature into a center-connected graph convolutional network
Zhongyu Bai, Qichuan Ding, Hongli Xu 0003, Jianning Chi, Xiangyue Zhang, Tiansheng Sun
Neurocomputing5
2022 An Infrared Small Target Detection Method Based on Gradient Correlation Measure
abstract
To overcome the interference of complex background and improve the detection ability of infrared small target under low signal-to-clutter ratio (SCR) scenes, a novel detection method based on gradient correlation measure (GCM) is proposed in this letter. Initially, the infrared gradient vector field (IGVF) of the original image is constructed based on the facet model. Then, a gradient correlation template is designed to distinguish the difference of local gradient between small targets and background. Finally, an adaptive threshold is adopted to extract small targets from background clutter. The proposed GCM method can identify the unique gradient characteristics of small targets. Experimental evaluations prove that the proposed method can achieve higher SCR scores in complex backgrounds. Especially in the scene where the gray contrast of small targets is low, the proposed GCM method shows a more robust detection performance.
Xiangyue Zhang, Jingyu Ru, Chengdong Wu 0001
IEEE Geosci. Remote. Sens. Lett.1