Xianghu Yue

dblp:199/5796 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-3527-6034ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding Alignment
abstract
Dysarthric speech reconstruction (DSR) aims to enhance the intelligibility of dysarthric speech. Compared with normal speech, the dysarthric speech is characterized by its pathological features, including discontinuous pronunciation, slow speech, hoarseness, and improper pauses. Significant disparities in the feature space between normal and dysarthric speech may result in suboptimal speech reconstruction, thereby degrading speech intelligibility. To enhance the reconstruction ability of speech feature spaces, this paper proposes a DSR model named the Encoding-Aligned Variational Autoencoder (EA-VAE). By incorporating alignment modules of frame-level embedding features, prior distributions, and duration into the encoder of the VAE, the model explicitly aligns the dysarthric speech encoding with a representation of the parallel normal speech. A shared decoder is then used to generate speech with improved intelligibility. Experimental results on the UASpeech benchmark confirm that EA-VAE achieves state-of-the-art performance, with a 31.7% relative word error rate reduction and the highest subjective MOS score (4.48), thoroughly validating the effectiveness and advancements of the proposed method in dysarthric speech reconstruction.
Daipeng Zhang 0001, Wenhuan Lu, Xianghu Yue, Hongcheng Zhang, Jianguo Wei
AAAI3
2026 NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language Models
abstract
LLM serving is limited by provider-side resources: longer generations consume more GPU time, increase latency, and reduce throughput in multi-tenant systems.This creates a denial-of-service (DoS) risk, where attackers degrade service by inducing excessive generation.Prior work on LLM DoS primarily relies on adversarial perturbations that delay end-of-sequence termination.We show perturbations are often unnecessary: natural, benignlooking instructions that specify impractical and meaningless tasks can already trigger excessive generation.To study this overlooked vulnerability, we introduce NaturalSloth, an adversarial dataset of natural, instruction-based DoS prompts.Starting from a human-curated seed set spanning diverse attack categories, we design a multi-agent synthesis framework to scale the dataset while preserving malicious intent and increasing semantic diversity.Experiments across a wide range of proprietary and open-source LLMs show that NaturalSloth consistently induces excessive generation, with attack effectiveness further amplified when combined with jailbreak techniques.Our analysis also reveals significant limitations of existing defenses, highlighting the need for dedicated protections against natural DoS attacks. 1
Yiming Chen 0010, Zexin Li 0001, Xianghu Yue, Robby T. Tan, Haizhou Li 0001
ACL (1)3
2026 HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
abstract
Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks.However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain.Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music.HalluAudio comprises over 5K humanverified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA.To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions.Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes.We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music.Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs.
Feiyu Zhao, Wenhuan Lu, Daipeng Zhang 0001, Xianghu Yue, Jianguo Wei
ACL (1)5
2026 PAL: Prompting analytic learning with missing modality for multi-modal class-incremental learning
Xianghu Yue, Yiming Chen 0010, Xueyi Zhang 0001, Xiaoxue Gao, Mengling Feng, Mingrui Lao, Huiping Zhuang, Haizhou Li 0001
Pattern Recognit.1
2026 Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction
abstract
Image retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task and a framework that learns the relationship between images and multi-speaker speech signals in the presence of a target speaker. Our method integrates pre-trained self-supervised audio encoders with vision models via target speaker-aware contrastive learning, conditioned on a Target Speaker Retrieval Extractor (TSRE) module. This method enables the extraction of semantic content from the target speaker's speech and aligns it with images representing the corresponding semantic meaning. Experiments on SpokenCOCO2Mix and SpokenCOCO3Mix show that TSRE significantly outperforms existing methods, achieving 36.3% and 29.9% Recall@1 in 2- and 3-speaker scenarios, respectively-substantial improvements over single-speaker baselines and state-of-the-art models. Our approach demonstrates potential for real-world deployment in assistive robotics and multimodal interaction systems.
Jianguo Wei, Wenhuan Lu, Xinyue Song, Xianghu Yue
IEEE Signal Process. Lett.5
2026 VoiceBench: Benchmarking LLM-Based Voice Assistants
abstract
Abstract Recent advancements in large language models (LLMs) like GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering an improved user experience over text-based interactions. However, a suitable benchmark to rigorously evaluate such speech interactions systems is currently lacking. To bridge this gap, we introduce VoiceBench, the first benchmark specifically designed to assess LLM-based voice assistants. VoiceBench comprises 6,783 synthetic and real spoken instructions recorded from diverse speakers across eight distinct tasks. These instructions are meticulously crafted to assess three crucial capability areas: general knowledge, instruction-following, and safety compliance. Furthermore, VoiceBench systematically incorporates realistic variations common in spoken interactions, including differences in speaker characteristics (e.g., accents), heterogeneous environmental conditions (e.g., reverberation), and content complexities such as mispronunciations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.1
Yiming Chen 0010, Xianghu Yue, Chen Zhang 0020, Xiaoxue Gao, Robby T. Tan, Haizhou Li 0001
Trans. Assoc. Comput. Linguistics2
2025 UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
abstract
Yidi Jiang, Qian Chen, Shengpeng Ji, Yu Xi, Wen Wang, Chong Zhang, Xianghu Yue, ShiLiang Zhang, Haizhou Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yidi Jiang, Qian Chen 0003, Shengpeng Ji, Yu Xi, Wen Wang 0001, Chong Zhang 0003, Xianghu Yue, Shiliang Zhang, Haizhou Li 0001
ACL (1)7
2025 EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph Modeling
abstract
Event cameras, with their microsecond-level temporal resolution and sparse visual encoding, provide a transformative paradigm for automatic lip reading (ALR). However, event data inherently lack explicit spatial structure and exhibit a pronounced frequency-domain bias. The low-frequency components fail to capture crucial lip structural information, which fundamentally impedes the modeling of intra-frame topological dependencies and inter-frame semantic evolution-both of which are critical for robust lip reading. To this end, we propose FAST-HG, a Frequency-Aware SpatioTemporal HyperGraph framework specifically designed for event-based lip reading. First, we apply low-frequency perturbation to improve the model's robustness for capturing discriminative features, and integrate adaptive high-frequency filtering to enhance edge-aware representations. Then, we construct a Spatial Region Hypergraph (SRH) and a Temporal Semantic Hypergraph (TSH). The former captures intra-frame topological dependencies among lip regions, while the latter explicitly models inter-frame structural associations throughout the lip movement process, enabling the model to capture discriminative patterns in lip dynamics. Furthermore, we propose a viseme-aware label smoothing strategy, where a novel viseme-level edit distance is designed to quantify visual similarities between classes and guide the construction of soft labels. FAST-HG achieves 79.85% and 84.03% accuracy on the DVS-Lip and DVS-LRW100 datasets, respectively, significantly outperforming prior methods and establishing a new benchmark for event-based lip reading.
Xueyi Zhang 0001, Jialu Sun, Xianghu Yue, Tianfang Xiao, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001
ACM Multimedia4
2025 Analytic Class Incremental Learning for Sound Source Localization With Privacy Protection
abstract
Sound Source Localization (SSL) enabling technology for applications such as surveillance and robotics. While traditional Signal Processing (SP)-based Sound Source Localization (SSL) methods provide analytic solutions under specific signal and noise assumptions, recent Deep Learning (DL)-based methods have significantly outperformed them. However, their success depends on extensive training data and substantial computational resources. Moreover, they often rely on large-scale annotated spatial data and may struggle to adapt to evolving sound classes. To mitigate these challenges, we propose a novel Class Incremental Learning (CIL) approach, termed SSL-CIL, which avoids serious accuracy degradation due tocatastrophic forgettingby incrementally updating the DL-based SSL model through a closed-form analytic solution. In particular, data privacy is ensured since the learning process does not revisit any historical data (exemplar-free), which is more suitable for smart home scenarios. Empirical results in the public SSLR dataset demonstrate the superior performance of our proposal, achieving a localization accuracy of 90.9%, surpassing other competitive methods.
Xinyuan Qian 0001, Xianghu Yue, Huiping Zhuang, Haizhou Li 0001
IEEE Signal Process. Lett.2
2024 MMAL: Multi-Modal Analytic Learning for Exemplar-Free Audio-Visual Class Incremental Tasks
abstract
Class-incremental learning poses a significant challenge under an exemplar-free constraint, leading to catastrophic forgetting and sub-par incremental accuracy. Previous attempts have focused primarily on single-modality tasks, such as image classification or audio event classification. However, in the context of Audio-Visual Class-Incremental Learning (AVCIL), the effective integration and utilization of heterogeneous modalities, with their complementary and enhancing characteristics, remains largely unexplored. To bridge this gap, we propose the Multi-Modal Analytic Learning (MMAL) framework, an exemplar-free solution for AVCIL that employs a closed-form, linear approach. To be specific, MMAL introduces a modality fusion module that re-formulates the AVCIL problem through a Recursive Least-Square (RLS) perspective. Complementing this, a Modality-Specific Knowledge Compensation (MSKC) module is designed to further alleviate the under-fitting limitation intrinsic to analytic learning by harnessing individual knowledge from audio and visual modality in tandem. Comprehensive experimental comparisons with existing methods show that our proposed MMAL demonstrates superior performance with the accuracy of 76.71%, 78.98%, and 76.19% on AVE, Kinetics-Sounds, and VGGSounds100 datasets, respectively, setting new state-of-the-art AVCIL performance. Notably, compared to those memory-based methods, our MMAL, being an exemplar-free approach, provides good data privacy and can better leverage multi-modal information for improved incremental accuracy.
Xianghu Yue, Xueyi Zhang 0001, Yiming Chen 0010, Mingrui Lao, Huiping Zhuang, Xinyuan Qian 0001, Haizhou Li 0001
ACM Multimedia1
2024 PD-Refiner: An Underlying Surface Inheritance Refiner with Adaptive Edge-Aware Supervision for Point Cloud Denoising
abstract
Point clouds from real-world scenarios inevitably contain complex noise, significantly impairing the accuracy of downstream tasks. To tackle this challenge, cascading encoder-decoder architecture has become a conventional technical route to iterative denoise. However, circularly feeding the output of denoiser as its input again involves the re-extraction of underlying surface, leading to unstable denoising process and over-smoothed geometric details. To address these issues, we propose a novel denoising paradigm dubbed PD-Refiner that employs a single encoder to model the underlying surface. Then, we leverage several lightweight hierarchical Underlying Surface Inheritance Refiners (USIRs) to inherit and strengthen it, thereby avoiding the re-extraction from the intermediate point cloud. Furthermore, we design adaptive edge-aware supervision to improve the edge awareness of the USIRs, allowing for the adjustment of the denoising preferences from global structure to local details. The results demonstrate that our method not only achieves state-of-the-art performance in terms of denoising stability and efficacy, but also enhances edge clarity and point cloud uniformity.
Xueyi Zhang 0001, Xianghu Yue, Mingrui Lao, Tao Jiang 0062, Fubo Zhang, Longyong Chen
ACM Multimedia3
2024 Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading
abstract
Lip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy environments. With the increasing interpersonal communications in social media owing to globalization, the existing monolingual datasets for lip reading may not be sufficient to meet the exponential proliferation of bilingual and even multilingual users. However, to our best knowledge, research on code-switching is only explored in speech recognition, while the attempts in lip reading are seriously neglected. To bridge this gap, we have collected a bilingual code-switching lip reading benchmark composed of Chinese and English, dubbed CSLR. As the pioneering work, we recruited 62 speakers with proficient foundations in bothspoken Chinese and English to express sentences containing both involved languages. Through rigorous criteria in data selection, CSLR benchmark has accumulated 85,560 video samples with a resolution of 1080x1920, totaling over 71.3 hours of high-quality code-switching lip movement data. To systematically evaluate the technical challenges in CSLR, we implement commonly-used lip reading backbones, as well as competitive solutions in code-switching speech for benchmark testing. Experiments show CSLR to be a challenging and under-explored lip reading task. We hope our proposed benchmark will extend the applicability of code-switching lip reading, and further contribute to the communities of cross-lingual communication and collaboration. Our dataset and benchmark are accessible at https://github.com/cslr-lipreading/CSLR.
Xueyi Zhang 0001, Mingrui Lao, Jun Tang 0001, Yanming Guo, Siqi Cai 0002, Xianghu Yue, Haizhou Li 0001
NeurIPS7
2024 Text-Guided HuBERT: Self-Supervised Speech Pre-Training via Generative Adversarial Networks
abstract
Human language can be expressed in either written or spoken form, i.e. text or speech. Humans can acquire knowledge from text to improve speaking and listening. However, the quest for speech pre-trained models to leverage unpaired text has just started. In this letter, we investigate a new way to pre-train such a joint speech-text model to learn enhanced speech representations and benefit various speech-related downstream tasks. Specifically, we propose a novel pre-training method, text-guided HuBERT, or T-HuBERT, which performs self-supervised learning over speech to derive phoneme-like discrete representations. And these phoneme-like pseudo-label sequences are firstly derived from speech via the generative adversarial networks (GAN) to be statistically similar to those from additional unpaired textual data. In this way, we build a bridge between unpaired speech and text in a unsupervised manner. Extensive experiments demonstrate the significant superiority of our proposed method over various strong baselines, which achieves up to 15.3% relative Word Error Rate (WER) reduction on the LibriSpeech dataset.
Duo Ma, Xianghu Yue, Junyi Ao, Xiaoxue Gao, Haizhou Li 0001
IEEE Signal Process. Lett.2
2023 Self-Transriber: Few-Shot Lyrics Transcription With Self-Training
abstract
The current lyrics transcription approaches heavily rely on supervised learning with labeled data, but such data are scarce and manual labeling of singing is expensive. How to benefit from unlabeled data and alleviate limited data problem have not been explored for lyrics transcription. We propose the first semi-supervised lyrics transcription paradigm, Self-Transcriber, by leveraging on unlabeled data using selftraining with noisy student augmentation. We attempt to demonstrate the possibility of lyrics transcription with a few amount of labeled data. Self-Transcriber generates pseudo labels of the unlabeled singing using teacher model, and augments pseudo-labels to the labeled data for student model update with both self-training and supervised training losses. This work closes the gap between supervised and semi- supervised learning as well as opens doors for few-shot learning of lyrics transcription. Our experiments show that our approach using only 12.7 hours of labeled data achieves competitive performance compared with the supervised approaches trained on 149.1 hours of labeled data for lyrics transcription.
Xiaoxue Gao, Xianghu Yue, Haizhou Li 0001
ICASSP2
2023 Token2vec: A Joint Self-Supervised Pre-Training Framework Using Unpaired Speech and Text
abstract
Self-supervised pre-training has been successful in both text and speech processing. Speech and text offer different but complementary information. The question is whether we are able to perform a speech-text joint pre-training on unpaired speech and text. In this paper, we take the idea of self-supervised pre-training one step further and propose token2vec, a novel joint pre-training framework for unpaired speech and text based on discrete representations of speech. Specifically, we introduce two modality-specific tokenizers for speech and text. Based on these tokenizers, we convert speech/text sequences into discrete speech/text token sequences consisting of similar language units, thus mitigating the domain mismatch problem and length mismatch problem, which are caused by the distinct characteristics between speech and text. Finally, we feed the discrete speech and text tokens into a modality-agnostic Transformer encoder and pre-train with token-level masking language modeling (tMLM). Experiments show that token2vec is significantly superior to various speech-only pre-training baselines, with up to 17.7% relative WER reduction. Token2vec model is also validated on a non-ASR task, i.e., spoken intent classification, and shows good transferability.
Xianghu Yue, Junyi Ao, Xiaoxue Gao, Haizhou Li 0001
ICASSP1
2023 Self-Supervised Acoustic Word Embedding Learning via Correspondence Transformer Encoder
Jingru Lin, Xianghu Yue, Junyi Ao, Haizhou Li 0001
INTERSPEECH2
2021 Phonetically Motivated Self-Supervised Speech Representation Learning
Xianghu Yue, Haizhou Li 0001
Interspeech1
2020 Data augmentation in fault diagnosis based on the Wasserstein generative adversarial network with gradient penalty
Fang Deng, Xianghu Yue
Neurocomputing3
2019 End-to-End Code-Switching ASR for Low-Resourced Language Pairs
abstract
Despite the significant progress in end-to-end (E2E) automatic speech recognition (ASR), E2E ASR for low resourced code-switching (CS) speech has not been well studied. In this work, we describe an E2E ASR pipeline for the recognition of CS speech in which a low-resourced language is mixed with a high resourced language. Low-resourcedness in acoustic data hinders the performance of E2E ASR systems more severely than the conventional ASR systems. To mitigate this problem in the transcription of archives with code-switching Frisian-Dutch speech, we integrate a designated decoding scheme and perform rescoring with neural network-based language models to enable better utilization of the available textual resources. We first incorporate a multi-graph decoding approach which creates parallel search spaces for each monolingual and mixed recognition tasks to maximize the utilization of the textual resources from each language. Further, language model rescoring is performed using a recurrent neural network pre-trained with cross-lingual embedding and further adapted with the limited amount of in-domain CS text. The ASR experiments demonstrate the effectiveness of the described techniques in improving the recognition performance of an E2E CS ASR system in a low-resourced scenario.
Xianghu Yue, Grandee Lee, Emre Yilmaz 0001, Fang Deng, Haizhou Li 0001
ASRU1
2019 Linguistically Motivated Parallel Data Augmentation for Code-Switch Language Modeling
Grandee Lee, Xianghu Yue, Haizhou Li 0001
INTERSPEECH2
2019 Multi-Graph Decoding for Code-Switching ASR
abstract
\n Contains fulltext :\n 214606.pdf (Publisher’s version ) (Open Access)\n
Emre Yilmaz 0001, Samuel Cohen, Xianghu Yue, David A. van Leeuwen, Haizhou Li 0001
INTERSPEECH3
2019 Multidimensional zero-crossing interval points: a low sampling rate acoustic fingerprint recognition method
Xianghu Yue, Fang Deng
Sci. China Inf. Sci.1
2019 Multisource Energy Harvesting System for a Wireless Sensor Network Node in the Field Environment
abstract
This paper presents the design, implementation, and characterization of a hardware platform applicable to a self-powered wireless sensor network (WSN) node. Its primary design objective is to devise a hybrid energy harvesting system to extend the operational lifetime of WSN node after they are deployed in the field environment. Besides the implementation of optimal components (microcontroller, sensor, radio frequency (RF) transceiver, and others) to achieve the lowest power consumption, it is also necessary to consider the sources of energy instead of the frequent recharging or replacement of batteries. Therefore, the platform incorporates a multisource energy harvesting module to collect energy from the surrounding environment, including wind, solar radiation, and thermal energy. The platform also includes an energy storage module through a super-capacitor, RF transceiver module, and the primary microcontroller module. Experimental results showed that the WSN node system with appropriate integration will reserve sufficient energy and meet the long-term power supply requirements of the WSN node without batteries in the field environment. The experimental results and empirical measurements taken over nine days demonstrated that the average daily generating capacity was 7805.09 J, which is far more than the energy consumption of the WSN node (about 2972.88 J).
Fang Deng, Xianghu Yue, Shengpan Guan, Jie Chen 0003
IEEE Internet Things J.2