Xinhui Hu

dblp:04/1555 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2025 Speaker Normalization and Content Restoration for Zero-Shot Voice Conversion with Attention-Enhanced Discriminator
Desheng Hu, Xinhui Hu, Xinkang Xu
INTERSPEECH4
2025 RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Haoqin Sun, Jingguang Tian, Jiaming Zhou 0001, Hui Wang 0075, Jiabei He 0001, Shiwan Zhao, Xiangyu Kong 0001, Desheng Hu, Xinkang Xu, Xinhui Hu
INTERSPEECH10
2025 Discrete Audio Representations for Automated Audio Captioning
Jingguang Tian, Haoqin Sun, Xinhui Hu, Xinkang Xu
INTERSPEECH3
2025 A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
Canan Huang, Desheng Hu, Jingguang Tian, Xinhui Hu
INTERSPEECH5
2024 Learning Emotion-Invariant Speaker Representations for Speaker Verification
abstract
In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To address this issue, we propose multiple improvements to train speaker encoders to increase emotion robustness. Firstly, we utilize CopyPaste-based data augmentation to gather additional parallel data, which includes different emotional expressions from the same speaker. Secondly, we apply cosine similarity loss to restrict parallel sample pairs and minimize intraclass variation of speaker representations to reduce their correlation with emotional information. Finally, we use emotion-aware masking (EM) based on the speech signal energy on the input parallel samples to further strengthen the speaker representation and make it emotion-invariant. We conduct a comprehensive ablation study to demonstrate the effectiveness of these various components. Experimental results show that our proposed method achieves a relative 19.29% drop in EER compared to the baseline system.
Jingguang Tian, Xinhui Hu, Xinkang Xu
ICASSP2
2024 A Deep Representation Learning-Based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder
abstract
Generally, the performance of deep neural networks (DNNs) heavily depends on the quality of data representation learning. Our preliminary work has emphasized the significance of deep representation learning (DRL) in the context of speech enhancement (SE) applications. Specifically, our initial SE algorithm employed a gated recurrent unit variational autoencoder (VAE) with a Gaussian distribution to enhance the performance of certain existing SE systems. Building upon our preliminary framework, this paper introduces a novel approach for SE using deep complex convolutional recurrent networks with a VAE (DCCRN-VAE). DCCRN-VAE assumes that the latent variables of signals follow complex Gaussian distributions that are modeled by DCCRN, as these distributions can better capture the behaviors of complex signals. Additionally, we propose the application of a residual loss in DCCRN-VAE to further improve the quality of the enhanced speech. Compared to our preliminary work, DCCRN-VAE introduces a more sophisticated DCCRN structure and probability distribution for DRL. Furthermore, in comparison to DCCRN, DCCRN-VAE employs a more advanced DRL strategy. The experimental results demonstrate that the proposed SE algorithm outperforms both our preliminary SE framework and the state-of-the-art DCCRN SE method in terms of scale-invariant signal-to-distortion ratio, speech quality, and speech intelligibility.
Jingguang Tian, Xinhui Hu, Xinkang Xu, Zhaohui Yin
ICASSP3
2024 SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
Shuaishuai Ye, Shunfei Chen, Xinhui Hu, Xinkang Xu
INTERSPEECH3
2023 LE-SSL-MOS: Self-Supervised Learning MOS Prediction with Listener Enhancement
abstract
Recently, researchers have shown an increasing interest in automatically predicting the subjective evaluation for speech synthesis systems. This prediction is a challenging task, especially on the out-of-domain test set. In this paper, we proposed a novel fusion model for MOS prediction that combines both supervised and unsupervised approaches. In the supervised aspect, we developed a SSL-based predictor called LE-SSL-MOS. The LE-SSL-MOS utilizes pre-trained self-supervised learning models and further improves prediction accuracy by utilizing the opinion scores of each utterance in the listener enhancement branch. In the unsupervised aspect, two steps are contained: one is that we fine-tuned unit language model (ULM) using highly-intelligible domain data to improve the correlation of an unsupervised metric SpeechLMScore. Another is that we utilized ASR confidence as a new metric with the help of ensemble learning. To the best of our knowledge, this is the first architecture that fuses supervised and unsupervised methods for MOS prediction.With these approaches, our experimental results on the VoiceMOS Challenge 2023 show that LE-SSL-MOS performs better than the baseline. Our fusion system achieved an absolute improvement of 13 % over LE-SSL-MOS on the noisy and enhanced speech track. And our system ranked 1st and 2 nd respectively in the French speech synthesis track and the noisy and enhanced speech track of the challenge.
Zili Qi, Xinhui Hu, Wangjin Zhou, Sheng Li 0010, Xinkang Xu
ASRU2
2023 Using Experience-Based Participatory Approach to Design Interactive Voice User Interfaces for Delivering Physical Activity Programs with Older Adults
abstract
Voice User Interfaces (VUIs) are popular among older adults, who find them easy to use and perceive them as social companions. However, there is a lack of research on voice-based applications to support physical activities for older adults. To address this gap, we present "Workout Pal," a voice agent designed to deliver physical activities to older adults. We conducted a mixed-methods study involving ten older adults to understand their perceptions and design priorities when interacting with Workout Pal. Questionnaires and semi-structured interviews were used to assess their experience, while experience-based co-design sessions facilitated collaboration in exploring the design space and identifying design requirements. Our findings highlight the feasibility and design preferences for smart speaker-based physical activity programs. The study contributes by sharing the design and development of Workout Pal, illustrating a novel co-design approach with older adults, and providing design implications tailored to the specific needs of older adults. The design guidelines emphasize the importance of sociability, voice design, and individualization. This research supports the development of elder-friendly VUIs helping older adults live independently and engage in physical activities.
Smit Desai, Xinhui Hu, Morgan Lundy, Jessie Chin
HAI2
2022 Toward Designing Trustworthy Autonomous Systems: Probing the Role of Humans' Ethical Perspectives
abstract
This study explores whether there is a predictive relationship between humans’ ethical perspectives and their trust in autonomous systems (AS). Whether AS can make ethically acceptable decisions, especially in safety-critical situations involving value trade-offs, has become a significant determinant of how humans will trust these systems. However, knowledge about these relation-based trust dimensions is largely absent from the current theoretical framework of human-AS trust. Addressing this gap, this study used MTurk and Qualtrics to perform an online survey that assessed people’s ethical perspectives, trust in automation, and propensity to trust. The results showed that: (1) Significant differences in trust in automation were seen across four ethical perspectives, confirming the predictive relationship between human ethical perspectives and AS trust. (2) There was no significant difference in propensity to trust among ethical orientations. (3) There was a positive but weak association between trust in automation and willingness to trust. The latter two observations jointly show that trust in automation and trust propensity may be regulated by distinct mechanisms. This study has contributed to existing knowledge by (1) validating the predictive relationship between human ethical perspectives and how they trust AS, (2) revealing potential mechanisms underlying such discrepancies, and (3) highlighting how these differences could help the design of trustworthy AS.
Xinhui Hu, Masooda N. Bashir
CSCWD1
2022 The Royalflush System of Speech Recognition for M2met Challenge
abstract
This paper describes our RoyalFlush system for the track of multi-speaker automatic speech recognition (ASR) in the M2MeT challenge. We adopted the serialized output training (SOT) based multi-speakers ASR system with large-scale simulation data. Firstly, we investigated a set of front-end methods, including multi-channel weighted predicted error (WPE), beamforming, speech separation, speech enhancement, etc., to process training, evaluation, and test sets. However, according to their experimental results, we only selected the WPE and beamforming approach as our front-end methods. Secondly, we made great efforts in the data augmentation for multi-speaker ASR, including adding noise and reverberation, over-lapped speech simulation, multi-channel speech simulation, speed perturbation, front-end processing, etc., which brought us a significant performance improvement. Finally, to make full use of the performance complementary of different model architecture, we trained the standard conformer based joint CTC/Attention (Conformer) and U2++ ASR model with a bidirectional attention decoder, a modification of Conformer, to fuse their results. Compared with the official baseline system, our system got a 12.22% absolute Character Error Rate (CER) reduction on the evaluation set and 12.11% on the test set.
Shuaishuai Ye, Shunfei Chen, Xinhui Hu, Xinkang Xu
ICASSP4
2022 Multiple Enhancements to LSTM for Learning Emotion-Salient Features in Speech Emotion Recognition
Desheng Hu, Xinhui Hu, Xinkang Xu
INTERSPEECH2
2021 An Investigation of Using Hybrid Modeling Units for Improving End-to-End Speech Recognition System
abstract
The acoustic modeling unit is crucial for an end-to-end speech recognition system, especially for the Mandarin language. Until now, most of the studies on Mandarin speech recognition focused on individual units, and few of them paid attention to using a combination of these units. This paper uses a hybrid of the syllable, Chinese character, and subword as the modeling units for the end-to-end speech recognition system based on the CTC/attention multi-task learning. In this approach, the character-subword unit is assigned to train the transformer model in the main task learning stage. In contrast, the syllable unit is assigned to enhance the transformer’s shared encoder in the auxiliary task stage with the Connectionist Temporal Classification (CTC) loss function. The recognition experiments were conducted on AISHELL-1 and an open data set of 1200-hour Mandarin speech corpus collected from the OpenSLR, respectively. The experimental results demonstrated that using the syllable-char-subword hybrid modeling unit can achieve better performances than the conventional units of char-subword, and 6.6% relative CER reduction on our 1200-hour data. The substitution error also achieves a considerable reduction.
Shunfei Chen, Xinhui Hu, Sheng Li 0010, Xinkang Xu
ICASSP2
2021 An End-to-End Dialect Identification System with Transfer Learning from a Multilingual Automatic Speech Recognition Model
Shuaishuai Ye, Xinhui Hu, Sheng Li 0010, Xinkang Xu
Interspeech3
2020 Data Augmentation for Code-Switch Language Modeling by Fusing Multiple Text Generation Methods
Xinhui Hu, Binbin Gu, Xinkang Xu
INTERSPEECH1
2016 Combination of multiple acoustic models with unsupervised adaptation for lecture speech transcription
Xugang Lu, Xinhui Hu, Naoyuki Kanda, Masahiro Saiko, Chiori Hori, Hisashi Kawai
Speech Commun.3
2015 Competitive Strategies for Online Cloud Resource Allocation with Discounts: The 2-Dimensional Parking Permit Problem
abstract
Cloud computing heralded an era where resources can be scaled up and down elastically and in an online manner. This paper initiates the study of cost-effective cloud resource allocation algorithms under price discounts, using a competitive analysis approach. We show that for a single resource, the online resource renting problem can be seen as a 2-dimensional variant of the classic online parking permit problem, and we formally introduce the PPP2problem accordingly. Our main contribution is an online algorithm for PPP2which achieves a deterministic competitive ratio of k (under a certain set of assumptions), where k is the number of resource bundles. This is almost optimal, as we also prove a lower bound of k/3 for any deterministic online algorithm. Our online algorithm makes use of an optimal offline algorithm, which may be of independent interest since it is the first optimal offline algorithm for the 1D and 2D versions of the parking permit problem. Finally, we show that our algorithms and results also generalize to multiple resources (i.e., Multi-dimensional parking permit problems).
Xinhui Hu, Arne Ludwig, Andréa W. Richa, Stefan Schmid 0001
ICDCS1
2014 Translating TED speeches by recurrent neural network based translation model
abstract
This paper presents our recent progress on translating TED speeches1, a collection of public lectures covering a variety of topics. Specially, we use word-to-word alignment to compose translation units of bilingual tuples and present a recurrent neural network-based translation model (RNNTM) to capture long-span context during estimating translation probabilities of bilingual tuples. However, this RNNTM has severe data sparsity problem due to large tuple vocabulary and limited training data. Therefore, a factored RNNTM, which takes bilingual tuples in addition to source and target phrases of the tuples as input features, is proposed to partially address the problem. Our experimental results on the IWSLT2012 test sets show that the proposed models significantly improve the translation quality over state-of-the-art phrase-based translation systems.
Youzheng Wu, Xinhui Hu, Chiori Hori
ICASSP2
2013 Multilingual Speech-to-Speech Translation System: VoiceTra
abstract
This study presents an overview of VoiceTra, which was developed by NICT and released as the world's first network-based multilingual speech-to-speech translation system for smartphones, and describes in detail its multilingual speech recognition, its multilingual translation, and its multilingual speech synthesis in regards to field experiments. We show the effects of system updates using the data collected from field experiments to improve our acoustic and language models.
Shigeki Matsuda, Xinhui Hu, Yoshinori Shiga, Hideki Kashioka, Chiori Hori, Keiji Yasuda, Hideo Okuma, Masao Uchiyama, Eiichiro Sumita, Hisashi Kawai, Satoshi Nakamura 0001
MDM (2)2
2010 Cluster-based language model for spoken document retrieval using NMF-based document clustering
Xinhui Hu, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001
INTERSPEECH1
2010 Construction and evaluations of an annotated Chinese conversational corpus in travel domain for the language model of speech recognition
Xinhui Hu, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001
INTERSPEECH1
2010 Constructing Japanese test collections for spoken term detection
abstract
Spoken Document Retrieval (SDR) and Spoken Term Detection (STD) have been two of the most intensively investigated topics in spoken document processing research according to the establishment of the SDR and STD test collections by the Text REtrieval Conference (TREC) and NIST. Because Japanese spoken document processing researchers also requires such test collections for SDR and STD, we have established a working group to develop these collections in Special Interest Group-Spoken Language Processing (SIG-SLP) of the Information Processing Society of Japan. The working group has constructed and made available a test collection for SDR, and is now constructing new test collections for STD that will be open to researchers. The present paper introduces the policies, outline, and schedule of the new test collections. Then, the new test collections are compared with the NIST STD test collections. Index Terms: spoken term detection, test collection 1.
Yoshiaki Itoh 0001, Hiromitsu Nishizaki, Xinhui Hu, Hiroaki Nanjo, Tomoyosi Akiba, Tatsuya Kawahara, Seiichi Nakagawa, Tomoko Matsui, Yoichi Yamashita, Kiyoaki Aikawa
INTERSPEECH3
2007 Mining redundancy in candidate-bearing snippets to improve web question answering
abstract
Conventional question answering (QA) techniques independently process candidate-bearing snippets to select an exact answer to a question from candidate answers. This paper presents two novel ways of utilizing redundancy in candidate-bearing snippets to help select an exact answer to a question in our Web QA system, i.e., cluster-based language model (CLM-M) and unsupervised SVM classifier (U-SVM) techniques. The comparative experiments demonstrate that the proposed methods significantly outperform the language model-based (LM-M) and supervised SVM-based (S-SVM) techniques that do not utilize this redundancy in the candidate-bearing snippets. Using the CLM-M, the top_1 score is increased from 36.03% (LM-M) to 46.96%; and the top_1 improvement in the U-SVM over the S-SVM is about 23%. Moreover, a cross-model comparison shows that the performance ranking of these models is: U-SVM > CLM-LM > LM-M > S-SVM > R-M (the retrieval-based model).
Youzheng Wu, Xinhui Hu, Hideki Kashioka
CIKM2
2007 Learning Unsupervised SVM Classifier for Answer Selection in Web Question Answering
Youzheng Wu, Ruiqiang Zhang, Xinhui Hu, Hideki Kashioka
EMNLP-CoNLL3
2007 A Priority MAC Protocol for Ad Hoc Networks with Multiple Channels
abstract
Priority scheduling has been widely used in mobile ad hoc networks. However, most of the prior work related to priority scheduling is designed only for a single data channel. In addition, without providing certain mechanisms to incorporate priority scheduling, existing multi-channel MAC protocols can not provide differentiated service. In this paper, we propose a multi-channel priority MAC protocol which can significantly increase the throughput of high priority flows and reduce their average delay. Our protocol consists of four mechanisms: control phase contention mechanism, priority- oriented channel assignment mechanism, data phase contention mechanism and flow interpolation mechanism. Simulation results demonstrate the effectiveness of the proposed protocol.
Xinhui Hu, Jianhua Zhang 0001, Ping Zhang 0003
PIMRC2
2006 Automatic Derivation of a Phoneme Set with Tone Information for Chinese Speech Recognition Based on Mutual Information Criterion
abstract
An appropriate approach to model tone information is helpful for building Chinese large vocabulary continuous speech recognition system. We propose to derive an efficient phoneme set of tone-dependent sub-word units to build a recognition system, by iteratively merging a pair of tone-dependent units according to the principle of minimal loss of the mutual information. The mutual information is measured between the word tokens and their phoneme transcriptions in a training text corpus, based on the system lexical and language model. The approach has the capability to keep discriminative tonal (and phoneme) contrasts that are most helpful for disambiguating homophone words due to lack of tones, and merge those tonal (and phoneme) contrasts that are not important for word disambiguation for the recognition task. This enable a flexible selection of phoneme set according to a balance between the MI information amount and the number of phonemes. We applied the method to traditional phoneme set of Initial/Finals, and derived several phoneme sets with different number of units. Speech recognition experiments using the derived sets showed their effectiveness.
Jinsong Zhang 0001, Xinhui Hu, Satoshi Nakamura 0001
ICASSP (1)2
2006 Language modeling of Chinese personal names based on character units for continuous Chinese speech recognition
Xinhui Hu, Hirofumi Yamamoto, Gen-ichiro Kikui, Yoshinori Sagisaka
INTERSPEECH1
1995 HMM-based tone recognition of Chinese trisyllables using double codebooks on fundamental frequency and waveform power
Keikichi Hirose, Xinhui Hu
EUROSPEECH2
1994 Recognition of Chinese tones in monosyllabic and disyllabic speech using HMM
Xinhui Hu, Keikichi Hirose
ICSLP1