EDBT 2026 Demo / reviewers in the wild / expert
Xihong Wu
dblp:49/6048
· DBLP profile ↗
87ranked-venue papers
3as first author
26since 2021 · last 2026
0009-0004-5236-7469ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 51 · 3 first-author · 15 since 2021Systems, architecture and hardware · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrajDiff: Diffusion-Based Trajectory Generation with 3D ESDF Scene Understanding
Xinmiao Du, Zhuoyu Jiang, Xihong Wu |
ICIC (28) | 3 |
| 2025 | Cross-attention Inspired Selective State Space Models for Target Sound ExtractionabstractThe Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state space models, notably the latest work Mamba, have shown comparable performance to Transformer-based methods while significantly reducing computational complexity in various tasks. However, Mamba’s applicability in target sound extraction is limited due to its inability to capture dependencies between different sequences as the cross-attention does. In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture. The calculation of Mamba can be divided to the query, key and value. We utilize the clue to generate the query and the audio mixture to derive the key and value, adhering to the principle of the cross-attention mechanism in Transformers. Experimental results from two representative target sound extraction methods validate the efficacy of the proposed CrossMamba. Donghang Wu, Yiwen Wang 0009, Xihong Wu, Tianshu Qu |
ICASSP | 3 |
| 2025 | TA-V2A: Textually Assisted Video-to-Audio GenerationabstractAs artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While Transformer and Diffusion models have advanced audio generation, a significant challenge persists in extracting precise semantic information from videos, as current models often lose sequential context by relying solely on frame-based features. To address this, we present TA-V2A, a method that integrates language, audio, and video features to improve semantic representation in latent space. By incorporating large language models for enhanced video comprehension, our approach leverages text guidance to enrich semantic expression. Our diffusion model-based system utilizes automated text modulation to enhance inference quality and efficiency, providing personalized control through text-guided interfaces. This integration enhances semantic expression while ensuring temporal alignment, leading to more accurate and coherent video-to-audio generation. Yuhuan You, Xihong Wu, Tianshu Qu |
ICASSP | 2 |
| 2025 | A Novel Multimodal Method for Decoding Speech Perception from Brain ActivitiesabstractDecoding speech from neural recordings has critical importance in application and scientific research. However, this task is still challenging with non-invasive recordings. Previous research has shown significant improvement in speech perception decoding task by leveraging wav2vec vectors and gives the potential for applications. To further explore this problem, we proposed a novel multimodal method by using functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG). In our method, separate encoders for fMRI and MEG are considered, then features extracted from both modalities are integrated and aligned with wav2vec vectors that were extracted from the speech. The multimodal method reaches averaged performance of 72.6% in top-10 accuracy with a negative sample size of 128. Performance evaluated with various metrics achieves steady improvement across subjects, demonstrating the effectiveness of the proposed data fusion method. Interpretation of the performance increment was also investigated by testing the correlation between encoder hidden outputs and different level of features extracted from the speech. Results demonstrate that MEG encoder learns more low-level information and fMRI encoder learns more high-level information, which indicates both complementary characteristics lead to the improvement. The result of this work shows the potential of multimodal methods for speech decoding. Aoke Zhang, Bo Wang 0110, Xihong Wu, Jing Chen 0019 |
ICASSP | 3 |
| 2025 | Using Ear-EEG to Decode Auditory Attention in Multiple-speaker EnvironmentabstractAuditory Attention Decoding (AAD) can help to determine the identity of the attended speaker during an auditory selective attention task, by analyzing and processing measurements of electroencephalography (EEG) data. Most studies on AAD are based on scalp-EEG signals in two-speaker scenarios, which are far from real application. Ear-EEG has recently gained significant attention due to its motion tolerance and invisibility during data acquisition, making it easy to incorporate with other devices for applications. In this work, participants selectively attended to one of the four spatially separated speakers’ speech in an anechoic room. The EEG data were concurrently collected from a scalp-EEG system and an ear-EEG system (cEEGrids). Temporal response functions (TRFs) and stimulus reconstruction (SR) were utilized using ear-EEG data. Results showed that the attended speech TRFs were stronger than each unattended speech and decoding accuracy was 41.3% in the 60s (chance level of 25%). To further investigate the impact of electrode placement and quantity, SR was utilized in both scalp-EEG and ear-EEG, revealing that while the number of electrodes had a minor effect, their positioning had a significant influence on the decoding accuracy. One kind of auditory spatial attention detection (ASAD) method, STAnet, was testified with this ear-EEG database, resulting in 93.1% in 1-second decoding window. The implementation code and database for our work are available on GitHub: https://github.com/zhl486/Ear_EEG_code.git and Zenodo: https://zenodo.org/records/10803261. Haolin Zhu, Yujie Yan, Xiran Xu, Zhongshu Ge, Pei Tian, Xihong Wu, Jing Chen 0019 |
ICASSP | 6 |
| 2025 | Position also matters! Separating Same Instruments in String Quartet using Timbral and Positional Cues
Yuetonghui Xu, Yiwen Wang 0009, Xihong Wu |
INTERSPEECH | 3 |
| 2025 | Overestimated performance of auditory attention decoding caused by experimental design in EEG recordings
Yujie Yan, Xiran Xu, Haolin Zhu, Songyi Li, Bo Wang 0110, Xihong Wu, Jing Chen 0019 |
INTERSPEECH | 6 |
| 2025 | DPGP: A Hybrid 2D-3D Dual Path Potential Ghost Probe Zone Prediction Framework for Safe Autonomous DrivingabstractModern robots must coexist with humans in dense urban environments. A key challenge is the ghost probe problem, where pedestrians or objects unexpectedly rush into traffic paths. This issue affects both autonomous vehicles and human drivers. Existing works propose vehicle-to-everything (V2X) strategies and non-line-of-sight (NLOS) imaging for ghost probe zone detection. However, most require high computational power or specialized hardware, limiting real-world feasibility. Additionally, many methods do not explicitly address this issue. To tackle this, we propose DPGP, a hybrid 2D-3D fusion framework for ghost probe zone prediction using only a monocular camera during training and inference. With unsupervised depth prediction, we observe ghost probe zones align with depth discontinuities, but different depth representations offer varying robustness. To exploit this, we fuse multiple feature embeddings to improve prediction. To validate our approach, we created a 12K-image dataset annotated with ghost probe zones, carefully sourced and cross-checked for accuracy. Experimental results show our framework outperforms existing methods while remaining cost-effective. To our knowledge, this is the first work extending ghost probe zone prediction beyond vehicles, addressing diverse non-vehicle objects. We will open-source our code and dataset for community benefit. Weiming Qu, Shenghai Yuan 0001, Shengyi Liu, Yuanhao Zhu, Jiayi Rao, Xihong Wu, Dingsheng Luo |
IROS | 13 |
| 2025 | Online Iterative Learning with Forward Simulation for Sub-minimum End-effector Displacement PositioningabstractPrecision is a crucial performance indicator for robot arms. During interacting with human, high precision enables a robot arm to be used effectively and safely, while low precision may lead to safety issues. Traditional methods for improving robot arm precision rely on error compensation. However, these methods are often not robust and lack adaptability. Learning-based methods offer greater flexibility and adaptability, while current researches show that they often fall short in achieving high precision and struggle to handle many scenarios requiring high precision. In this paper, we propose a novel high-precision robot arm manipulation framework based on online iterative learning and forward simulation, which can achieve positioning error (precision) less than end-effector physical minimum displacement. In other words, our proposed method can compensate for the precision-limitation of the hardware structure of the robot arms. Furthermore, we consider the joint angular resolution of the real robot arm, which is usually neglected in related works. A series of experiments on both simulation and real UR3 robot arm platforms demonstrate that our proposed method is effective and promising. The related code will be available soon. Weiming Qu, Tianlin Liu, Xihong Wu, Dingsheng Luo |
IROS | 4 |
| 2025 | SILM: A Subjective Intent Based Low-Latency Framework for Multiple Traffic Participants Joint Trajectory PredictionabstractTrajectory prediction is a fundamental technology for advanced autonomous driving systems and represents one of the most challenging problems in the field of cognitive intelligence. Accurately predicting the future trajectories of each traffic participant is a prerequisite for building high safety and high reliability decision-making, planning, and control capabilities in autonomous driving. However, existing methods often focus solely on the motion of other traffic participants without considering the underlying intent behind that motion, which increases the uncertainty in trajectory prediction. Autonomous vehicles operate in real-time environments, meaning that trajectory prediction algorithms must be able to process data and generate predictions in real-time. While many existing methods achieve high accuracy, they often struggle to effectively handle heterogeneous traffic scenarios. In this paper, we propose a Subjective Intent-based Low-latency framework for Multiple traffic participants joint trajectory prediction. Our method explicitly incorporates the subjective intent of traffic participants based on their key points, and predicts the future trajectories jointly without map, which ensures promising performance while significantly reducing the prediction latency. Additionally, we introduce a novel dataset designed specifically for trajectory prediction. Related code and dataset will be available soon. Weiming Qu, Yuanhao Zhu, Xihong Wu, Dingsheng Luo |
IROS | 8 |
| 2024 | Semantic Reconstruction of Continuous Language from Meg SignalsabstractDecoding language from neural signals holds considerable theoretical and practical importance. Previous research has indicated the feasibility of decoding text or speech from invasive neural signals. However, when using non-invasive neural signals, significant challenges are encountered due to their low quality. In this study, we proposed a data-driven approach for decoding semantic of language from Magnetoencephalography (MEG) signals recorded while subjects were listening to continuous speech. First, a multi-subject decoding model was trained using contrastive learning to reconstruct continuous word embeddings from MEG data. Subsequently, a beam search algorithm was adopted to generate text sequences based on the reconstructed word embeddings. Given a candidate sentence in the beam, a language model was used to predict the subsequent words. The word embeddings of the subsequent words were correlated with the reconstructed word embedding. These correlations were then used as a measure of the probability for the next word. The results showed that the proposed continuous word embedding model can effectively leverage both subject-specific and subject-shared information. Additionally, the decoded text exhibited significant similarity to the target text, with an average BERTScore of 0.816. Bo Wang 0110, Xiran Xu, Longxiang Zhang, Boda Xiao, Xihong Wu, Jing Chen 0019 |
ICASSP | 5 |
| 2024 | A Hybrid Deep-Online Learning Based Method for Active Noise Control in Wave DomainabstractThe traditional feedback Active Noise Control (ANC) algorithms are built upon linear filters, which leads to reduced performance when dealing with real-world noise. Deep learning-based feedback ANC algorithms have been proposed to overcome this problem. However, methods relying on pre-trained neural networks exhibit performance degradation when encountering noise from unseen scenes in the training dataset. This paper proposed a hybrid deep-online learning based spatial ANC system which combines online learning with pre-trained deep neural networks. The proposed method can keep the performance on noise from the trained scenes while improve the performance of cancelling noise from new scenes. Additionally, by incorporating wave domain decomposition, this paper achieves noise cancellation over a control spatial region. Simulation experiments validate the effectiveness of the combination of online learning and deep learning in handling previously unseen noise. Furthermore, the efficiency of wave domain decomposition in spatial noise cancellation is also verified. Donghang Wu, Xihong Wu, Tianshu Qu |
ICASSP | 2 |
| 2024 | A DenseNet-Based Method for Decoding Auditory Spatial Attention with EEGabstractAuditory spatial attention detection (ASAD) aims to decode the attended spatial location with EEG in a multiple-speaker setting. ASAD methods are inspired by the brain lateralization of cortical neural responses during the processing of auditory spatial attention, and show promising performance for the task of auditory attention decoding (AAD) with neural recordings. In the previous ASAD methods, the spatial distribution of EEG electrodes is not fully exploited, which may limit the performance of these methods. In the present work, by transforming the original EEG channels into a two-dimensional (2D) spatial topological map, the EEG data is transformed into a three-dimensional (3D) arrangement containing spatial-temporal information. And then a 3D deep convolutional neural network (DenseNet-3D) is used to extract temporal and spatial features of the neural representation for the attended locations. The results show that the proposed method achieves higher decoding accuracy than the state-of-the-art (SOTA) method (94.3% compared to XANet's 90.6%) with 1-second decision window for the widely used KULeuven (KUL) dataset, and the code to implement our work is available on Github: https://github.com/xuxiran/ASAD_DenseNet Xiran Xu, Bo Wang 0110, Yujie Yan, Xihong Wu, Jing Chen 0019 |
ICASSP | 4 |
| 2024 | TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
Yiwen Wang 0009, Xihong Wu |
INTERSPEECH | 2 |
| 2024 | Auditory Attention Decoding in Four-Talker Environment with EEG
Yujie Yan, Xiran Xu, Haolin Zhu, Pei Tian, Zhongshu Ge, Xihong Wu, Jing Chen 0019 |
INTERSPEECH | 6 |
| 2024 | Employing feature mixture for active learning of object detection
Licheng Zhang 0004, Siew-Kei Lam, Dingsheng Luo, Xihong Wu |
Neurocomputing | 4 |
| 2023 | PGSS: Pitch-Guided Speech SeparationabstractMonaural speech separation aims to separate concurrent speakers from a single-microphone mixture recording. Inspired by the effect of pitch priming in auditory scene analysis (ASA) mechanisms, a novel pitch-guided speech separation framework is proposed in this work. The prominent advantage of this framework is that both the permutation problem and the unknown speaker number problem existing in general models can be avoided by using pitch contours as the primary means to guide the target speaker. In addition, adversarial training is applied, instead of a traditional time-frequency mask, to improve the perceptual quality of separated speech. Specifically, the proposed framework can be divided into two phases: pitch extraction and speech separation. The former aims to extract pitch contour candidates for each speaker from the mixture, modeling the bottom-up process in ASA mechanisms. Any pitch contour can be selected as the condition in the second phase to separate the corresponding speaker, where a conditional generative adversarial network (CGAN) is applied. The second phase models the effect of pitch priming in ASA. Experiments on the WSJ0-2mix corpus reveal that the proposed approaches can achieve higher pitch extraction accuracy and better separation performance, compared to the baseline models, and have the potential to be applied to SOTA architectures. Xiang Li 0072, Yiwen Wang 0009, Xihong Wu, Jing Chen 0019 |
AAAI | 4 |
| 2023 | A Model-Based Hearing Compensation Method Using a Self-Supervised FrameworkabstractHearing aids can improve auditory perception for hearing-impaired (HI) listeners, but even state-of-art devices provide only limited benefits if not configured correctly for the listeners. The prescriptive fittings of hearing aids ignore the individual difference among HI listeners with identical hearing thresholds. This paper proposes a model-based hearing compensation method using a self-supervised framework with a given auditory model. The influence of outer/inner hair cells dysfunction was simulated in the auditory model. And then, a neural network was trained to compensate for the given hearing impairment. Both objective and subjective experiments were conducted to evaluate the present method, and the results showed that listeners are sensitive to the parameter controlling the contribution of outer hair cells dysfunction. Additionally, the result indicated that listeners significantly preferred the speech processed by the proposed method to the traditional perspective fitting. Yadong Niu, Xihong Wu, Jing Chen 0019 |
ICASSP | 3 |
| 2023 | TT-Net: Dual-Path Transformer Based Sound Field Translation in the Spherical Harmonic DomainabstractIn the current method for the sound field translation tasks based on spherical harmonic (SH) analysis, the solution based on the additive theorem usually faces the problem of singular values caused by large matrix condition numbers. The influence of different distances and frequencies of the spherical radial function on the stability of the translation matrix will affect the accuracy of the SH coefficients at the selected point. Due to the problems mentioned above, we propose a neural network scheme based on the dual-path transformer. More specifically, the dual-path network is constructed by the selfattention module along the two dimensions of the frequency and order axes. The transform-average-concatenate layer and upscaling layer are introduced in the network, which provides solutions for multiple sampling points and upscaling. Numerical simulation results indicate that both the working frequency range and the distance range of the translation are extended. More accurate higher-order SH coefficients are obtained with the proposed dual-path network. Yiwen Wang 0009, Zijian Lan, Xihong Wu, Tianshu Qu |
ICASSP | 3 |
| 2023 | Emotion Classification with EEG Responses Evoked by Emotional Prosody of Speech
Xihong Wu |
INTERSPEECH | 2 |
| 2023 | A Physical Model-Based Self-Supervised Learning Method for Signal Enhancement Under Reverberant EnvironmentabstractIn a reverberant environment, interferences such as reflections and background noise can degrade the perception of the sound source signal. Although the DNN-based methods have made a tremendous breakthrough in addressing this issue, the performance of these models is highly dependent on the completeness of the training dataset, which will limit its generalization under unknown environments. In this paper, we propose a physical model-based self-supervised learning (PMSSL) method to realize the DNN model optimization under unknown scenarios. This method incorporates a room reverberation physical model into the sound source enhancement model optimization process, realizing the self-learning of the DNN model under physical constraints. In this process, the time-frequency characteristics of the input signal and the spatial feature of the reverberation environment are utilized for parameter optimization, improving the adaptability of the DNN model under unknown scenarios. Experimental results based on simulated and measured data prove that the proposed method can obtain much more accurate source signal enhancement results compared with the pre-trained models, verifying its effectiveness and adaptability in new environments. Xihong Wu, Tianshu Qu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Multi-Speaker Pitch Tracking via Embodied Self-Supervised LearningabstractPitch is a critical cue in human speech perception. Although the technology of tracking pitch in single-talker speech succeeds in many applications, it’s still a challenging problem to extract pitch information from mixtures. Inspired by the motor theory of speech perception, a novel multi-speaker pitch tracking approach is proposed in this work, based on an embodied self-supervised learning method (EMSSL-Pitch). The conceptual idea is that speech is produced through an underlying physical process (i.e., human vocal tract) given the articulatory parameters (articulatory-to-acoustic), while speech perception is like the inverse process, aiming at perceiving the intended articulatory gestures of the speaker from acoustic signals (acoustic-to-articulatory). Pitch value is part of the articulatory parameters, corresponding to the vibration frequency of vocal folders. The acoustic-to-articulatory inversion is modeled in a self-supervised manner to learn an inference network by iteratively sampling and training. The learned representations from this inference network can have explicit physical meanings, i.e., articulatory parameters where pitch information can be further extracted. Experiments on GRID database show that EMSSL-Pitch can achieve a reachable performance compared with supervised baselines and be generalized to unseen speakers. Xihong Wu |
ICASSP | 3 |
| 2022 | Advanced Face Anti-Spoofing with Depth SegmentationabstractFace anti-spoofing (FAS) plays a vital role in securing face recognition systems. In state-of-the-art FAS methods, face depth is determined for every position in a facial image. However, face depth varies at different positions, which leads to low accuracy when predicting face depth. As we observe, spoof faces have depth values that are 0, while live faces have depths that are equal to or greater than 0. As a result, for faces that have a depth greater than 0, if they are estimated as merely a positive value, instead of an accurate value, the prediction of whether they are real or not will not change. Further, if a range of depths are considered as one category, then there are more samples per category for the network training. Based on the above observation, in this paper, we propose to aggregate simple depth values to the same category and perform classification to optimize the FAS network. To evaluate the performance of the proposed approach, we perform extensive experiments on four benchmark databases, respectively, OULU-NPU, SiW, CASIA-FASD, and Replay-Attack. The results demonstrate that the proposed approach outperforms state-of-the-art methods on intra-database testing. Furthermore, our proposed approach shows advanced performance on cross-database testing. Nan Sun 0002, Xihong Wu, Dingsheng Luo |
IJCNN | 3 |
| 2022 | Unsupervised Acoustic-to-Articulatory Inversion with Variable Vocal Tract Anatomy
Qinlong Huang, Xihong Wu |
INTERSPEECH | 3 |
| 2022 | Unsupervised Inference of Physiologically Meaningful Articulatory Trajectories with VocalTractLab
Qinlong Huang, Xihong Wu |
INTERSPEECH | 3 |
| 2022 | Sparse DNN Model for Frequency Expanding of Higher Order Ambisonics Encoding ProcessabstractThe performance of higherorder Ambisonics (HOA) signals obtained using spherical harmonics decomposition method is disturbed by two primary sources of errors, the noise pollution in low-frequency band and the spatial aliasing in high-frequency band. Inspired by the HOA signals upscale method, which is performed using the sparse character of the sound field, this paper propose a sound field decomposition model based on a sparse deep neural network that offers HOA signals with wider frequency bandwidth. We use the frequency domain multi-scale convolutional network to realize the spherical harmonics decomposition, as well as learning the spatial aliasing pattern, based on which the aliasing-free HOA signals can be derived. Besides, we apply a sparse encoding network to cpature the sparse feature of the sound field which will improve the model performance when the sparse condition is satisfied. The experiments results prove that the proposed model can obtain HOA signals with wider frequency range of operation under multiple sources (up to 10 sources) and low reverberant environments ($T_{60}\le$400 ms). When the sparsity feature cannot be satisfied ($T_{60} =$800 ms), the proposed network model still maintain the same performance as the traditional methods. Xihong Wu, Tianshu Qu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Single-Channel Speech Separation Integrating Pitch Information Based on a Multi Task Learning FrameworkabstractPitch is a critical cue for speech separation in humans' auditory perception. Although the technology of tracking pitch in single-talker speech succeeds in many applications, it's still a challenging problem to extract pitch information from speech mixtures in machine perception. In this paper, we aimed to combine speech separation and pitch tracking together to let them benefit from each other. A multi-task learning framework was proposed, in which a unified objective that considered both speech separation and pitch tracking was used, based on the utterance-level permutation invariant training (uPIT) as well as deep clustering (DPCL). In such framework, two tasks were optimized simultaneously and could benefit from each other through the sharing layers in the networks. Experimental results indicated the proposed multi-task framework outperformed the corresponding single-task framework, in terms of both speech separation and pitch tracking. The improvement was more significant for challenging same-gender mixtures. Xiang Li 0072, Xihong Wu, Jing Chen 0019 |
ICASSP | 4 |
| 2020 | Individual Distance-Dependent HRTFS Modeling Through A Few Anthropometric MeasurementsabstractThe lack of data is a major problem in individual HRTF modeling. There are many HRTF databases, but each database only has limited HRTFs with different characteristics, such as distance-dependent HRTFs or individual HRTFs. How to effectively model HRTFs through several different databases is an important task. In this paper, a method for predicting individual distance-dependent HRTFs using a few anthropometric parameters is proposed. By modeling the HRTFs in CIPIC database, which contains individual HRTFs in 1 meter, and the PKU&IOA database, which contains KEMAR HRTFs in eight distances, we predict the individual HRTFs in arbitrary directions and distances. The objective experiments show that the proposed model has less spectral distortions than distance variation function model. The subjective experiments show that the proposed model can predict the individual HRTFs in arbitrary directions and distances. Mengfan Zhang, Xihong Wu, Tianshu Qu |
ICASSP | 2 |
| 2020 | Competing Speaker Count Estimation on the Fusion of the Spectral and Spatial Embedding Space
Xihong Wu, Tianshu Qu |
INTERSPEECH | 2 |
| 2020 | Modeling of Individual HRTFs Based on Spatial Principal Component AnalysisabstractHead-related transfer function (HRTF) plays an important role in the construction of 3D auditory display. This article presents an individual HRTF modeling method using deep neural networks based on spatial principal component analysis. The HRTFs are represented by a small set of spatial principal components combined with frequency and individual-dependent weights. By estimating the spatial principal components using deep neural networks and mapping the corresponding weights to a quantity of anthropometric parameters, we predict individual HRTFs in arbitrary spatial directions. The objective and subjective experiments evaluate the HRTFs generated by the proposed method, the principal component analysis (PCA) method, and the generic method. The results show that the HRTFs generated by the proposed method and PCA method perform better than the generic method. For most frequencies the spectral distortion of the proposed method is significantly smaller than the PCA method in the high frequencies but significantly larger in the low frequencies. The evaluation of the localization model shows the PCA method is better than the proposed method. The subjective localization experiments show that the PCA and the proposed methods have similar performances in most conditions. Both the objective and subjective experiments show that the proposed method can predict HRTFs in arbitrary spatial directions. Mengfan Zhang, Zhongshu Ge, Tiejun Liu, Xihong Wu, Tianshu Qu |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Improvements to the Matching Projection Decoding Method for Ambisonic System with Irregular Loudspeaker LayoutsabstractThe Ambisonic technique has been widely used for sound field recording and reproduction recently. However, the basic Ambisonic decoding method will break down when the playback loudspeakers distribute unevenly. Various methods have been proposed to solve this problem. This paper introduces several improvements to a recently proposed Ambisonic decoding method, the matching projection method, for uneven loudspeaker layouts. The first improvement is energy preserving; the second is introducing the "in-phase" weight, and the third is introducing partial projection coefficients. To evaluate the improved method, we compared it with the original one and the all-round Ambisonic decoding method with a 2-dimension unevenly arranged loudspeaker array. The result shows our method greatly improves the original method where the loudspeaker arranges very sparsely or densely. Zhongshu Ge, Xihong Wu, Tianshu Qu |
ICASSP | 2 |
| 2019 | A Spectral-change-aware Loss Function for DNN-based Speech SeparationabstractSpeech separation can be treated as a mask estimation problem where supervised learning is employed to construct the mapping from acoustic features to a mask. Interference can be reduced by applying the estimated mask on a time-frequency (T-F) representation of noisy speech, resulting in improved speech intelligibility. Most of existing learning networks for speech separation aim to minimize the Mean Square Error (MSE) over the training set, where the loss from each T-F representation is equally weighted. In this paper, we proposed a spectral-change-aware loss function, where loss from the T-F units with large spectral changes over time were assigned higher weights compared to the T-F units with minor spectral changes. Such spectral-change-aware loss function was evaluated on speech separation performance in terms of mask estimation accuracy, short-time objective intelligibility (STOI) and SNR gain of unvoiced segments. The results indicated that the proposed loss function could further improve the speech intelligibility and increase SNR gain of unvoiced segments even in the cost of increased error rate of estimated mask. Xiang Li 0072, Xihong Wu, Jing Chen 0019 |
ICASSP | 2 |
| 2019 | Integrating Spectrotemporal Context into Features Based on Auditory Perception for Classification-based Speech SeparationabstractSpeech separation, which has been a challenging task for decades, especially at low signal-to-noise ratios (SNRs), can be cast as a classification problem. In such adverse acoustic environment, extracting robust features from noisy mixtures is crucial for successful classification. In the past studies, features representing temporal dynamics, known as delta features, have been widely used. Combining basic features with their deltas yields better speech separation results than using basic features alone. In this study, the commonly used delta feature was modified according to the characteristics of auditory perception, which included auditory processing on spectral change and spectral contrast. Therefore, we proposed a feature which integrated spectrotemporal context via replacing the commonly used delta feature by spectral change feature and spectral contrast feature. Experimental results showed that the proposed feature could produce better speech segregation performance than the common delta feature. Xiang Li 0072, Xihong Wu, Jing Chen 0019 |
ICASSP | 2 |
| 2019 | Distance-dependent Modeling of Head-related Transfer FunctionsabstractIn this paper, a method for modeling distance dependent head-related transfer functions is presented. The HRTFs are first decomposed by spatial principal component analysis. Using deep neural networks, we model the spatial principal component weights of different distances. Then we realize the prediction of HRTFs in arbitrary spatial distances. The objective and subjective experiments are conducted to evaluate the proposed distance model and the distance variation function model, and the results have shown that the proposed model has less spectral distortions than distance variation function model, and the virtual sound generated by the proposed model has better performance in terms of distance localization. Mengfan Zhang, Xihong Wu, Tianshu Qu |
ICASSP | 3 |
| 2019 | A Hierarchical Model for StarCraft II Mini-GameabstractStarCraft II is one of the most challenging real-time strategy games, due to huge action space, large observation space, imperfect information, etc. Therefore, it is hard to learn the full game of StarCraft II. To reduce the learning complexity, DeepMind and Blizzard released several mini-games, in which the BuildMarines mini-game is most challenging, due to long time horizons, partially-observed state, high-dimensional, continuous action space and observation space. In this paper, we propose a hierarchical modeling method to solve those challenges in BuildMarines mini-game. Our approach consists of two levels, combining learning-based (high-level) and rule-based (low-level) method. The learning-based method leverages DQN reinforcement learning algorithm, while the rule-based method leverages script to realize. Experimental results show that the proposed approach is effective for an agent to learn the long planning horizon game, BuildMarines. Tianlin Liu, Xihong Wu, Dingsheng Luo |
ICMLA | 2 |
| 2019 | Effects of Spectral and Temporal Cues to Mandarin Concurrent-Vowels Identification for Normal-Hearing and Hearing-Impaired Listeners
Zhen Fu, Xihong Wu, Jing Chen 0019 |
INTERSPEECH | 2 |
| 2018 | Matching Projection Decoding Method for Ambisonics SystemabstractThe basic Ambisonics decoding method will break down when the playback loudspeakers distribute unevenly. This paper proposes a modified Ambisonics method, the matching projection decoding method, for solving this problem. The matching projection decoding method is a kind of the greedy algorithm. It firstly calculates the projection value of the object Ambisonics signal over each Ambisonics signal of loudspeakers, then the maximum projection value is assigned to the corresponding loudspeaker. This process is repeated until all the loudspeakers have been assigned a gain value. The objective and subjective experiments were performed to evaluate the proposed system and the basic system. Objective evaluation results show that the accuracy of the sound field generated by the matching projection decoding method is better than that of the basic method; and the subjective evaluation results show a more correct directional perception of the matching projection decoding method than the basic one. Tianshu Qu, Xihong Wu |
ICASSP | 4 |
| 2018 | A Time-Weighted Method for Predicting the Intelligibility of Speech in the Presence of Interfering SoundsabstractThe speech intelligibility index (SII) has been widely used as an objective method of predicting speech intelligibility, but its traditional form is most effective predicting speech intelligibility scores under stationary noise but not more challenging conditions (e.g., competing noise interference). To address this limitation, the present work extended the SII model to predict the intelligibility of speech in both steady speech-spectral noise (SSN) and dual-talker speech (DTS), by using a time-weighted function that accounted for the relative perceptual importance of vowels and consonants in speech intelligibility. The performance of the new time-weighted SII (TW-SII) was compared to the other two well-known methods, i.e., the time-averaged SII (TA-SII) and coherence SII (CSII). Experimental results showed the intelligibility prediction accuracy of the three methods was similar for speech in SSN, but the prediction by TW-SII was more accurate than those by TA-SII and CSII for speech in DTS. The possible applications and limitations of the present intelligibility model were analyzed and discussed. Mingjie Song, Fei Chen 0011, Xihong Wu, Jing Chen 0019 |
ICASSP | 3 |
| 2018 | Measuring the Band Importance Function for Mandarin Chinese with a Bayesian Adaptive Procedure
Yufan Du, Yi Shen 0008, Hongying Yang, Xihong Wu, Jing Chen 0019 |
INTERSPEECH | 4 |
| 2017 | Multi-scale feature based convolutional neural networks for large vocabulary speech recognitionabstractDeep learning has brought a breakthrough to the performance of speech recognition. The speech recognition systems based on deep neural networks have obtained the state-of-the-art performance on various speech recognition tasks. These systems almost utilize the Mel-frequency cepstral coefficients or the Mel-scale log-filterbank coefficients, which are based on short-time Fourier transform. Although these features are designed based on the auditory characteristics of the human, it is a problem that the inherent tradeoff of the temporal and frequency resolution still exists in spectral representations based on short-time Fourier transform. In this paper, we propose a multi-scale method to mitigate the tradeoff and a model architecture that enables to analyze speech at multiple scale. Experiments are conducted on TIMIT and HKUST corpus. We compare the proposed multi-scale features and traditional features at various number of configurations. Experimental results show that the proposed model architecture can obtain significant performance improvement. Tong Fu, Xihong Wu |
ICME | 2 |
| 2017 | A hierarchical inverse model based on proprioception and DNN for robot reachingabstractRobot reaching ability serves as one of essential basis for many other manipulation skills, such as grasping, placing etc., and has been widely focused for decades. Inverse model, which plays a fundamental role within robot reaching ability, aims to produce motor commands to drive the system to the desired state. However, since the inverse model is always an one-to-many mapping, it suffers from the multi-solution issue and the adaptation problem for changing circumstances when traditional kinematic/dynamic models are employed. And thus, learning based approaches are investigated, including the deep learning based models. In this paper, to further improve the performance of deep neural networks (DNN) methods, a novel model is proposed, where both the proprioception and a hierarchical structure are involved. Here, the employed concept of proprioception is aimed to follow human mechanism, while the hierarchical structure is expected to mimic the fact that different joints usually play different effects in a manipulation process. Experiments are performed with respect to PKU-HR6.0 II humanoid robot, and the results illustrate the effectiveness and superiority of the proposed model. Tao Zhang 0071, Yian Deng, Jun-Hai Zhai, Xihong Wu, Dingsheng Luo |
IECON | 5 |
| 2017 | Towards human-like and transhuman perception in AI 2.0: a reviewabstractPerception is the interaction interface between an intelligent system and the real world. Without sophisticated and flexible perceptual capabilities, it is impossible to create advanced artificial intelligence (AI) systems. For the next-generation AI, called ‘AI 2.0’, one of the most significant features will be that AI is empowered with intelligent perceptual capabilities, which can simulate human brain’s mechanisms and are likely to surpass human brain in terms of performance. In this paper, we briefly review the state-of-the-art advances across different areas of perception, including visual perception, auditory perception, speech perception, and perceptual information processing and learning engines. On this basis, we envision several R&D trends in intelligent perception for the forthcoming era of AI 2.0, including: (1) human-like and transhuman active vision; (2) auditory perception and computation in an actual auditory setting; (3) speech perception and computation in a natural interaction setting; (4) autonomous learning of perceptual information; (5) large-scale perceptual information processing and learning platforms; and (6) urban omnidirectional intelligent perception and reasoning engines. We believe these research directions should be highlighted in the future plans for AI 2.0. Yonghong Tian 0001, Xilin Chen 0001, Hongkai Xiong, Li-Rong Dai 0001, Jing Chen 0002, Junliang Xing, Jing Chen 0003, Xihong Wu, Weiming Hu 0004, Yu Hu 0003, Tiejun Huang 0001, Wen Gao 0001 |
Frontiers Inf. Technol. Electron. Eng. | 9 |
| 2016 | Biped robot falling motion control with human-inspired active complianceabstractProtecting robot from broken of falling is always a challenge issue for a bipedal humanoid robot in dealing with various locomotion related tasks to serve human society, especially as the assigned tasks turns increasingly complicated and the corresponding real environment gets more and more complex. Unlike several previous successful approaches on humanoid falling control, in this study, a new approach is suggested in the light of how human do when a fall happens. The proposed approach takes a tripod-like falling controller followed by a human-inspired active compliance strategy using joints' active flexion and torque increment. Therefore, robot falling action covers both stages of a fall, i.e. before and after landing impact, so that to reduce the fall damage as far as possible. And the tripod like posture prevents accumulation of kinetic energy, while the active compliance absorbs the impact energy in a tender way. Meanwhile, considering the complexity of robot dynamics, other than taking expert experiences, the proposed human-inspired falling control strategy is parametrically modelled and optimized with policy gradient reinforcement learning. Experiments on both simulation and real robot PKU-HR5.1 are performed, and the results demonstrate this approach is effective and promising. Dingsheng Luo, Yian Deng, Xiaoqiang Han, Xihong Wu |
IROS | 4 |
| 2016 | Frequency importance function of the speech intelligibility index for Mandarin Chinese
Jing Chen 0019, Xihong Wu |
Speech Commun. | 3 |
| 2015 | Constructing long short-term memory based deep recurrent neural networks for large vocabulary speech recognitionabstractLong short-term memory (LSTM) based acoustic modeling methods have recently been shown to give state-of-the-art performance on some speech recognition tasks. To achieve a further performance improvement, in this research, deep extensions on LSTM are investigated considering that deep hierarchical model has turned out to be more efficient than a shallow one. Motivated by previous research on constructing deep recurrent neural networks (RNNs), alternative deep LSTM architectures are proposed and empirically evaluated on a large vocabulary conversational telephone speech recognition task. Meanwhile, regarding to multi-GPU devices, the training process for LSTM networks is introduced and discussed. Experimental results demonstrate that the deep LSTM networks benefit from the depth and yield the state-of-the-art performance on this task. Xiangang Li, Xihong Wu |
ICASSP | 2 |
| 2015 | Improving long short-term memory networks using maxout units for large vocabulary speech recognitionabstractLong short-tem memory (LSTM) recurrent neural networks have been shown to give state-of-the-art performance on many speech recognition tasks. To achieve a further performance improvement, in this paper, maxout units are proposed to be integrated with the LSTM cells, considering those units have brought significant improvements to deep feed-forward neural networks. A novel architecture was constructed by replacing the input activation units (generally tanh) in the LSTM networks with maxout units. We implemented the LSTM network training on multi-GPU devices with truncated BPTT, and empirically evaluated the proposed designs on a large vocabulary Mandarin conversational telephone speech recognition task. The experimental results support our claim that the performance of LSTM based acoustic models can be further improved using the maxout units. Xiangang Li, Xihong Wu |
ICASSP | 2 |
| 2015 | Recognizing Human Activities from Raw Accelerometer Data Using Deep Neural NetworksabstractActivity recognition from wearable sensor data has been researched for many years. Previous works usually extracted features manually, which were hand-designed by the researchers, and then were fed into the classifiers as the inputs. Due to the blindness of manually extracted features, it was hard to choose suitable features for the specific classification task. Besides, this heuristic method for feature extraction could not generalize across different application domains, because different application domains needed to extract different features for classification. There was also work that used auto-encoders to learn features automatically and then fed the features into the K-nearest neighbor classifier. However, these features were learned in an unsupervised manner without using the information of the labels, thus might not be related to the specific classification task. In this paper, we recommend deep neural networks (DNNs) for activity recognition, which can automatically learn suitable features. DNNs overcome the blindness of hand-designed features and make use of the precious label information to improve activity recognition performance. We did experiments on three publicly available datasets for activity recognition and compared deep neural networks with traditional methods, including those that extracted features manually and auto-encoders followed by a K-nearest neighbor classifier. The results showed that deep neural networks could generalize across different application domains and got higher accuracy than traditional methods. Xihong Wu, Dingsheng Luo |
ICMLA | 2 |
| 2015 | Convolutional Networks Based Edge Detector Learned via Contrast Sensitivity Function
Haobin Dou, Wentao Liu 0002, Xihong Wu |
ICONIP (1) | 4 |
| 2015 | Learning to Reconstruct 3D Structure from Object Motion
Wentao Liu 0002, Haobin Dou, Xihong Wu |
ICONIP (1) | 3 |
| 2015 | Coarse-to-fine trained multi-scale Convolutional Neural Networks for image classificationabstractConvolutional Neural Networks (CNNs) have become forceful models in feature learning and image classification. They achieve translation invariance by spatial convolution and pooling mechanisms, while their ability in scale invariance is limited. To tackle the problem of scale variation in image classification, this work proposed a multi-scale CNN model with depth-decreasing multi-column structure. Input images were decomposed into multiple scales and at each scale image, a CNN column was instantiated with its depth decreasing from fine to coarse scale for model simplification. Scale-invariant features were learned by weights shared across all scales and pooled among adjacent scales. Particularly, a coarse-to-fine pre-training method imitating the human's development of spatial frequency perception was proposed to train this multi-scale CNN, which accelerated the training process and reduced the classification error. In addition, model averaging technique was used to combine models obtained during pre-training and further improve the performance. With these methods, our model achieved classification errors of 15.38% on CIFAR-10 dataset and 41.29% on CIFAR-100 dataset, i.e. 1.05% and 2.97% reduction compared with single-scale CNN model. Haobin Dou, Xihong Wu |
IJCNN | 2 |
| 2015 | Modeling speaker variability using long short-term memory networks for speech recognition
Xiangang Li, Xihong Wu |
INTERSPEECH | 2 |
| 2015 | Long short-term memory based convolutional recurrent neural networks for large vocabulary speech recognitionabstractLong short-term memory (LSTM) recurrent neural networks (RNNs) have been shown to give state-of-the-art performance on many speech recognition tasks, as they are able to provide the learned dynamically changing contextual window of all sequence history.On the other hand, the convolutional neural networks (CNNs) have brought significant improvements to deep feed-forward neural networks (FFNNs), as they are able to better reduce spectral variation in the input signal.In this paper, a network architecture called as convolutional recurrent neural network (CRNN) is proposed by combining the CNN and LSTM RNN.In the proposed CRNNs, each speech frame, without adjacent context frames, is organized as a number of local feature patches along the frequency axis, and then a LSTM network is performed on each feature patch along the time axis.We train and compare FFNNs, LSTM RNNs and the proposed LSTM CRNNs at various number of configurations.Experimental results show that the LSTM CRNNs can exceed stateof-the-art speech recognition performance. Xiangang Li, Xihong Wu |
INTERSPEECH | 2 |
| 2015 | I-vector dependent feature space transformations for adaptive speech recognition
Xiangang Li, Xihong Wu |
INTERSPEECH | 2 |
| 2015 | A comparative study on selecting acoustic modeling units in deep neural networks based large vocabulary Chinese speech recognition
Xiangang Li, Zaihu Pang, Xihong Wu |
Neurocomputing | 4 |
| 2014 | Exploiting limited data for parsingabstractData sparsity issues are extremely severe for parser due to the flexibility of tree structures. Many tags and productions appears a little, nevertheless, they are crucial for the parse disambiguation where it occurs. Besides, when a common tag somewhat regularly occurs in a non-canonical position, its distribution is usually distinct. In this paper, we propose a metric that measures the scarcity of any phrase with arbitrary span size. To make a better compromise between training trees with high confidence and scarcity, we try to catch some constraints in response to rare but articulating categories when training latent variable grammar. We exploits the limited data more sufficiently by capturing the depicting power of rate tree structure configuration in Expectation & Maximization procedure and Split & Merge framework. The resulting grammars are interpretable as our intension. Based on this approach, we further propose a method that exploits the limited training date from multiple perspectives, and accumulates their advantages in a product model. Despite its limited training data, out model improves parsing performance on Penn Chinese Treebank Fifth Edition, even higher than some systems with extra unlabeled data and external resources. Furthermore, this method is easy to generalized to cope with data sparsity in other natural language processing tasks. Xihong Wu |
ICIS | 3 |
| 2014 | Learning the Taxonomy of Function Words for Parsing
Dingsheng Luo, Xihong Wu |
COLING | 4 |
| 2014 | Query-based composition for large-scale language model in LVCSRabstractThis paper describes a query-based composition algorithm that can integrate an ARPA format language model in the unified WFST framework, which avoids the memory and time cost of converting the language models to WFST and optimizing the WFST of language models. The proposed algorithm is applied to on-the-fly one-pass decoder and rescoring decoder. Both modified decoder require less memory during decoding on different scale of language models. What's more, query-based on-the-fly one-pass decoder nearly has the same decoding speed as standard one and query-based rescoring decoder even use less time to rescore the lattice. Because of these advantages, large-scale language models can be applied by query-based composition algorithm to improve the performance of large vocabulary continuous speech recognition. Xiangang Li, Xihong Wu |
ICASSP | 5 |
| 2014 | A Cyclic Contrastive Divergence Learning Algorithm for High-Order RBMsabstractThe Restricted Boltzmann Machine (RBM), a special case of general Boltzmann Machines and a typical Probabilistic Graphical Models, has attracted much attention in recent years due to its powerful ability in extracting features and representing the distribution underlying the training data. A most commonly used algorithm in learning RBMs is called Contrastive Divergence (CD) proposed by Hinton, which starts a Markov chain at a data point and runs the chain for only a few iterations to get a low variance estimator. However, when referring to a high-order RBM, since there are interactions among its visible layers, the gradient approximation via CD learning usually becomes far from the log-likelihood gradient and even may cause CD learning to fall into an infinite loop with high reconstruction error. In this paper, a new algorithm named Cyclic Contrastive Divergence (CCD) is introduced for learning high-order RBMs. Unlike the standard CD algorithm, CCD updates the parameters according to each visible layer in turn, by borrowing the idea of Cyclic Block Coordinate Descent method. To evaluate the performance of the proposed CCD algorithm, regarding to high-order RBMs learning, both algorithms CCD and standard CD are theoretically analyzed, including convergence, estimate upper bound and both biases comparison, from which the superiority of CCD learning is revealed. Experiments on MNIST dataset for the handwritten digit classification task are performed. The experimental results show that CCD is more applicable and consistently outperforms the standard CD in both convergent speed and performance. Dingsheng Luo, Xiaoqiang Han, Xihong Wu |
ICMLA | 4 |
| 2014 | Parsing named entity as syntactic structure
Xihong Wu |
INTERSPEECH | 3 |
| 2014 | Visual gesture recognition for human robot interaction using dynamic movement primitivesabstractIn this paper a method to address the efficiency and robustness of dynamic hand gesture recognition for human robot interaction is proposed. By using on-board monocular camera and specialized gesture detection algorithms, the humanoid robot is able to detect gestures fast. To model the dynamics of gestures, the dynamic movement primitives (DMP) model is employed, which well characterizes both spatial and temporal evolutions of gestures. The invariance properties of the DMP model against different spatiotemporal scales also offer expected robustness to handle the variances in gestures. To cope with the diversity and noise of gestures, an efficient adaptive DMP learning method is further proposed. Since the learnt weights of the DMP compactly represent the original gestures, they serve as ideal feature vectors for building a classifier to recognize new gestures. To evaluate the proposed method, a nine-class human gestures recognition task on a real humanoid robot is performed and 98.06% accuracy is obtained. Experimental results demonstrate the effectiveness of our method. Dingsheng Luo, Xihong Wu |
SMC | 4 |
| 2014 | A comparative study of RPCL and MCE based discriminative training methods for LVCSR
Zaihu Pang, Shikui Tu, Xihong Wu, Lei Xu 0001 |
Neurocomputing | 3 |
| 2013 | Discriminative Apprenticeship Learning with Both Preference and Non-preference BehaviorabstractConsidering that expert's demonstrations are usually sub optimal and failed demonstrations often have some useful guidance, in this paper, a Discriminative Apprenticeship Learning algorithm is proposed, where the apprentice is taught with the join of failed attempts to acquire the ability that could discriminate the preference and non-preference cases so that to actively take a corresponding action. Since robot usually encounters changing environments, generalization ability is taken into account in the algorithm through which the reward function is recovered under the evaluation of generalization error. The problem of the representation error is also analyzed and involved in the algorithm. To ensure performance of the algorithm, theoretical guarantee is presented. Experiments on a simple car-driving robot and the comparison with a variety of inverse reinforcement learning methods are performed, which illustrate the proposed method is an effective and promising alternative. Dingsheng Luo, Xihong Wu |
ICMLA (1) | 3 |
| 2012 | Probabilistic Speaker-Class based Acoustic Modeling for Large Vocabulary Continuous Speech Recognition
Xiangang Li, Zaihu Pang, Xihong Wu |
INTERSPEECH | 4 |
| 2012 | Effects of aging on the ability to benefit from prior knowledge of message content in masked speech recognition
Meihong Wu, Huahui Li, Zhiling Hong, Xinchi Xian, Xihong Wu |
Speech Commun. | 6 |
| 2010 | GMM-HMM acoustic model training by a two level procedure with Gaussian components determined by automatic model selectionabstractThis paper investigates the Bayesian Ying-Yang (BYY) learning for speech recognition via Gaussian mixture models (GMMs) based Hidden Markov models (HMMs). A two level procedure is proposed with the hidden Markov level trained still under the maximum likelihood principle by the Baum-Welch algorithm but with the GMMs level trained under the BYY best harmony. We proposed a new batch way EM-like Ying-Yang alternation algorithm and used it as a plug-in block to the Baum-Welch algorithm. The advantage is that number of GMM components can be automatically determined during this BYY harmony learning and that the resulted model parameters become less affected than EM-ML training by the problem of overfitting and singular solution. In comparison with the standard EM-ML training and classical model selection criterions, including BIC and AIC, speech recognition experiments in a large vocabulary task on the Hub4 broadcast news database shown that the proposed algorithm provides an improved performance and also good convergence. Xihong Wu, Lei Xu 0001 |
ICASSP | 2 |
| 2010 | Maximum entropy based tone modeling for mandarin speech recognitionabstractTo explore the potential of prosody for Mandarin speech recognition, this paper addresses the tone modeling problem and its integration issue. This study adopts the maximum entropy approach to capture both acoustic and lexical characteristics of tones due to its flexibility in handling multiple interacting features. Moreover, considering the phoneme factor, besides a tone model, a phoneme dependent model is also constructed. With regard to the model integration, the presented models are integrated into the recognizer under the one-pass decoding framework, where they are used to prune the active word-final states during beam search. Experimental results on the HUB-4 evaluation material reveal the effectiveness of the presented models. They significantly improve the performance of speech recognition with 7.6% and 11.1% relative reduction of character error rate. Yansuo Yu, Xihong Wu, Huisheng Chi |
ICASSP | 3 |
| 2009 | Refining Grammars for Parsing with Hierarchical Semantic Knowledge
Xiaojun Lin 0002, Xihong Wu, Huisheng Chi |
EMNLP | 4 |
| 2009 | PHMM based asynchronous acoustic model for Chinese large vocabulary continuous speech recognitionabstractIn this paper, we presented an asynchronous multiple stream based Chinese tonal acoustic modeling framework. In this framework, toneless phonetic units and tones are modeled separately with different acoustic features. During the training and decoding process, a set of models are coupled together with a product hidden Markov models (PHMM) to represent whole tonal phonetic units. Through this, a compound context dependent tonal model can be generated from a few simple models. Experiments show that such model scheme generates more compact and accurate model presentation and brings improvement on the performance for large vocabulary speech recognition tasks. Xihong Wu, Huisheng Chi |
ICASSP | 2 |
| 2009 | Distance-Dependent Head-Related Transfer Functions Measured With High Spatial Resolution Using a Spark GapabstractA measurement of head-related transfer functions (HRTFs) with high spatial resolution was carried out in this study. HRTF measurement is difficult in the proximal region because of the lack of an appropriate acoustic point source. In this paper, a modified spark gap was used as the acoustic sound source. Our evaluation experiments showed that the spark gap was more like an acoustic point source than others previously used from the viewpoints of frequency response, directivity, power attenuation, and stability. Using this spark gap, high spatial resolution HRTFs were measured at 6344 spatial points, with distances from 20 to 160 cm, elevations from -40deg to 90deg, and azimuths from 0deg to 360deg. Based on these measurements, an HRTF database was obtained and its reliability was confirmed by both objective and subjective evaluations. Tianshu Qu, Mei Gong, Xihong Wu |
IEEE Trans. Speech Audio Process. | 6 |
| 2008 | Integrating Multi-level Linguistic Knowledge with a Unified Framework for Mandarin Speech Recognition
Jiazhong Nie, Dingsheng Luo, Xihong Wu |
EMNLP | 4 |
| 2008 | Exploiting prosodic and lexical features for tone modeling in a conditional random field frameworkabstractTonal cues play an important role in distinguishing ambiguous words in Mandarin speech recognition. This paper explores an innovative tone modeling framework using prosodic and lexical features, as well as syllable context information. A discriminative model, namely a Conditional Random Field (CRF), is adopted, which is sufficiently flexible to handle multiple interacting features and long-range dependencies of observations. After the first pass search of a recognition system, the CRF based tone models are employed to rerank N-best hypotheses according to the tonal scores which can represent the correctness of the tone sequence given each candidate hypothesis and the observed speech signal. Experiments results show that the tonal cues help to achieve 7.8% and 8.6% relative reductions of character error rate on two widely used Mandarin speech recognition tasks, Hub-4 test and 863 test. Hongxiu Wei, Dingsheng Luo, Xihong Wu |
ICASSP | 5 |
| 2008 | Monaural speech separation based on multi-scale Fan-Chirp TransformabstractA novel method for monaural speech separation is presented in this paper. Instead of the traditional short time Fourier transform (STFT) for time-frequency analysis in speech separation, the fan-chirp transform (FChT) has been applied to track the pitch and harmonics of the target speech. This method has two advantages over STFT. Firstly, the spectrum spread of dynamic harmonics within each analysis frame has been alleviated. Secondly, the FChT bases with proper chirp rate could be chosen according to different frequency modulation rates in the simultaneous speech. Furthermore, considering the changeability of frequency modulation rates, a multi-scale FChT is proposed to adaptively adjust the frame length of spectrum analysis. Experimental results prove the validity of the approach in monaural speech separation. Xihong Wu |
ICASSP | 3 |
| 2008 | An Improved CRF based Chinese Language Processing System for SIGHAN Bakeoff 2007
Xihong Wu, Xiaojun Lin 0002, Chunyao Wu, Dianhai Yu |
IJCNLP | 1 |
| 2008 | Probabilistic latent speaker training for large vocabulary speech recognitionabstractIn this paper, we describe an improvement on probabilistic latent speaker analysis method and investigate the use of probabilistic latent speaker analysis for acoustic model training. By performing co-occurrence analysis between speaker and dominant components, speaker variation is dealt with based on different trajectories. Speech recognition experiment results show that our method, although with a general acoustic model and one-pass decoding, outperform the gender-dependent acoustic model with each gender is given for test set. Further experiment shows that the probabilistic latent speaker training method, although with no adaptation stage and no adaptation data, has outperformed the eigenMLLR adaptation method. Xihong Wu, Huisheng Chi |
INTERSPEECH | 2 |
| 2008 | A Joint Segmenting and Labeling Approach for Chinese Lexical Analysis
Jiazhong Nie, Dingsheng Luo, Xihong Wu |
ECML/PKDD (2) | 4 |
| 2007 | Refine bigram PLSA model by assigning latent topics unevenlyabstractAs an important component in many speech and language processing applications, statistical language model has been widely investigated. The bigram topic model, which combines advantages of both the traditional n-gram model and the topic model, turns out to be a promising language modeling approach. However, the original bigram topic model assigns the same topic number for each context word but ignores the fact that there are different complexities to the latent semantics of context words. we present a new bigram topic model, the bigram PLSA model, and propose a modified training strategy that unevenly assigns latent topics to context words according to an estimation of their latent semantic complexities. As a consequence, a refined bigram PLSA model is reached. Experiments on HUB4 Mandarin test transcriptions reveal the superiority over existing models and further performance improvements on perplexity are achieved through the use of the refined bigram PLSA model. Jiazhong Nie, Runxin Li, Dingsheng Luo, Xihong Wu |
ASRU | 4 |
| 2007 | Probabilistic latent speaker analysis for large vocabulary speech recognitionabstractTrajectory folding problem is intrinsic for HMM-based speech recognition systems in which each state is modeled by a mixture of Gaussian components. In this paper, a probabilistic latent semantic analysis (PLSA)-based approach is proposed for use in speech recognition systems to alleviate this problem. The basic idea is that different speech trajectories are strongly correlated with speaker variation, and different speakers may have high scores on certain Gaussian components consistently. Thus, PLSA is adopted to perform co-occurrence analysis between Gaussian components and speakers and provide additional source of information to constrain searching path during decoding procedure. Experimental results show that 11.2% and 2.7% relative reduction on word error rate can be achieved on a homogeneous test set and the 2004 863 evaluation set, respectively. Xihong Wu, Huisheng Chi |
INTERSPEECH | 2 |
| 2007 | Effect of number of masking talkers on speech-on-speech masking in Chinese
Xihong Wu, Jing Chen 0019 |
INTERSPEECH | 1 |
| 2007 | Context dependent syllable acoustic model for continuous Chinese speech recognitionabstractThe choice of basic modeling unit in building acoustic model for a continuous Mandarin speech recognition task is a very important issue [1]. Unlike traditional phoneme or Initial/Finals (IFs) units based acoustic modeling methods, which usually suffer from the limitations of less accuracy in modeling intrasyllable variations and long scale temporal dependencies, in this paper, a practicable syllable based approach is presented. In contrast with IFs, syllable can implicitly model the intrasyllable variations in good accuracy. Also, by carefully choosing context modeling schemes and parameter tying methods, syllable based acoustic model can capture longer temporal variations while keeping the complexity of model well controlled. Meanwhile, considering the data unbalanced problem, multiple sized unit model based approaches are also implemented in this research. The experiment result shows the acoustic model based on the presented syllable based approach is effective in improving the performance of the Chinese continuous speech recognition. Xihong Wu |
INTERSPEECH | 2 |
| 2007 | The effect of voice cuing on releasing Chinese speech from informational masking
Jing Chen 0019, Xihong Wu, Bruce A. Schneider |
Speech Commun. | 4 |
| 2006 | CASA based speech separation for robust speech recognitionabstractThis paper introduces a speech separation system as a front-end processing step for automatic speech recognition (ASR). It employs computational auditory scene analysis (CASA) to separate the target speech from the interference speech. Specifically, the mixed speech is preprocessed based on auditory peripheral model. Then a pitch tracking is conducted and the dominant pitch is used as a main cue to find the target speech. Next, the time frequency (TF) units are merged into many segments. These segments are then combined into streams via CASA initial grouping. A regrouping strategy is employed to refine these streams via amplitude modulate (AM) cues, which are finally organized by the speaker recognition techniques into corresponding speakers. Finally, the output streams are reconstructed to compensate the missing data in the abovementioned processing steps by a cluster based feature reconstruction. Experimental results of ASR show that at low TMR (<-6dB) the proposed method offers significantly higher recognition accuracy. Index term: speech separation, CASA, speaker recognition, pitch tracking, units grouping, reconstruction Runqiang Han, Qin Gao, Xihong Wu |
INTERSPEECH | 6 |
| 2006 | Just-in-Time Latent Semantic Adaptation on Language Model for Chinese speech Recognition Using Web DataabstractA novel method is proposed, which is for performing just-in- time adaptation on language models in Chinese speech recognition using Web search engines. Latent semantic analysis (LSA) is employed to change the probability distribution of N-gram language model. The method has two advantages. First, it needs relatively small amount of data which can be obtained from Web on-the-fly. Second, comparing to traditional adaptation formula of LSA, the proposed approach is more efficient, which ensures second pass decoding to be performed with high speed. Experiments show that the perplexity of language model is reduced by over 13% after adaptation. A 4.29% relative reduction on WER is achieved in large vocabulary Chinese speech recognition over standard test set. Qin Gao, Xiaojun Lin 0002, Xihong Wu |
SLT | 3 |
| 2003 | Biomimetics speaker identification systems for network security gatekeepersabstractThe perception mechanisms of the human auditory periphery and cochlear nucleus were simulated and the potential application to the voice password gatekeeper was discussed. A biomimetics speaker identification system was implemented based on the auditory processing. Obvious improvement in the robustness was shown under a noisy environment. Xihong Wu, Dingsheng Luo, Huisheng Chi, Harold Szu |
IJCNN | 1 |
| 2000 | An auditory feature extraction method based on forward-masking and its application in robust speaker identification and speech recognition
Xihong Wu, Bin Zhen, Huisheng Chi |
INTERSPEECH | 2 |
| 2000 | On the importance of components of the MFCC in speech and speaker recognition
Bin Zhen, Xihong Wu, Huisheng Chi |
INTERSPEECH | 2 |
| 2000 | On the use of bandpass liftering in speaker recognition
Bin Zhen, Xihong Wu, Huisheng Chi |
INTERSPEECH | 2 |
| 1999 | Auditory model based speech feature extraction and its application to speaker identificationabstractAccording to the characteristics of the auditory periphery and cochlear nucleus, as well as attempting to simulate the mechanism of auditory system as a whole, two kinds of novel speech feature are presented in this paper, and a framework of neural network has been adopted. The two features considered are: the weighted average localized synchronized rate cepstrum, and the weighted firing rate cepstrum. Both of them are applied to speaker identification. The modular tree and modified linear opinion pools are used as classifiers to simulate the parallel processing mechanism of the upper level function of auditory system. Good recognition accuracy is obtained under both clean and noisy environments. Bing Xiang, Xihong Wu, Huisheng Chi |
IJCNN | 2 |