Yu Hu 0003

dblp:08/6001-3 · DBLP profile ↗
← Back
39ranked-venue papers
2as first author
9since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding
abstract
Accurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-level annotations. This limitation overlooks the hierarchically structured representation of documents, constraining comprehensive understanding of complex forms. To address this issue, we present the SRFUND, a hierarchically structured multi-task form understanding benchmark. SRFUND provides refined annotations on top of the original FUNSD and XFUND datasets, encompassing five tasks: (1) word to text-line merging, (2) text-line to entity merging, (3) entity category classification, (4) item table localization, and (5) entity-based full-document hierarchical structure recovery. We meticulously supplemented the original dataset with missing annotations at various levels of granularity and added detailed annotations for multi-item table regions within the forms. Additionally, we introduce global hierarchical structure dependencies for entity relation prediction tasks, surpassing traditional local key-value associations. The SRFUND dataset includes eight languages including English, Chinese, Japanese, German, French, Spanish, Italian, and Portuguese, making it a powerful tool for cross-lingual form understanding. Extensive experimental results demonstrate that the SRFUND dataset presents new challenges and significant opportunities in handling diverse layouts and global hierarchical structures of forms, thus providing deep insights into the field of form understanding. The original dataset and implementations of baseline methods are available at https://sprateam-ustc.github.io/SRFUND.
Jiefeng Ma, Jun Du 0002, Yu Hu 0003, Pengfei Hu 0006, Qing Wang 0008, Jianshu Zhang 0001
NeurIPS5
2023 Joint optimization for attention-based generation and recognition of chinese characters using tree position embedding
Mobai Xue, Jun Du 0002, Bin Wang 0070, Bo Ren 0002, Yu Hu 0003
Pattern Recognit.5
2023 QDM-SSD: Quality-Aware Dynamic Masking for Separation-Based Speaker Diarization
abstract
We improve iterative separation-based speaker diarization (ISSD) with quality-aware dynamic masking (QDM). We call the proposed framework QDM-SSD. Compared with ISSD, QDM-SSD enhances the simulated data used for model adaptation through QDM to alleviate the influence of errors in speaker priors. In addition to data quality purification, QDM-SSD also makes the adaptation data sparse by automatically adjusting speaker overlap ratios according to data quality. Furthermore, using a sliding window over the adaptation data, clean regions in speech segments can be better localized. Experiments on the two-speaker conversational telephone speech (CTS) corpus show that the proposed QDM-SSD framework can reduce the diarization error rate (DER) by 18.56% relatively compared with ISSD. Moreover, QDM-SSD is shown to generalize to other two-speaker non-conversation telephone speech data sets where ISSD fails to work. Finally, we demonstrate that QDM-SSD can serve as a front-end to improve the performances of back-end automatic speech recognition.
Shutong Niu, Jun Du 0002, Lei Sun 0010, Yu Hu 0003, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Tracking Interaction States for Multi-Turn Text-to-SQL Semantic Parsing
abstract
The task of multi-turn text-to-SQL semantic parsing aims to translate natural language utterances in an interaction into SQL queries in order to answer them using a database which normally contains multiple table schemas. Previous studies on this task usually utilized contextual information to enrich utterance representations and to further influence the decoding process. While they ignored to describe and track the interaction states which are determined by history SQL queries and are related with the intent of current utterance. In this paper, two kinds of interaction states are defined based on schema items and SQL keywords separately. A relational graph neural network and a non-linear layer are designed to update the representations of these two states respectively. The dynamic schema-state and SQL-state representations are then utilized to decode the SQL query corresponding to current utterance. Experimental results on the challenging CoSQL dataset demonstrate the effectiveness of our proposed method, which achieves better performance than other published methods on the task leaderboard.
Zhen-Hua Ling, Jingbo Zhou 0003, Yu Hu 0003
AAAI4
2021 Automatic Lip-Reading with Hierarchical Pyramidal Convolution and Self-Attention for Image Sequences with No Word Boundaries
Hang Chen 0001, Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Chin-Hui Lee 0001
Interspeech3
2021 Adversarial Voice Conversion Against Neural Spoofing Detectors
Yi-Yang Ding, Li-Juan Liu, Yu Hu 0003, Zhen-Hua Ling
Interspeech3
2021 Correlating subword articulation with lip shapes for embedding aware audio-visual speech enhancement
Hang Chen 0001, Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Chin-Hui Lee 0001
Neural Networks3
2021 Robustness of Speech Spoofing Detectors Against Adversarial Post-Processing of Voice Conversion
abstract
With the development of speech synthesis and voice conversion techniques, the quality of artificially generated speech has been significantly improved and detecting such spoofing speech becomes crucial to practical applications, such as automatic speaker verification (ASV). State-of-the-art neural-network-based spoofing detection models can distinguish most artificial utterances from natural ones effectively in the latest ASVspoof 2019 evaluation. Motivated by recent progresses of adversarial example generation, this paper studies the robustness of neural-network-based speech spoofing detectors against adversarial attacks. To this end, an adversarial post-processing network (APN) is proposed which generates adversarial examples against a white-box anti-spoofing model by post-processing the speech waveforms produced by a baseline voice conversion system. Experimental results demonstrate the adversarial ability of our proposed APNs against the white-box anti-spoofing models which were used as the adversarial targets of APNs at the training stage. For example, the equal error rate (EER) of a fused detection model based on light convolution neural networks (LCNNs) increased from 0.278% to 12.743% under the white-box condition without degrading the subjective quality of converted speech. Furthermore, the trained APNs can also perform against the detectors with either unseen structures or unseen features by raising their EERs in our experiments. All these results indicate the threat of adversarial speech generation to the performance of state-of-the-art spoofing detection models.
Yi-Yang Ding, Hao-Jian Lin, Li-Juan Liu, Zhen-Hua Ling, Yu Hu 0003
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 A Multiple-Integration Encoder for Multi-Turn Text-to-SQL Semantic Parsing
abstract
This paper studies multi-turn text-to-SQL generation, which is a new but important task in semantic parsing. In order to deal with its two challenges, i.e., multi-turn interaction and cross-domain evaluation, this paper proposes a multiple-integration encoder, which derives the vector representations of user utterances and database schemas using three custom-designed modules for information integration. First, an utterance representation enhancing module is built to integrate the information of history utterances into the representation of each token in current utterance by attentive selection. Second, a schema discrepancy enhancing module is designed to integrate previous predicted SQL query into the representation of schema items. Third, a latent schema linking module is employed to integrate schema information into utterance representations for better dealing with unseen database schemas. These three modules are all implemented based on a lightweight multi-head attention mechanism, which reduces the number of parameters in conventional multi-head attention. Experimental results on the SParC dataset show that our method achieved better accuracy of multi-turn text-to-SQL generation than the most advanced benchmarks. Further ablations studies and analysis also demonstrate the effectiveness of the three modules designed for information integration in the encoder.
Zhen-Hua Ling, Jing-Bo Zhou, Yu Hu 0003
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Cause-Effect Knowledge Acquisition and Neural Association Model for Solving A Set of Winograd Schema Problems
abstract
This paper focuses on the investigations in Winograd Schema (WS), a challenging problem which has been proposed for measuring progress in commonsense reasoning.Due to the lack of commonsense knowledge and training data, very little work has been found on the WS problems in recent years.Actually, there is no shortcut to solve this problem except to collect more commonsense knowledge and design suitable models.Therefore, this paper addresses a set of WS problems by proposing a knowledge acquisition method and a general neural association model.To avoid the sparseness issue, the knowledge we aim to collect is the cause-effect relationships between thousands of commonly used words.The knowledge acquisition method supports us to extract hundreds of thousands of cause-effect pairs from large text corpus automatically.Meanwhile, a neural association model (NAM) is proposed to encode the association relationships between any two discrete events.Based on the extracted knowledge and the NAM models, in this paper, we successfully build a system for solving WS problems from scratch and achieve 70.0% accuracy.Most importantly, this paper provides a flexible framework to solve WS problems based on event association and neural network methods.
Quan Liu 0003, Hui Jiang 0001, Andrew Evdokimov, Zhen-Hua Ling, Xiaodan Zhu 0001, Si Wei, Yu Hu 0003
IJCAI7
2017 Towards human-like and transhuman perception in AI 2.0: a review
abstract
Perception is the interaction interface between an intelligent system and the real world. Without sophisticated and flexible perceptual capabilities, it is impossible to create advanced artificial intelligence (AI) systems. For the next-generation AI, called ‘AI 2.0’, one of the most significant features will be that AI is empowered with intelligent perceptual capabilities, which can simulate human brain’s mechanisms and are likely to surpass human brain in terms of performance. In this paper, we briefly review the state-of-the-art advances across different areas of perception, including visual perception, auditory perception, speech perception, and perceptual information processing and learning engines. On this basis, we envision several R&D trends in intelligent perception for the forthcoming era of AI 2.0, including: (1) human-like and transhuman active vision; (2) auditory perception and computation in an actual auditory setting; (3) speech perception and computation in a natural interaction setting; (4) autonomous learning of perceptual information; (5) large-scale perceptual information processing and learning platforms; and (6) urban omnidirectional intelligent perception and reasoning engines. We believe these research directions should be highlighted in the future plans for AI 2.0.
Yonghong Tian 0001, Xilin Chen 0001, Hongkai Xiong, Li-Rong Dai 0001, Jing Chen 0002, Junliang Xing, Jing Chen 0003, Xihong Wu, Weiming Hu 0004, Yu Hu 0003, Tiejun Huang 0001, Wen Gao 0001
Frontiers Inf. Technol. Electron. Eng.11
2017 Nonrecurrent Neural Structure for Long-Term Dependence
abstract
In this paper, we propose a novel neural network structure, namely feedforward sequential memory networks (FSMN), to model long-term dependence in time series without using recurrent feedback. The proposed FSMN is a standard fully connected feedforward neural network equipped with some learnable memory blocks in its hidden layers. The memory blocks use a tapped-delay line structure to encode the long context information into a fixed-size representation as short-term memory mechanism which are somehow similar to the time-delay neural networks layers. We have evaluated the FSMNs in several standard benchmark tasks, including speech recognition and language modeling. Experimental results have shown that FSMNs outperform the conventional recurrent neural networks (RNN) while can be learned much more reliably and faster in modeling sequential signals like speech or language. Moreover, we also propose a compact feedforward sequential memory networks (cFSMN) by combining FSMN with low-rank matrix factorization and make a slight modification to the encoding method used in FSMNs in order to further simplify the network architecture. On the speech recognition Switchboard task, the proposed cFSMN structures can reduce the model size by 60% and speed up the learning by more than seven times while the model can still significantly outperform the popular bidirectional LSTMs for both frame-level cross-entropy criterion-based training and MMI-based sequence training.
Shiliang Zhang, Cong Liu 0006, Hui Jiang 0001, Si Wei, Li-Rong Dai 0001, Yu Hu 0003
IEEE ACM Trans. Audio Speech Lang. Process.6
2016 Exploring Semantic Representation in Brain Activity Using Word Embeddings
abstract
In this paper, we utilize distributed word representations (i.e., word embeddings) to analyse the representation of semantics in brain activity.The brain activity data were recorded using functional magnetic resonance imaging (fMRI) when subjects were viewing words.First, we analysed the functional selectivity of different cortex areas by calculating the correlations between neural responses and several types of word representations, including skipgram word embeddings, visual semantic vectors, and primary visual features.The results demonstrated consistency with existing neuroscientific knowledge.Second, we utilized behavioural data as the semantic ground truth to measure their relevance with brain activity.A method to estimate word embeddings under the constraints of brain activity similarities is further proposed based on the semantic word embedding (SWE) model.The experimental results show that the brain activity data are significantly correlated with the behavioural data of human judgements on semantic similarity.The correlations between the estimated word embeddings and the semantic ground truth can be effectively improved after integrating the brain activity data for learning, which implies that semantic patterns in neural representations may exist that have not been fully captured by state-of-the-art word embeddings derived from text corpora.
Yu-Ping Ruan, Zhen-Hua Ling, Yu Hu 0003
EMNLP3
2016 Modulation spectrum compensation for HMM-based speech synthesis using line spectral pairs
abstract
In previous work, a method to compensate the divergence between the distributions of natural and generated modulation spectra (MS) has been proposed for hidden Markov model (HMM) based speech synthesis. This method can alleviate the over-smoothing effect of parameter generation when Mel-cepstral coefficients (MCC) are used as spectral features. This paper further investigates the MS compensation method for line spectral pairs (LSP). Four approaches to extract MS from LSPs are implemented and compared. These approaches calculate MS vectors using original LSP sequences, log power spectra (LPS) derived from LSPs, MCCs derived from LSPs, and MCCs derived from speech waveforms, respectively. Experimental results show that the naturalness of synthetic speech gets improved after MS compensation when LSPs are used as spectral features for HMM modeling. The degree of improvement depends on the type of spectral features for MS calculation significantly. MCCs derived from LSPs are more suitable for MS compensation than original LSPs and LPS derived from LSPs. Besides, using MCCs derived from speech waveforms also achieves satisfactory performance. This means that MS compensation can also be implemented as a post-filter to synthetic waveforms which does not rely on the type of spectral features and vocoders adopted in the synthesis system.
Zhen-Hua Ling, Xiao-Hui Sun, Li-Rong Dai 0001, Yu Hu 0003
ICASSP4
2016 Intra-Topic Variability Normalization based on Linear Projection for Topic Classification
abstract
This paper proposes a variability normalization algorithm to reduce the variability between intra-topic documents for topic classification.Firstly, an optimization problem is constructed based on linear variability removable assumption.Secondly, a new feature space for document representation is found by solving the optimization problem with kernel principle component analysis (KPCA).Finally, effective feature transformation is taken through linear projection.As for experiments, state-of-the-art SVM and KNN algorithm are adopted for topic classification respectively.Experimental results on a free-style conversational corpus show that the proposed variability normalization algorithm for topic classification achieves 3.8% absolute improvement for micro-F 1 measure.
Quan Liu 0003, Wu Guo, Zhen-Hua Ling, Hui Jiang 0001, Yu Hu 0003
HLT-NAACL5
2015 Learning Semantic Word Embeddings based on Ordinal Knowledge Constraints
abstract
Quan Liu, Hui Jiang, Si Wei, Zhen-Hua Ling, Yu Hu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Quan Liu 0003, Hui Jiang 0001, Si Wei, Zhen-Hua Ling, Yu Hu 0003
ACL (1)5
2015 State-Clustering Based Multiple Deep Neural Networks Modeling Approach for Speech Recognition
abstract
The hybrid deep neural network (DNN) and hidden Markov model (HMM) has recently achieved dramatic performance gains in automatic speech recognition (ASR). The DNN-based acoustic model is very powerful but its learning process is extremely time-consuming. In this paper, we propose a novel DNN-based acoustic modeling framework for speech recognition, where the posterior probabilities of HMM states are computed from multiple DNNs (mDNN), instead of a single large DNN, for the purpose of parallel training towards faster turnaround. In the proposed mDNN method all tied HMM states are first grouped into several disjoint clusters based on data-driven methods. Next, several hierarchically structured DNNs are trained separately in parallel for these clusters using multiple computing units (e.g. GPUs). In decoding, the posterior probabilities of HMM states can be calculated by combining outputs from multiple DNNs. In this work, we have shown that the training procedure of the mDNN under popular criteria, including both frame-level cross-entropy and sequence-level discriminative training, can be parallelized efficiently to yield significant speedup. The training speedup is mainly attributed to the fact that multiple DNNs are parallelized over multiple GPUs and each DNN is smaller in size and trained by only a subset of training data. We have evaluated the proposed mDNN method on a 64-hour Mandarin transcription task and the 320-hour Switchboard task. Compared to the conventional DNN, a 4-cluster mDNN model with similar size can yield comparable recognition performance in Switchboard (only about 2% performance degradation) with a greater than 7 times speed improvement in CE training and a 2.9 times improvement in sequence training, when 4 GPUs are used.
Hui Jiang 0001, Li-Rong Dai 0001, Yu Hu 0003, Qingfeng Liu
IEEE ACM Trans. Audio Speech Lang. Process.4
2011 Boosted Mixture Learning of Gaussian Mixture Hidden Markov Models Based on Maximum Likelihood for Speech Recognition
abstract
In this paper, we apply the well-known boosted mixture learning (BML) method to learn Gaussian mixture HMMs in speech recognition. BML is an incremental method to learn mixture models for classification problems. In each step of BML, one new mixture component is estimated according to the functional gradient of an objective function to ensure that it is added along the direction that maximizes the objective function. Several techniques have been proposed to extend BML from simple mixture models like the Gaussian mixture model (GMM) to the Gaussian mixture hidden Markov model (HMM), including Viterbi approximation for state segmentation, weight decay and sampling boosting to initialize sample weights to avoid overfitting, combination between partial updating and global updating to refine model parameters in each BML iteration, and use of the Bayesian Information Criterion (BIC) for parsimonious modeling. Experimental results on two large-vocabulary continuous speech recognition tasks, namely the WSJ-5k and Switchboard tasks, have shown that the proposed BML yields significant performance gain over the conventional training procedure, especially for small model sizes.
Jun Du 0002, Yu Hu 0003, Hui Jiang 0001
IEEE Trans. Speech Audio Process.2
2011 Trust Region-Based Optimization for Maximum Mutual Information Estimation of HMMs in Speech Recognition
abstract
In this paper, we have proposed two novel optimization methods for discriminative training (DT) of hidden Markov models (HMMs) in speech recognition based on an efficient global optimization algorithm used to solve the so-called trust region (TR) problem, where a quadratic function is minimized under a spherical constraint. In the first method, maximum mutual information estimation (MMIE) of Gaussian mixture HMMs is formulated as a standard TR problem so that the efficient global optimization method can be used in each iteration to maximize the auxiliary function of discriminative training for speech recognition. In the second method, we propose to construct a new auxiliary function for DT of HMMs by adding a quadratic penalty term. The new auxiliary function is constructed to serve as first-order approximation as well as lower bound of the original discriminative objective function within a locality constraint. Due to the lower-bound property, the found optimal point of the new auxiliary function is guaranteed to improve the original discriminative objective function until it converges to a local optimum or stationary point of the objective function. Both TR-based optimization methods have been investigated on two standard large-vocabulary continuous speech recognition tasks, using the WSJ0 and Switchboard databases. Experimental results have shown that the proposed TR methods outperform the conventional EBW method in terms of convergence behavior as well as recognition performance.
Cong Liu 0006, Yu Hu 0003, Li-Rong Dai 0001, Hui Jiang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2010 HMM-based pseudo-clean speech synthesis for splice algorithm
abstract
In this paper, we present a novel approach to relax the constraint of stereo-data which is needed in a series of algorithms for noise-robust speech recognition. As a demonstration in SPLICE algorithm, we generate the pseudo-clean features to replace the ideal clean features from one of the stereo channels, by using HMM-based speech synthesis. Experimental results on aurora2 database show that the performance of our approach is comparable with that of SPLICE. Further improvements are achieved by concatenating a bias adaptation algorithm to handle unknown environments. Relative word error rate reductions of 66% and 24% are achieved over the baseline systems in the clean-training and multi-training conditions, respectively.
Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Renhua Wang
ICASSP2
2010 A bounded trust region optimization for discriminative training of HMMS in speech recognition
abstract
In this paper, we have proposed a new method to construct an auxiliary function for the discriminative training of HMMs in speech recognition. The new auxiliary function serves as a first-order approximation of the original objective function but more importantly it remains as a lower bound of the original objective function as well. Furthermore, the trust region (TR) method in [1] is applied to find the globally optimal point of the new auxiliary function. Due to its lower-bound property, the found optimal point is theoretically guaranteed to increase the original discriminative objective function. The proposed bounded trust region method has been investigated on two LVCSR tasks, namely WSJ-5k and Switchboard 60-hour subset tasks. Experimental results show that the bounded TR method yields much better convergence behavior than both the conventional EBW method and the original TR method.
Cong Liu 0006, Yu Hu 0003, Hui Jiang 0001, Li-Rong Dai 0001
ICASSP2
2010 Boosted mixture learning of Gaussian mixture HMMs for speech recognition
abstract
In this paper, we propose a novel boosted mixture learning (BML) framework for Gaussian mixture HMMs in speech recognition. BML is an incremental method to learn mixture models for classification problem. In each step of BML, one new mixture component is calculated according to functional gradient of an objective function to ensure that it is added along the direction to maximize the objective function the most. Several techniques have been proposed to extend BML from simple mixture models like Gaussian mixture model (GMM) to Gaussian mixture hidden Markov model (HMM), including Viterbi approximation to obtain state segmentation, weight decay to initialize sample weights to avoid overfitting, combining partial updating with global updating of parameters and using Bayesian information criterion (BIC) for parsimonious modeling. Experimental results on the WSJ0 task have shown that the proposed BML yields relative word and sentence error rate reduction of 10.9% and 12.9%, respectively, over the conventional training procedure.
Jun Du 0002, Yu Hu 0003, Hui Jiang 0001
INTERSPEECH2
2010 Global variance modeling on the log power spectrum of LSPs for HMM-based speech synthesis
Zhen-Hua Ling, Yu Hu 0003, Li-Rong Dai 0001
INTERSPEECH2
2009 A trust region based optimization for maximum mutual information estimation of HMMS in speech recognition
abstract
In this paper, we present a new optimization method for MMIE-based discriminative training of HMMs in speech recognition. In our method, the MMIE training of Gaussian mixture HMMs is formulated as a so-called trust region problem, where a quadratic objective function is minimized under a spherical constraint, so that an efficient global optimization method for the trust region problem can be used to solve the MMIE training problem of HMMs. Experimental results on the WSJ0 Nov'92 evaluation task demonstrate that the trust region based optimization significantly outperforms the conventional EBWmethod in terms of optimization convergence behavior as well as speech recognition performance. It has been observed that the trust region method achieves up to 23.3% relative recognition error reduction over a well-trained MLE system while the EBW method gives only 13.3% relative error reduction.
Zhijie Yan, Cong Liu 0006, Yu Hu 0003, Hui Jiang 0001
ICASSP3
2009 A new method for mispronunciation detection using Support Vector Machine based on Pronunciation Space Models
Si Wei, Yu Hu 0003, Renhua Wang
Speech Commun.3
2008 Heteroscedastic discriminant analysis with two-dimensional constraints
abstract
Heteroscedastic discriminant analysis (HDA) with two-dimensional (2D) constraints is proposed in this paper. HDA suffers from the small sample size problem and instability when lack of training data or feature dimension is high, even when the number of dimension is in a suitable range. Two-dimensional HDA is first proposed, then we show that 2D methods are actually a kind of structure-constrained 1D methods, and lastly, HDA with 2D constraints is proposed. Experiments on TIMIT and WSJ0 show that the proposed method outperforms other methods.
Sibao Chen 0001, Yu Hu 0003, Bin Luo 0001, Renhua Wang
ICASSP2
2008 Minimum word classification error training of HMMS for automatic speech recognition
abstract
This paper presents a novel discriminative training criterion, minimum word classification error (MWCE). By localizing conventional string-level MCE loss function to word-level, a more direct measure of empirical word classification error is approximated and minimized. Because the word-level criterion better matches performance evaluation criteria such as WER, an improved word recognition performance can be achieved. We evaluated and compared MWCE criterion in a unified DT framework, with other commonly-used criteria including MCE, MMI, MWE, and MPE. Experiments on TIMIT and WS JO evaluation tasks suggest that word-level MWCE criterion can achieve consistently better results than string-level MCE. MWCE even outperforms other substring-level criteria on the above two tasks, including MWE and MPE.
Zhijie Yan, Yu Hu 0003, Renhua Wang
ICASSP3
2006 Word structure and tone perception in Mandarin
abstract
This paper presents results concerning the relationship between word structure in terms of number of syllables and tonal realization in Mandarin. It examines whether the fact that a word (in our context a prosodic word) is more complex implies certain tonal reductions. Our hypothesis is that a monosyllabic word will be uttered more carefully than a polysyllabic word due to the potentially larger number of possibly confusable words. We also examine whether the total number of syllables in a word has an effect, creating more tonal reductions in longer than in shorter words. A database of Mandarin originally designed for concatenative speech synthesis and segmented into prosodic words was statistically analyzed regarding the occurrences of syllable/tone combinations in prosodic words of varying length. 10 sets of syllables were selected comprising all four tones of Mandarin and occurring as monosyllabic words as well as in varying positions in two- to five-syllable prosodic words. The target syllables were then extracted from their original context and presented to native speakers of Mandarin who had to decide which tone they perceived. The results of the perception test indicate, inter alia, that perception of syllables taken from polysyllables indeed is more error prone than that of monosyllabic words. The number of syllables in a word, however, has only a weak influence. Furthermore, reductions mostly appear for syllables in certain locations in a word and are related with underlying syllables' durations. Index Terms: Speech production and perception, tone languages
Hansjörg Mixdorff, Yu Hu 0003
INTERSPEECH2
2006 Automatic Mandarin pronunciation scoring for native learners with dialect accent
Si Wei, Qing-Sheng Liu, Yu Hu 0003, Renhua Wang
INTERSPEECH3
2005 A Novel Source Analysis Method by Matching Spectral Characters of LF Model with STRAIGHT Spectrum
Zhen-Hua Ling, Yu Hu 0003, Renhua Wang
ACII2
2005 Cross-language perception of word stress
abstract
This paper presents a study of the perception of Mandarin disyllabic words by native speakers of German. It examines how speakers of an accent language perceive word stress in words from a tone language. A corpus of 15 sets of words with all possible combinations of the four tones of Mandarin was recorded by a professional speaker. In addition monotonized versions of the words were created. In a forced-choice listening experiment native speakers of German were asked to assess whether they perceived the word stress on the first or second syllable. Results include, inter alia, that words with two high tones, as well as the monotonized stimuli were predominantly perceived as carrying the word stress on the first syllable. Words with a falling tone on the second syllable were mostly classified as carrying stress on the second syllable, with the combination of low and falling tone yielding the highest score. Many combinations of tones, however, could not be identified as any of the two kinds. This suggests that though some tonal configurations in Mandarin are similar to German two-syllable word accent patterns and can be associated with the latter, others might be rather interpreted as pertaining to two mono-syllabic words, both of which are stressed.
Hansjörg Mixdorff, Yu Hu 0003
INTERSPEECH2
2005 Visual cues in Mandarin tone perception
abstract
This paper presents results concerning the exploitation of visual cues in the perception of Mandarin tones. The lower part of a female speaker's face was recorded on digital video as she uttered 25 sets of syllabic tokens covering the four different tones of Mandarin. Then in a perception study the audio sound track alone, as well an audio plus video condition were presented to native Mandarin speakers who were required to decide which tone they perceived. Audio was presented in various conditions: clear, babble-noise masked at different SNR levels, as well as devoiced and amplitudemodulated noise conditions using LPC resynthesis. In the devoiced and the clear audio conditions, there is little augmentation of audio alone due to the addition of video. However, the addition of visual information did significantly improve perception in the babble-noise masked condition, and this effect increased with decreasing SNR. This outcome suggests that the improvement in noise-masked conditions is not due to additional information in the video per se, but rather to an effect of early integration of acoustic and visual cues facilitating auditory-visual speech perception.
Hansjörg Mixdorff, Yu Hu 0003, Denis Burnham
INTERSPEECH2
2004 Polynomial regression model for duration prediction in Mandarin
Yu Hu 0003, Renhua Wang
INTERSPEECH1
2004 Compression of speech database by feature separation and pattern clustering using STRAIGHT
abstract
This paper presents an alternative solution for speech database compression aiming at the embedded application of concatenative synthesis systems. The waveform of a speech segment is firstly decomposed into a prosodic pattern and a spectral pattern by STRAIGHT – a powerful speech analysissynthesis algorithm. Then all the prosodic and spectral patterns are clustered respectively to remove the redundant acoustic information within database. The clustering process is controllable and can export flexible compression ratio to meet the actual footprint requirement of various embedded devices. Besides, some labeling and contextual information are utilized to improve the performance of pattern clustering. Subjective listening test shows that our Mandarin synthesis system with corpus compressed by proposed method at about 2.7kbps perform corresponding to the same system compressed by G.723.1 at 5.3kps and the quality degradation is not serious as the compression ratio increases.
Zhen-Hua Ling, Yu Hu 0003, Zhiwei Shuang, Renhua Wang
INTERSPEECH2
2003 Towards the automatic extraction of fujisaki model parameters for Mandarin
abstract
The generation of naturally-sounding F0 contours in TTS enhances the intelligibility and perceived naturalness of synthetic speech. In earlier works the first author developed a linguistically motivated model of German intonation based on the quantitative Fujisaki model of the production process of F0, and an automatic procedure for extracting the parameters from the F0 contour which, however, was specific to German. As has been shown by Fujisaki and his co-workers, parametrization of F0 contours of Mandarin requires negative tone commands, as well as a more precise control of F0 associated with the syllabic tones. This paper presents an approach to the automatic parameter estimation for Mandarin, as well as first results concerning the accuracy of estimation. The paper also introduces a recently developed tool for editing Fujisaki parameters featuring resynthesis which will soon be publicly available.
Hansjörg Mixdorff, Hiroya Fujisaki, Gao Peng Chen, Yu Hu 0003
INTERSPEECH4
2002 A miniature Chinese TTS system based on tailored corpus
Zhiwei Shuang, Yu Hu 0003, Zhen-Hua Ling, Renhua Wang
INTERSPEECH2
2002 A new method of building decision tree based on target information
Yi-Jian Wu, Yu Hu 0003, Xiaoru Wu, Renhua Wang
INTERSPEECH2
2000 KD2000 Chinese Text-To-Speech System
Renhua Wang, Qingfeng Liu, Yu Hu 0003, Xiaoru Wu
ICMI3
2000 Prosody generation in Chinese synthesis using the template of quantified prosodic unit and base intonation contour
Yu Hu 0003, Qingfeng Liu, Renhua Wang
INTERSPEECH1