Ying Li 0125

dblp:22/1805-125 · DBLP profile ↗
← Back
7ranked-venue papers
7as first author
0since 2021 · last 2014
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Speech recognition and synthesis · 38% Language models and text generation · 38% Information extraction and text analysis · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › multilingual speech recognition
code-switching speech recognition
0.212014
Language Modeling with Functional Head Constraint for Code Switching Speech Recognition · EMNLP 2014
Natural language and speech › Language models and text generation › language modeling › language model architecture
structured language model
0.212014
Language Modeling with Functional Head Constraint for Code Switching Speech Recognition · EMNLP 2014
Natural language and speech › Information extraction and text analysis › syntactic parsing
lattice parsing
0.112014
Language Modeling with Functional Head Constraint for Code Switching Speech Recognition · EMNLP 2014
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.112014
Language Modeling with Functional Head Constraint for Code Switching Speech Recognition · EMNLP 2014

Methods — techniques the papers use, named apart from their topics

weighted finite-state transducer · 0.2partial parsing · 0.2lattice parsing · 0.2functional head constraint · 0.2
YearPublicationVenuePosition
2014 Language Modeling with Functional Head Constraint for Code Switching Speech Recognition
abstract
In this paper, we propose novel structured language modeling methods for code mixing speech recognition by incorporating a well-known syntactic constraint for switching code, namely the Functional Head Constraint (FHC).Code mixing data is not abundantly available for training language models.Our proposed methods successfully alleviate this core problem for code mixing speech recognition by using bilingual data to train a structured language model with syntactic constraint.Linguists and bilingual speakers found that code switch do not happen between the functional head and its complements.We propose to learn the code mixing language model from bilingual data with this constraint in a weighted finite state transducer (WFST) framework.The constrained code switch language model is obtained by first expanding the search network with a translation model, and then using parsing to restrict paths to those permissible under the constraint.We implement and compare two approacheslattice parsing enables a sequential coupling whereas partial parsing enables a tight coupling between parsing and filtering.We tested our system on a lecture speech dataset with 16% embedded second language, and on a lunch conversation dataset with 20% embedded language.Our language models with lattice parsing and partial parsing reduce word error rates from a baseline mixed language model by 3.8% and 3.9% in terms of word error rate relatively on the average on the first and second tasks respectively.It outperforms the interpolated language model by 3.7% and 5.6% in terms of word error rate relatively, and outperforms the adapted language model by 2.6% and 4.6% relatively.Our proposed approach avoids making early decisions on codeswitch boundaries and is therefore more robust.We address the code switch data scarcity challenge by using bilingual data with syntactic structure.
Ying Li 0125, Pascale Fung
EMNLP1
2014 Code switch language modeling with Functional Head Constraint
abstract
In this paper, we propose for the first time to incorporate the linguistically well-known Functional Head Constraint into a code switch language model for speech recognition. Under this constraint, code switch cannot occur between the functional head and its complements. The constrained code switching language model is obtained by first expanding the search network with a translation model; and then using lattice-based parsing to restrict paths to those permissible under the Functional Head Constraint. We tested our system on two tasks of code switch speech recognition, namely lecture speech recognition and lunch conversation recognition. Our system reduces word error rates (WER) from a baseline mixed language model by 3.72% relative in the first task; and by 5.85% in the second task. It reduces WER from an interpolated language model by 2.51% in the first task; and by 4.57% in the second task. All results are statistically significant. In addition, our method reduces WER for both the matrix language and the embedded language.
Ying Li 0125, Pascale Fung
ICASSP1
2013 Improved mixed language speech recognition using asymmetric acoustic model and language model with code-switch inversion constraints
abstract
We propose an integrated framework for large vocabulary continuous mixed language speech recognition that handles the accent effect in the bilingual acoustic model and the inversion constraint well known to linguists in the language model. Our asymmetric acoustic model with phone set extension improves upon previous work by striking a balance between data and phonetic knowledge. Our language model improves upon previous work by (1) using the inversion constraint to predict code switching points in the mixed language and (2) integrating a code-switch prediction model, a translation model and a reconstruction model together. This integration means that our language model avoids the pitfall of propagated error that could arise from decoupling these steps. Finally, a WFST-based decoder integrates the acoustic models, code-switch language model and a monolingual language model in the matrix language all together. Our system reduces word error rate by 1.88% on a lecture speech corpus and by 2.43% on a lunch conversation corpus, with statistical significance, over the conventional bilingual acoustic model and interpolated language model.
Ying Li 0125, Pascale Fung
ICASSP1
2013 Language modeling for mixed language speech recognition using weighted phrase extraction
abstract
To train a code switching language model for mixed language speech recognition, we propose to assign weights to the sentence pairs in the parallel text data. The code switching language model which is composed of the code switching boundary prediction model, code switching translation model and reconstruction model is incorporated with a language for mixed language speech recognition. The code switching translation model which is trained using selected subsets of the sentence pairs in the parallel text data allows the decoder to make the decision whether a phrase is in the matrix language or in the embedded language. Moreover, we propose a weighting procedure while training the code switching translation model. We evaluate our methods on Mandarin-English code switching lecture speech and lunch conversations. Our proposed method reduces word error rate by a statistically significant 1.74% on the lecture speech, and by 1.29% on the lunch conversation over the conventional interpolated language model. Copyright © 2013 ISCA.
Ying Li 0125, Pascale Fung
INTERSPEECH1
2012 Code-Switch Language Model with Inversion Constraints for Mixed Language Speech Recognition
Ying Li 0125, Pascale Fung
COLING1
2012 A Mandarin-English Code-Switching Corpus
Ying Li 0125, Pascale Fung
LREC1
2011 Asymmetric acoustic modeling of mixed language speech
abstract
We propose to improve speech recognition performance on speaker-independent, mixed language speech by asymmetric acoustic modeling. Mixed language is either inter-sentential code switching from the source matrix language to a foreign language or intra-sentential code mixing between the matrix language and embedded foreign words or phrases. In either case, the foreign phrases are pronounced by the matrix language speaker with varying degrees of accent. Our pro posed system using selective decision tree merging between a bilingual model and an accented embedded speech model outperforms previous approaches of either using a bilingual model with model retraining by 21.51%, or using adaptation by 15.88%. It outperforms all models on both code mixing and code switching cases. We successfully improved recognition on embedded foreign speech without degrading the performance on the matrix language speech.
Ying Li 0125, Pascale Fung, Yi Liu 0022
ICASSP1