Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jung-Kuei Chen

dblp:85/9346 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 1995
0000-0003-4964-2031ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorArtificial intelligence and machine learning · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Speech recognition and synthesis · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.011994
An N-best candidates-based discriminative training for speech recognition applications · IEEE Trans. Speech Audio Process. 1994
Natural language and speech › Speech recognition and synthesis › acoustic modeling
discriminative acoustic model training
0.011994
An N-best candidates-based discriminative training for speech recognition applications · IEEE Trans. Speech Audio Process. 1994

Methods — techniques the papers use, named apart from their topics

tree-trellis search · 0.0n-best decoding · 0.0hidden markov model · 0.0gradient descent · 0.0frame-level loss function · 0.0
YearPublicationVenuePosition
1995 Large vocabulary, word-based Mandarin dictation system
Jung-Kuei Chen, Lin-Shan Lee, Frank K. Soong
EUROSPEECH1
1994 Discriminative training of high performance speech recognizer using N best candidates
abstract
Proposes an N-best candidates based, discriminative training procedure for constructing high performance HMM speech recognizers. The algorithm has two features: (1) a new frame-level loss function; (2) N best candidates are used for training. The new frame-level loss function, defined as a rectified log likelihood difference between the correct and other competing hypotheses, is minimized over all training utterances. Two speech recognition applications have been tested: speaker independent, small vocabulary (10 Mandarin Chinese digits), continuous speech recognition; and a speaker-trained, large vocabulary (5,000 commonly used Chinese words), isolated word recognition. Significant performance improvement over the traditional maximum likelihood trained HMMs has been obtained. In the connected Chinese digit recognition experiment, the string error rate is reduced from 17% to 10.8% for unknown length decoding and from 8.2% to 5.2% for known length decoding. In the large vocabulary, isolated word recognition experiment, the recognition error rate is improved from 6.8% to 3.8%.>
Jung-Kuei Chen, Frank K. Soong
ICASSP (1)1
1994 Large vocabulary word recognition based on tree-trellis search
abstract
In this paper we propose a large vocabulary (90000 words), Chinese (Mandarin) word recognizer based on the tree-trellis fast search algorithm. The recognizer is divided into 3 modules: local likelihood computation, a forward trellis search and a backward tree search. In the forward trellis search, a free syllable decoding is performed without a language model and a partial path map is created. The best-first tree search is then applied backward along a lexicon, which is arranged as a syllabic tree, to find the N-best word candidates. In the experiment, context-dependent subsyllabic HMMs were trained with a new discriminative training method. When it is evaluated on a speaker-trained database, the recognizer achieved a word error rate of 5% for the full size (90000 words) vocabulary and 1.7% for a smaller subset (5000 words) vocabulary. A real-time demo system has also been implemented on an SGI R-4000 workstation.>
Jung-Kuei Chen, Frank K. Soong, Lin-Shan Lee
ICASSP (2)1
1994 An N-best candidates-based discriminative training for speech recognition applications
abstract
The authors propose an N-best candidates-based discriminative training procedure for constructing high-performance HMM speech recognizers. The algorithm has two distinct features: N-best hypotheses are used for training discriminative models; and a new frame-level loss function is minimized to improve the separation between the correct and incorrect hypotheses. The N-best candidates are decoded based on their recently proposed tree-trellis fast search algorithm. The new frame-level loss function, which is defined as a halfwave rectified log-likelihood difference between the correct and competing hypotheses, is minimized over all training tokens. The minimization is carried out by adjusting the HMM parameters along a gradient descent direction. Two speech recognition applications have been tested, including a speaker independent, small vocabulary (ten Mandarin Chinese digits), continuous speech recognition, and a speaker-trained, large vocabulary (5000 commonly used Chinese words), isolated word recognition. Significant performance improvement over the traditional maximum likelihood trained HMMs has been obtained. In the connected Chinese digit recognition experiment, the string error rate is reduced from 17.0 to 10.8% for unknown length decoding and from 8.2 to 5.2% for known length decoding. In the large vocabulary, isolated word recognition experiment, the recognition error rate is reduced from 7.2 to 3.8%. Additionally, they have found that using more relaxed decoding constraints in preparing N-best hypotheses yields better recognition results.>
Jung-Kuei Chen, Frank K. Soong
IEEE Trans. Speech Audio Process.1
1990 Discriminant methods for improving the robustness of Mandarin syllables recognition based upon hidden Markov model
abstract
Two types of hidden Markov modeling (HMM), discrete-type (vector quantization (VQ)-based) and mixture density, are used to overcome part of the problem of Mandarin syllable recognition. In the VQ-based HMM part, a two-stage process is used to generate a VQ codebook in the hope of avoiding the possible domination of the vowel signal vectors on the codewords in vector space. A polyobservation sequence approach is adopted to minimize the quantization error. In the mixture density approach, the context-dependent models are incorporated in the classification of different buffers in the memory. The 408 syllables are divided into 154 buffers, 115 for consonants and 39 for vowels. In this way a buffer with tied probabilities may be shared by several different syllables. Hence, the bias between the confusing syllables with the same vowel part is expected to be eliminated by tied probability techniques. Experimental results show that improvement can be achieved with both types of HMM.>
Fu-hua Liu, Jung-Kuei Chen, Eng-Fong Huang, Chi-Shi Liu
ICASSP2