Jinsong Zhang 0001

dblp:56/728-1 · DBLP profile ↗
← Back
52ranked-venue papers
12as first author
12since 2021 · last 2023
0000-0002-1603-3136ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 12 first-author · 9 since 2021Artificial intelligence and machine learning · 43 · 6 first-author · 11 since 2021
YearPublicationVenuePosition
2023 LIMI-VC: A Light Weight Voice Conversion Model with Mutual Information Disentanglement
abstract
Voice conversion(VC) model aims to convert the source timbre to the target one. Recently, many VC models utilize pre-trained models to enhance the performance and achieve good results. However, pre-trained models could not somehow disentangle the timbre and linguistic information, thus resulting in a redundancy, which may hurt the conversion performance. In this paper we proposed LIMI-VC, reducing the redundancy between the linguistic content and the timbre information with mutual information disentanglement. We design the model in a light weight form, for the sake of parameter and computation efficiency when pre-trained models are commonly used nowadays. Experiments show that the proposed model can still improve the performance, with 15 times smaller size, compared to baseline. An out-of-domain cross-lingual inference also shows that our model greatly outperforms the baseline. Our source code and audio examples will be available at: https://github.com/WongLaw/LIMI-VC.
Liangjie Huang, Yunming Liang, Can Wen, Yanlu Xie, Jinsong Zhang 0001, Dengfeng Ke
ICASSP7
2023 The effect of stress on Mandarin tonal perception in continuous speech for Spanish-speaking learners
Lixia Hao, Jinsong Zhang 0001
INTERSPEECH3
2023 Dual Audio Encoders Based Mandarin Prosodic Boundary Prediction by Using Multi-Granularity Prosodic Representations
Ruishan Li, Yingming Gao, Yanlu Xie, Dengfeng Ke, Jinsong Zhang 0001
INTERSPEECH5
2022 A study of production error analysis for Mandarin-speaking Children with Hearing Impairment
Jingwen Cheng, Yingming Gao, Xiaoli Feng, Yannan Wang, Jinsong Zhang 0001
INTERSPEECH6
2022 A VR Interactive 3D Mandarin Pronunciation Teaching Model
Yujia Jin, Yanlu Xie, Jinsong Zhang 0001
INTERSPEECH3
2022 Self-Supervised Learning with Multi-Target Contrastive Coding for Non-Native Acoustic Modeling of Mispronunciation Verification
Longfei Yang, Jinsong Zhang 0001, Takahiro Shinozaki
INTERSPEECH2
2022 Modeling Unsupervised Empirical Adaptation by DPGMM and DPGMM-RNN Hybrid Model to Extract Perceptual Features for Low-Resource ASR
abstract
Speech feature extraction is critical for ASR systems. Such successful features as MFCC and PLP use filterbank techniques to model log-scaled speech perception but fail to model the adaptation of human speech perception by hearing experiences. Infant perception that is adapted by hearing speech without text may cause permanent brain state modifications (engrams) that serve as a physical fundamental basis for lifetime speech perception formation. This realization motivates us to propose to model such an unsupervised adaptation process, where adaptation denotes perception that is affected or changed by the history of experiences, with the Dirichlet Process Gaussian Mixture Model (DPGMM) and the DPGMM-RNN hybrid model to extract perceptual features to improve ASR. Our proposed features extend MFCC features with posteriorgrams extracted from the DPGMM algorithm or the DPGMM-RNN hybrid model. Our analysis shows that the DPGMM and DPGMM-RNN model perplexities agree with infant auditory perplexity to support that the proposed features are perceptual. Our ASR results verify the effectiveness of the proposed unsupervised features in such tasks as LVCSR on WSJ and ASR on noisy low-resource telephone conversations, compared with the supervised bottleneck features from Kaldi in ASR performance.
Sakriani Sakti, Jinsong Zhang 0001, Satoshi Nakamura 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 A Study on Fine-Tuning wav2vec2.0 Model for the Task of Mispronunciation Detection and Diagnosis
Linkai Peng, Kaiqi Fu, Binghuai Lin, Dengfeng Ke, Jinsong Zhang 0001
Interspeech5
2021 A Preliminary Study on Discourse Prosody Encoding in L1 and L2 English Spontaneous Narratives
Yuqing Zhang 0003, Binghuai Lin, Jinsong Zhang 0001
Interspeech4
2021 Relationships Between Perceptual Distinctiveness, Articulatory Complexity and Functional Load in Speech Communication
Yuqing Zhang 0003, Yanlu Xie, Binghuai Lin, Jinsong Zhang 0001
Interspeech6
2021 Non-native acoustic modeling for mispronunciation verification based on language adversarial representation learning
Longfei Yang, Kaiqi Fu, Jinsong Zhang 0001, Takahiro Shinozaki
Neural Networks3
2021 Tackling Perception Bias in Unsupervised Phoneme Discovery Using DPGMM-RNN Hybrid Model and Functional Load
abstract
The human perception of phonemes is biased against speech sounds. The lack of correspondence between perceptual phonemes and acoustic signals forms a big challenge in designing unsupervised algorithms to distinguish phonemes from sound. We propose the DPGMM-RNN hybrid model that improves phoneme categorization by relieving the fragmentation problem. We also merge segments with low functional load, which is the work done by segment contrasts to differentiate between utterances, just like humans who convert unambiguous segments into phonemes as units for immediate perception. Our results show that the DPGMM-RNN hybrid model relieves the fragmentation problem and improves phoneme discriminability. The minimal functional load merge compresses a segment system, preserves information and keeps phoneme discriminability.
Sakriani Sakti, Jinsong Zhang 0001, Satoshi Nakamura 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Formant Tracking Using Dilated Convolutional Networks Through Dense Connection with Gating Mechanism
abstract
Formant tracking is one of the most fundamental problems in speech processing.Traditionally, formants are estimated using signal processing methods.Recent studies showed that generic convolutional architectures can outperform recurrent networks on temporal tasks such as speech synthesis and machine translation.In this paper, we explored the use of Temporal Convolutional Network (TCN) for formant tracking.In addition to the conventional implementation, we modified the architecture from three aspects.First, we turned off the "causal" mode of dilated convolution, making the dilated convolution see the future speech frames.Second, each hidden layer reused the output information from all the previous layers through dense connection.Third, we also adopted a gating mechanism to alleviate the problem of gradient disappearance by selectively forgetting unimportant information.The model was validated on the open access formant database VTR.The experiment showed that our proposed model was easy to converge and achieved an overall mean absolute percent error (MAPE) of 8.2% on speech-labeled frames, compared to three competitive baselines of 9.4% (LSTM), 9.1% (Bi-LSTM) and 8.9% (TCN).
Wang Dai, Jinsong Zhang 0001, Yingming Gao, Dengfeng Ke, Binghuai Lin, Yanlu Xie
INTERSPEECH2
2020 Perception and Production of Mandarin Initial Stops by Native Urdu Speakers
Dan Du, Xianjin Zhu, Jinsong Zhang 0001
INTERSPEECH4
2020 An Investigation of the Target Approximation Model for Tone Modeling and Recognition in Continuous Mandarin Speech
abstract
The complex f0 variations in continuous speech make it rather difficult to perform automatic recognition of tones in a language like Mandarin Chinese.In this study, we tested the use of target approximation model (TAM) for continuous tone recognition on two datasets.TAM simulates f0 production from the articulatory point of view and so allow to discover the underlying pitch targets from the surface f0 contour.The f0 contour of each tone represented by 30 equidistant points in the first dataset was simulated by the TAM model.Using a support vector machine (SVM) to classify tones showed that, compared to the representation by 30 f0 values, the estimated three-dimensional TAM parameters had a comparable performance in characterizing tone patterns.The TAM model was further tested on the second dataset containing more complex tonal variations.With equal or a fewer number of features, the TAM parameters provided better performance than the coefficients of the cosine transform and a slightly worse performance than the statistical f0 parameters for tone recognition.Furthermore, we investigated bidirectional LSTM neural network for modelling the sequential tonal variations, which proved to be more powerful than the SVM classifier.The BLSTM system incorporating TAM and statistical f0 parameters achieved the best accuracy of 87.56%.
Yingming Gao, Jinsong Zhang 0001, Peter Birkholz
INTERSPEECH4
2020 Automatic Scoring at Multi-Granularity for L2 Pronunciation
Binghuai Lin, Xiaoli Feng, Jinsong Zhang 0001
INTERSPEECH4
2020 Joint Detection of Sentence Stress and Phrase Boundary for Prosody
Binghuai Lin, Xiaoli Feng, Jinsong Zhang 0001
INTERSPEECH4
2020 A Mandarin L2 Learning APP with Mispronunciation Detection and Feedback
Yanlu Xie, Xiaoli Feng, Boxue Li, Jinsong Zhang 0001, Yujia Jin
INTERSPEECH4
2020 Pronunciation Erroneous Tendency Detection with Language Adversarial Represent Learning
Longfei Yang, Kaiqi Fu, Jinsong Zhang 0001, Takahiro Shinozaki
INTERSPEECH3
2019 The Production of Chinese Affricates /ts/ and /tsh/ by Native Urdu Speakers
Dan Du, Jinsong Zhang 0001
INTERSPEECH2
2019 Capturing L1 Influence on L2 Pronunciation by Simulating Perceptual Space Using Acoustic Features
Shuju Shi, Chilin Shih, Jinsong Zhang 0001
INTERSPEECH3
2018 Interactions between Vowels and Nasal Codas in Mandarin Speakers' Perception of Nasal Finals
Yanlu Xie, Jinsong Zhang 0001
INTERSPEECH5
2018 A Preliminary Study on Tonal Coarticulation in Continuous Speech
Lixia Hao, Wei Zhang 0190, Yanlu Xie, Jinsong Zhang 0001
INTERSPEECH4
2018 Analysis of L2 Learners' Progress of Distinguishing Mandarin Tone 2 and Tone 3
Win Thuzar Kyaw, Jinsong Zhang 0001, Yoshinori Sagisaka
INTERSPEECH3
2018 Improving Mandarin Tone Recognition Using Convolutional Bidirectional Long Short-Term Memory with Attention
Longfei Yang, Yanlu Xie, Jinsong Zhang 0001
INTERSPEECH3
2018 Emotional Prosody Perception in Mandarin-speaking Congenital Amusics
Tianzhu Geng, Jinsong Zhang 0001
INTERSPEECH3
2017 Effective articulatory modeling for pronunciation error detection of L2 learner without non-native training data
abstract
For effective articulatory feedback in computer-assisted pronunciation training (CAPT) systems, we address effective articulatory models of second language (L2) learners' speech without using such data, which is difficult to collect and annotate in a large scale. Context-dependent articulatory attributes (placement and manner of articulation) are modeled based on deep neural network (DNN). In order to efficiently train the non-native articulatory models, we exploit large speech corpora of native and target language to model inter-language phenomena. This multi-lingual learning is then combined with multi-task learning, which uses phone-classification as a sub-task. These methods are applied to Mandarin Chinese pronunciation learning by Japanese native speakers. Effects are confirmed in the native attribute classification and pronunciation error detection of non-native speech.
Richeng Duan, Tatsuya Kawahara, Masatake Dantsuji, Jinsong Zhang 0001
ICASSP4
2017 The Influence on Realization and Perception of Lexical Tones from Affricate's Aspiration
Yanlu Xie, Jinsong Zhang 0001
INTERSPEECH4
2017 Reanalyze Fundamental Frequency Peak Delay in Mandarin
Lixia Hao, Wei Zhang 0190, Yanlu Xie, Jinsong Zhang 0001
INTERSPEECH4
2016 Landmark of Mandarin nasal codas and its application in pronunciation error detection
abstract
L2 learners of Mandarin have difficulty learning native-like pronunciation of nasal codas. In order to help them learn native-like pronunciation, we propose to develop targeted classifiers for automatic pronunciation error detection. In this paper, perceptual experiments with modified speech are designed to analyze the exact position of the landmark of a nasal coda. Based on perceptual results from isolated words, we propose that information about nasal coda place of articulation is most dense near a landmark at the center of the nasalized vowel. Landmarks detected in a database of Japanese learners of Mandarin, and classified as correct vs. incorrect using an SVM. The result shows that the detection performance of the SVM+Landmark system is similar to that of a DNN-HMM+MFCC system. When the two systems are combined, an FRR of 4.6% is achieved at DA of 83.9%. This performance is comparable to that of previously developed classifiers for 16 common Mandarin pronunciation errors.
Yanlu Xie, Mark Hasegawa-Johnson, Leyuan Qu, Jinsong Zhang 0001
ICASSP4
2016 Automatic Pronunciation Evaluation of Non-Native Mandarin Tone by Using Multi-Level Confidence Measures
Ju Lin, Yanlu Xie, Jinsong Zhang 0001
INTERSPEECH3
2016 Analysis of Chinese Syllable Durations in Running Speech of Japanese L2 Learners
Shudon Hsiao, Yoshinori Sagisaka, Jinsong Zhang 0001
INTERSPEECH4
2015 A study on robust detection of pronunciation erroneous tendency based on deep neural network
Yingming Gao, Yanlu Xie, Jinsong Zhang 0001
INTERSPEECH4
2014 A preliminary study on ASR-based detection of Chinese mispronunciation by Japanese learners
Richeng Duan, Jinsong Zhang 0001, Yanlu Xie
INTERSPEECH2
2014 A preliminary study on acoustic correlates of tone2+tone2 disyllabic word stress in Mandarin
Shuju Shi, Jinsong Zhang 0001
INTERSPEECH3
2014 Phoneme Set Design Using English Speech Database by Japanese for Dialogue-Based English CALL Systems
Xiaoyun Wang 0002, Jinsong Zhang 0001, Masafumi Nishida, Seiichi Yamamoto
LREC2
2010 Developing a Chinese L2 speech database of Japanese learners with narrow-phonetic labels for computer assisted pronunciation training
Dongning Wang, Jinsong Zhang 0001, Ziyu Xiong
INTERSPEECH3
2006 Automatic Derivation of a Phoneme Set with Tone Information for Chinese Speech Recognition Based on Mutual Information Criterion
abstract
An appropriate approach to model tone information is helpful for building Chinese large vocabulary continuous speech recognition system. We propose to derive an efficient phoneme set of tone-dependent sub-word units to build a recognition system, by iteratively merging a pair of tone-dependent units according to the principle of minimal loss of the mutual information. The mutual information is measured between the word tokens and their phoneme transcriptions in a training text corpus, based on the system lexical and language model. The approach has the capability to keep discriminative tonal (and phoneme) contrasts that are most helpful for disambiguating homophone words due to lack of tones, and merge those tonal (and phoneme) contrasts that are not important for word disambiguation for the recognition task. This enable a flexible selection of phoneme set according to a balance between the MI information amount and the number of phonemes. We applied the method to traditional phoneme set of Initial/Finals, and derived several phoneme sets with different number of units. Speech recognition experiments using the derived sets showed their effectiveness.
Jinsong Zhang 0001, Xinhui Hu, Satoshi Nakamura 0001
ICASSP (1)1
2006 The ATR multilingual speech-to-speech translation system
abstract
In this paper, we describe the ATR multilingual speech-to-speech translation (S2ST) system, which is mainly focused on translation between English and Asian languages (Japanese and Chinese). There are three main modules of our S2ST system: large-vocabulary continuous speech recognition, machine text-to-text (T2T) translation, and text-to-speech synthesis. All of them are multilingual and are designed using state-of-the-art technologies developed at ATR. A corpus-based statistical machine learning framework forms the basis of our system design. We use a parallel multilingual database consisting of over 600 000 sentences that cover a broad range of travel-related conversations. Recent evaluation of the overall system showed that speech-to-speech translation quality is high, being at the level of a person having a Test of English for International Communication (TOEIC) score of 750 out of the perfect score of 990.
Satoshi Nakamura 0001, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, Jinsong Zhang 0001, Hirofumi Yamamoto, Eiichiro Sumita, Seiichi Yamamoto
IEEE Trans. Speech Audio Process.7
2005 Tone nucleus-based multi-level robust acoustic tonal modeling of sentential F0 variations for Chinese continuous speech tone recognition
Jinsong Zhang 0001, Satoshi Nakamura 0001, Keikichi Hirose
Speech Commun.1
2004 A study on robust segmentation and location of tone nuclei in Chinese continuous speech
abstract
Tone nuclei in continuous speech are regarded as efficient targets for either tone recognition or intonation function decomposition. The paper presents our statistically robust method to segment and locate tone nuclei in continuous speech. The method includes: an iterative segmental K-means segmentation of the tonal F0 contours, which is further aided with t-test based segment amalgamation; a linear discriminant function based tone nucleus discriminator, whose features are selected by the sequential feature selection method. The developed system achieved 97.5% correct tone nuclei on a speaker dependent task. The tone recognizer based on the detected tone nuclei improved tone recognition rate by over 6% more than the baseline ones using the full tonal syllable features.
Jinsong Zhang 0001, Keikichi Hirose
ICASSP (1)1
2004 Efficient tone classification of speaker independent continuous Chinese speech using anchoring based discriminating features
abstract
Anchoring based discriminating features were proposed efficient for tone discrimination of Chinese continuous speech, and have been successfully applied before to tone classification of speaker dependent experiment. This paper presents its application to speaker independent tone classification experiments. Furthermore, we made detailed comparison experiments on the efficiencies of three groups of features: the left context dependent, the right context dependent anchoring F0 features, and the conventional F0 features. Experimental results showed that a combination of all three groups achieved a significant improvement of absolute 6.4% from 82.6% by the baseline system to 89.0%. When the three groups of features are used individually, both groups of the anchoring features led to better results than the conventional features, and the left context dependent anchoring features led to the highest performance.
Jinsong Zhang 0001, Satoshi Nakamura 0001, Keikichi Hirose
INTERSPEECH1
2004 Tone nucleus modeling for Chinese lexical tone recognition
Jinsong Zhang 0001, Keikichi Hirose
Speech Commun.1
2003 A multilevel framework to model the inherently confounding nature of sentential F0sentential F0 contours contours for recognizing Chinese lexical tones
abstract
This paper presents a multilevel framework to cope with the complex variations in Chinese sentential F0 contours in order to recognize lexical tones. Tone nucleus model is to get rid of the influence of intrinsic F0 transition loci at sub-syllable level. The pitch anchoring concept is used to normalize tonal F0 contours at syllable level. The hypo- and hyper-intonation model is used to account for the interplay of tone coarticulation and higher level prosodic effects. The whole approach achieved significant higher performance than the conventional method.
Jinsong Zhang 0001, Keikichi Hirose, Satoshi Nakamura 0001
ICASSP (1)1
2002 Weighted graph based decision tree optimization for high accuracy acoustic modeling
Jinsong Zhang 0001, Satoshi Nakamura 0001, Chin-Hui Lee 0001, Tat-Seng Chua
INTERSPEECH2
2002 Modeling varying pauses to develop robust acoustic models for recognizing noisy conversational speech
Jinsong Zhang 0001, Satoshi Nakamura 0001
INTERSPEECH1
2001 A hybrid approach to enhance task portability of acoustic models in Chinese speech recognition
abstract
This paper presents our approach to enhance the portability of acoustic models by mitigating the phonetic mismatch arising from a new testing task which is rather different from the training data. The approach is a hybrid one which combines knowledge-based context categorization to generate a context rich set of subword units, and data-driven-based acoustic model clustering on the level of context category. Compared with the conventional approach of only phonetic decision tree based model clustering and unseen model generation, the new approach improved greatly the desired subword coverage for the new testing domain, and achieved an error rate reduction by 10.8% for Chinese character accuracy in the recognition experiments. Together with the effect of the newly adopted basic units of 9 glottal stops, we achieved a total 23.5% error rate reduction in the testing compared to the baseline system.
Jinsong Zhang 0001, Shuwu Zhang, Yoshinori Sagisaka, Satoshi Nakamura 0001
INTERSPEECH1
2000 Anchoring hypothesis and its application to tone recognition of Chinese continuous speech
abstract
We present in this paper a new Chinese lexical tone recognition approach based on our pitch anchoring hypothesis, which suggests that the tone offset of the preceding lexical tone and the tone onset of the succeeding lexical tone serve as anchor points for the pitch heights of the onset and offset of the sandwiched lexical tone. The new approach exploits relative F0 heights between neighboring tones as important discriminating features for the lexical tones. Experimental results revealed that the new approach could increase the tone recognition accuracy greatly: above 10% (from 75.3% to 85.5%) compared with the conventional one.
Jinsong Zhang 0001, Keikichi Hirose
ICASSP1
2000 Discriminating Chinese lexical tones by anchoring F0 features
Jinsong Zhang 0001, Satoshi Nakamura 0001, Keikichi Hirose
INTERSPEECH1
1999 Tone recognition of Chinese continuous speech using tone critical segments
Keikichi Hirose, Jinsong Zhang 0001
EUROSPEECH2
1998 A robust tone recognition method of Chinese based on sub-syllabic F0 contours
Jinsong Zhang 0001, Keikichi Hirose
ICSLP1
1996 Adaptive recognition method based on posterior use of distribution pattern of output probabilities
Jinsong Zhang 0001, Beiqian Dai, Changfu Wang, HingKeung Kwan, Keikichi Hirose
ICSLP1