VLDB 2026 Research / reviewers in the wild / expert
Ju Zhang 0001
dblp:39/3127-1
· DBLP profile ↗
13ranked-venue papers
1as first author
7since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Frequency Patterns of Individual Speaker Characteristics at Higher and Lower Spectral Ranges
Ju Zhang 0001, Yujie Chi, Kiyoshi Honda, Jianguo Wei |
INTERSPEECH | 2 |
| 2022 | Using Multiple Reference Audios and Style Embedding Constraints for Speech SynthesisabstractThe end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding must be manually selected during inference. Due to the fact that only the matched text and speech are used in the training process, using unmatched text and speech for inference would cause the model to synthesize speech with low content quality. In this study, we propose to mitigate these two problems by using multiple reference audios and style embedding constraints rather than using only the target audio. Multiple reference audios are automatically selected using the sentence similarity determined by Bidirectional Encoder Representations from Transformers (BERT). In addition, we use "target" style embedding from a pre-trained encoder as a constraint by considering the mutual information between the predicted and "target" style embedding. The experimental results show that the proposed model can improve the speech naturalness and content quality with multiple reference audios and can also outperform the baseline model in ABX preference tests of style similarity. Longbiao Wang, Zhen-Hua Ling, Ju Zhang 0001, Jianwu Dang 0001 |
ICASSP | 4 |
| 2022 | Vocal-Tract Area Functions with Articulatory Reality for Tract Opening
Ju Zhang 0001, Jianguo Wei, Kiyoshi Honda, Tatsuya Kitamura |
INTERSPEECH | 2 |
| 2021 | Learning Language and Speaker Information for Code-Switch Speech Synthesis with Limited DataabstractEnd-to-end speech synthesis demonstrates remarkable performance in monolingual speech, whereas code-switching (CS) speech synthesis remains a challenge owing to the sparsity of data and diverse syntactic structures across languages. Previous studies show that large mixed-lingual corpora are essential for effective learning text/language representations and target speaker information. In this study, we propose a method using three independent encoders (text, language, and speaker), which requires only a small amount of mixed-lingual data to realize the CS speech synthesis of Mandarin and English. Additionally, to distinguish between Mandarin and English, we investigate two text-representation methods: (1) the implicit method, which uses Pinyin and the CMU11http://www.speech.cs.cmu.edu/cgi-bin/cmudict dictionary to represent both languages; and (2) the explicit method, which uses language markers i.e., masks, to differentiate the languages. Through our proposed method, we can improve synthesized speech in terms of quality and speaker similarity using a small amount of mixed-lingual data. In addition, the experimental results demonstrate that the proposed method achieves performance improvement of 0.06 in terms of the mean opinion score and absolute improvement of 0.64% in terms of the character error rate compared to the baseline method. Mengxin Chai, Shaotong Guo, Longbiao Wang, Jianwu Dang 0001, Ju Zhang 0001 |
ASRU | 6 |
| 2021 | Improving Naturalness and Controllability of Sequence-to-Sequence Speech Synthesis by Learning Local Prosody RepresentationsabstractState-of-the-art neural text-to-speech (TTS) networks are trained with a large amount of speech data, which significantly improves the quality of synthetic speech compared with traditional approaches. However, the prosody and controllability of the generated speech is still insufficient, especially in tonal languages. Moreover, the generated prosody is solely defined by the input text, which does not allow for different styles for the same sentence or words. In this study, we extended Tacotron2 with a pitch prediction task to capture discrete pitch-related representations. Specifically, the learned pitch-related suprasegmental information is fed simultaneously with traditional character features into the decoder to generate final Mel spectrogram. Experiments show that the proposed method can improve the quality of the generated speech (mean opinion score of 4.37 vs. 4.22). Moreover, we demonstrated that we can easily achieve word-level pitch control during generation by changing local pitch-related representations before passing them to the decoder network. Longbiao Wang, Zhen-Hua Ling, Shaotong Guo, Ju Zhang 0001, Jianwu Dang 0001 |
ICASSP | 5 |
| 2021 | Exploring Effective Speech Representation via ASR for High-Quality End-to-End Multispeaker TTS
Longbiao Wang, Sheng Li 0010, Chenchen Ding, Ju Zhang 0001, Jianwu Dang 0001 |
ICONIP (6) | 6 |
| 2021 | TacoLPCNet: Fast and Stable TTS by Conditioning LPCNet on Mel Spectrogram Predictions
Longbiao Wang, Ju Zhang 0001, Shaotong Guo, Yuguang Wang 0003, Jianwu Dang 0001 |
Interspeech | 3 |
| 2020 | Investigation of Effectively Synthesizing Code-Switched Speech Using Highly Imbalanced Mix-Lingual Data
Shaotong Guo, Longbiao Wang, Sheng Li 0010, Ju Zhang 0001, Yuguang Wang 0003, Jianwu Dang 0001, Kiyoshi Honda |
ICONIP (1) | 4 |
| 2020 | Three-Dimensional Lip Motion Network for Text-Independent Speaker RecognitionabstractLip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a text-dependent context. However, 2D lip easily suffers from various face orientations. To this end, in this work, we present a novel end-to-end 3D lip motion Network (3LMNet) by utilizing the sentence-level 3D lip motion (S3DLM) to recognize speakers in both the text-independent and text-dependent contexts. A new regional feedback module (RFM) is proposed to obtain attentions in different lip regions. Besides, prior knowledge of lip motion is investigated to complement RFM, where landmark-level and frame-level features are merged to form a better feature representation. Moreover, we present two methods, i.e., coordinate transformation and face posture correction to pre-process the LSD-AV dataset, which contains 68 speakers and 146 sentences per speaker. The evaluation results on this dataset demonstrate that our proposed 3LMNet is superior to the baseline models, i.e., LSTM, VGG-16 and ResNet-34, and outperforms the state-of-the-art using 2D lip image as well as the 3D face. The code of this work is released at https://github.com/wutong18/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition. Jianrong Wang, Shanyu Wang, Mei Yu 0004, Qiang Fang 0003, Ju Zhang 0001, Li Liu 0036 |
ICPR | 6 |
| 2018 | Three-Dimensional Joint Geometric-Physiologic Feature for Lip-ReadingabstractLip-reading has been successfully demonstrated that it can improve the performance of automatic speech recognition system especially in the presence of acoustic noise. However, the information about lip movement is still insufficient as the lip features are obtained from discrete three-dimensional points and planar images. The internal mechanisms of lip movement are not described and reflected. In this paper, we employed a novel deepening technique, namely densely connected convolutional networks (DenseNets), to obtain visual representation from color images. In addition, a new 3D lip physiologic feature based on the position and structure of facial muscles was extracted to represent the similarity of the way people speak. The color image feature and 3D lip geometric-physiologic feature were coupled together in the last fully-connected layer of DenseNets. The experimental results show that DenseNets can handle spatial-temporal information of a whole image sequence and the lip feature integrating our proposed 3D geometric-physiological feature is sufficient to improve the recognition rate by as much as 3.91% (from 94.84%, with the color images only, to 98.75%). Jianguo Wei, Ju Zhang 0001, Mei Yu 0004, Jianrong Wang |
ICTAI | 3 |
| 2018 | Tooth visualization in vowel production MR images for three-dimensional vocal tract modeling
Ju Zhang 0001, Kiyoshi Honda, Jianguo Wei |
Speech Commun. | 1 |
| 2016 | Continuous ultrasound based tongue movement video synthesis from speechabstractThe movement of tongue plays an important role in pronunciation. Visualizing the movement of tongue can improve speech intelligibility and also helps learning a second language. However, hardly any research has been investigated for this topic. In this paper, a framework to synthesize continuous ultrasound tongue movement video from speech is presented. Two different mapping methods are introduced as the most important parts of the framework. The objective evaluation and subjective opinions show that the Gaussian Mixture Model (GMM) based method has a better result for synthesizing static image and Vector Quantization (VQ) based method produces more stable continuous video. Meanwhile, the participants of evaluation state that the results of both methods are visual-understandable. Jianrong Wang, Yalong Yang 0001, Jianguo Wei, Ju Zhang 0001 |
ICASSP | 4 |
| 2016 | Audio-visual speech recognition integrating 3D lip information obtained from the Kinect
Jianrong Wang, Ju Zhang 0001, Kiyoshi Honda, Jianguo Wei, Jianwu Dang 0001 |
Multim. Syst. | 2 |