VLDB 2026 Research / reviewers in the wild / expert
Qiang Fang 0003
dblp:17/3080-3
· DBLP profile ↗
22ranked-venue papers
5as first author
7since 2021 · last 2024
0000-0002-6548-980XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evaluation of an Improved Ultrasonic Imaging Helmet for Observing Articulatory DataabstractUltrasonic imaging is one of the most popular methods for tracking tongue motion. Imaging plane shift and contact variation are crucial factors affecting the consistency of the obtained ultrasonic images. To solve this issue, researchers proposed many different helmets. In this study, we propose an evaluation framework to quantitatively assess the helmet’s imaging plane shift and contact variation. The framework is applied to one helmet we designed. Compared to the baseline, the imaging plane shift using our helmet is less than 0.1°, and the contact variation is reduced by 68.5%, decreasing the mean difference of the extracted contours by 50.3% and increasing the contrast and sharpness of the obtained images by 9.0% and 28.4%. The proposed framework provides an approach for quantitatively proofing helmet structure with data accuracy and consistency. Jianguo Wei, Qiang Fang 0003, Xugang Lu |
ICASSP | 3 |
| 2024 | Multi-modal co-learning for silent speech recognition based on ultrasound tongue images
Jianguo Wei, Ruiteng Zhang, Qiang Fang 0003 |
Speech Commun. | 5 |
| 2023 | Two-Stream Joint-Training for Speaker Independent Acoustic-to-Articulatory InversionabstractAcoustic-to-articulatory inversion (AAI) aims to estimate the parameters of articulators from speech audio. There are two common challenges in AAI, which are the limited data and the unsatisfactory performance in speaker independent scenario. Most current works focus on extracting features directly from speech and ignoring the importance of phoneme information which may limit the performance of AAI. To this end, we propose a novel network called SPN that uses two different streams to carry out the AAI task. Firstly, to improve the performance of speaker-independent experiment, we propose a new phoneme stream network to estimate the articulatory parameters as the phoneme features. To the best of our knowledge, this is the first work that extracts the speaker-independent features from phonemes to improve the performance of AAI. Secondly, in order to better represent the speech information, we train a speech stream network to combine the local features and the global features. Compared with state-of-the-art (SOTA), the proposed method reduces 0.18mm on RMSE and increases 6.0% on Pearson correlation coefficient in the speaker-independent experiment. The code has been released at https://github.com/liujinyu123/AAINetwork-SPN. Jianrong Wang, Xuewei Li 0001, Mei Yu 0004, Jie Gao 0008, Qiang Fang 0003, Li Liu 0036 |
ICASSP | 6 |
| 2022 | Residual-Guided Personalized Speech Synthesis based on Face ImageabstractPrevious works derive personalized speech features by training the model on a large dataset composed of his/her audio sounds. It was reported that face information has a strong link with the speech sound. Thus in this work, we innovatively extract personalized speech features from human faces to synthesize personalized speech using neural vocoder. A Face-based Residual Personalized Speech Synthesis Model (FR-PSS) containing a speech encoder, a speech synthesizer and a face encoder is designed for PSS. In this model, by designing two speech priors, a residual-guided strategy is introduced to guide the face feature to approach the true speech feature in the training. Moreover, considering the error of feature’s absolute values and their directional bias, we formulate a novel tri-item loss function for face encoder. Experimental results show that the speech synthesized by our model is comparable to the personalized speech synthesized by training a large amount of audio data in previous works. Jianrong Wang, Xiaosheng Hu, Xuewei Li 0001, Qiang Fang 0003, Li Liu 0036 |
ICASSP | 5 |
| 2022 | MVNet: Memory Assistance and Vocal Reinforcement Network for Speech Enhancement
Jianrong Wang, Xuewei Li 0001, Mei Yu 0004, Qiang Fang 0003, Li Liu 0036 |
ICONIP (2) | 5 |
| 2021 | An Attention Self-Supervised Contrastive Learning Based Three-Stage Model for Hand Shape Feature Representation in Cued SpeechabstractCued Speech (CS) is a communication system for deaf people or hearing impaired people, in which a speaker uses it to aid a lipreader in phonetic level by clarifying potentially ambiguous mouth movements with hand shape and positions.Feature extraction of multi-modal CS is a key step in CS recognition.Recent supervised deep learning based methods suffer from noisy CS data annotations especially for hand shape modality.In this work, we first propose a self-supervised contrastive learning method to learn the feature representation of image without using labels.Secondly, a small amount of manually annotated CS data are used to fine-tune the first module.Thirdly, we present a module, which combines Bi-LSTM and self-attention networks to further learn sequential features with temporal and contextual information.Besides, to enlarge the volume and the diversity of the current limited CS datasets, we build a new British English dataset containing 5 native CS speakers.Evaluation results on both French and British English datasets show that our model achieves over 90% accuracy in hand shape recognition.Significant improvements of 8.75% (for French) and 10.09% (for British English) are achieved in CS phoneme recognition correctness compared with the state-of-the-art. Jianrong Wang, Nan Gu, Mei Yu 0004, Xuewei Li 0001, Qiang Fang 0003, Li Liu 0036 |
Interspeech | 5 |
| 2021 | Cross-Modal Knowledge Distillation Method for Automatic Cued Speech RecognitionabstractCued Speech (CS) is a visual communication system for the deaf or hearing impaired people. It combines lip movements with hand cues to obtain a complete phonetic repertoire. Current deep learning based methods on automatic CS recognition suffer from a common problem, which is the data scarcity. Until now, there are only two public single speaker datasets for French (238 sentences) and British English (97 sentences). In this work, we propose a cross-modal knowledge distillation method with teacher-student structure, which transfers audio speech information to CS to overcome the limited data problem. Firstly, we pretrain a teacher model for CS recognition with a large amount of open source audio speech data, and simultaneously pretrain the feature extractors for lips and hands using CS data. Then, we distill the knowledge from teacher model to the student model with frame-level and sequence-level distillation strategies. Importantly, for frame-level, we exploit multi-task learning to weigh losses automatically, to obtain the balance coefficient. Besides, we establish a five-speaker British English CS dataset for the first time. The proposed method is evaluated on French and British English CS datasets, showing superior CS recognition performance to the state-of-the-art (SOTA) by a large margin. Jianrong Wang, Ziyue Tang, Xuewei Li 0001, Mei Yu 0004, Qiang Fang 0003, Li Liu 0036 |
Interspeech | 5 |
| 2020 | Three-Dimensional Lip Motion Network for Text-Independent Speaker RecognitionabstractLip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a text-dependent context. However, 2D lip easily suffers from various face orientations. To this end, in this work, we present a novel end-to-end 3D lip motion Network (3LMNet) by utilizing the sentence-level 3D lip motion (S3DLM) to recognize speakers in both the text-independent and text-dependent contexts. A new regional feedback module (RFM) is proposed to obtain attentions in different lip regions. Besides, prior knowledge of lip motion is investigated to complement RFM, where landmark-level and frame-level features are merged to form a better feature representation. Moreover, we present two methods, i.e., coordinate transformation and face posture correction to pre-process the LSD-AV dataset, which contains 68 speakers and 146 sentences per speaker. The evaluation results on this dataset demonstrate that our proposed 3LMNet is superior to the baseline models, i.e., LSTM, VGG-16 and ResNet-34, and outperforms the state-of-the-art using 2D lip image as well as the 3D face. The code of this work is released at https://github.com/wutong18/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition. Jianrong Wang, Shanyu Wang, Mei Yu 0004, Qiang Fang 0003, Ju Zhang 0001, Li Liu 0036 |
ICPR | 5 |
| 2018 | A Nonlinear 3D Geometric Tongue ModelabstractThis study describes a nonlinear geometric tongue model based on MRI and Cone-beam CT (CBCT) data. Comparing with the conventional geometric tongue model, the proposed tongue model is controlled by several prototype vertices, and the relationship between tongue mesh vertices and prototype vertices are modeled with quadratic functions. The results indicate that: i) quadratic models do improve the reconstruction performance of tongue mesh, especially in the tongue root region; ii) the quadratic model which use the cross-prototype-vertex information achieves the best performance of tongue mesh reconstruction; iii) the reconstruction performance can be further improved if an extra prototype vertex TP in the tongue root region is taken into account, even if TP is estimated from the measured prototype vertices. Qiang Fang 0003, Hequn Li, Jianguo Wei, Jianrong Wang, Xiyu Wu |
ICASSP | 1 |
| 2018 | Tongue Segmentation with Geometrically Constrained Snake Model
Zhihua Su, Jianguo Wei, Qiang Fang 0003, Jianrong Wang, Kiyoshi Honda |
INTERSPEECH | 3 |
| 2018 | Study of articulators' contribution and compensation during speech by articulatory speech recognition
Jianguo Wei, Jingshu Zhang, Qiang Fang 0003, Wenhuan Lu, Kiyoshi Honda, Xugang Lu |
Multim. Tools Appl. | 4 |
| 2017 | Acoustic VR in the mouth: A real-time speech-driven visual tongue systemabstractWe propose an acoustic-VR system that converts acoustic signals of human language (Chinese) to realistic 3D tongue animation sequences in real time. It is known that directly capturing the 3D geometry of the tongue at a frame rate that matches the tongue's swift movement during the language production is challenging. This difficulty is handled by utilizing the electromagnetic articulography (EMA) sensor as the intermediate medium linking the acoustic data to the simulated virtual reality. We leverage Deep Neural Networks to train a model that maps the input acoustic signals to the positional information of pre-defined EMA sensors based on 1,108 utterances. Afterwards, we develop a novel reduced physics-based dynamics model for simulating the tongue's motion. Unlike the existing methods, our deformable model is nonlinear, volume-preserving, and accommodates collision between the tongue and the oral cavity (mostly with the jaw). The tongue's deformation could be highly localized which imposes extra difficulties for existing spectral model reduction methods. Alternatively, we adopt a spatial reduction method that allows an expressive subspace representation of the tongue's deformation. We systematically evaluate the simulated tongue shapes with real-world shapes acquired by MRI/CT. Our experiment demonstrates that the proposed system is able to deliver a realistic visual tongue animation corresponding to a user's speech signal. Ran Luo 0001, Qiang Fang 0003, Jianguo Wei, Wenhuan Lu, Weiwei Xu 0003, Yin Yang 0002 |
VR | 2 |
| 2016 | An Improved 3D Geometric Tongue Model
Qiang Fang 0003, Jianguo Wei, Jianrong Wang, Xiyu Wu |
INTERSPEECH | 1 |
| 2016 | Morphological normalization of vowel images for articulatory speech recognition
Jianguo Wei, Jingshu Zhang, Qiang Fang 0003, Wenhuan Lu |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Mapping ultrasound-based articulatory images and vowel sounds with a deep neural network framework
Jianguo Wei, Qiang Fang 0003, Xinyuan Zheng, Wenhuan Lu, Jianwu Dang 0001 |
Multim. Tools Appl. | 2 |
| 2016 | Multi-modal recording and modeling of vocal tract movements
Jianguo Wei, Song Wang 0005, Wenhuan Lu, Darcy Qingzhi Hou, Qiang Fang 0003, Jianwu Dang 0001 |
Multim. Tools Appl. | 5 |
| 2015 | Combined cine- and tagged-MRI for tracking landmarks on the tongue surface
Honghao Bao, Wenhuan Lu, Kiyoshi Honda, Jianguo Wei, Qiang Fang 0003, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2015 | Phonetic-phonological feature emerges by associating phonetic with semantic information - a GSOM-based modeling study
Mengxue Cao, Qiang Fang 0003, Bernd J. Kröger |
INTERSPEECH | 3 |
| 2014 | Reconstruction of mistracked articulatory trajectories
Qiang Fang 0003, Jianguo Wei |
INTERSPEECH | 1 |
| 2013 | An anisotropic diffusion filter based on multidirectional separability
Jianguo Wei, Xin Wang 0037, Wenhuan Lu, Qiang Fang 0003, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2009 | Feedforward control of a 3d physiological articulatory model for vowel production
Qiang Fang 0003, Akikazu Nishikido, Jianwu Dang 0001 |
INTERSPEECH | 1 |
| 2008 | A model based investigation of activation patterns of the tongue muscles for vowel production
Qiang Fang 0003, Satoru Fujita, Xugang Lu, Jianwu Dang 0001 |
INTERSPEECH | 1 |