VLDB 2026 Research / reviewers in the wild / expert
Xiao Chen 0012
dblp:05/3054-12
· DBLP profile ↗
20ranked-venue papers
0as first author
13since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | bert2BERT: Towards Reusable Pretrained Language ModelsabstractCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Yujia Qin, Zhi Wang 0001, Xiao Chen 0012, Zhiyuan Liu 0001, Qun Liu 0001 |
ACL (1) | 8 |
| 2022 | CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and DiagnosisabstractMispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems.End-to-end (e2e) approaches are becoming dominant in MDD.However an e2e MDD model usually requires entire speech utterances as input context, which leads to significant time latency especially for long paragraphs.We propose a streaming e2e MDD model called CoCA-MDD.We utilize conv-transformer structure to encode input speech in a streaming manner.A coupled cross-attention (CoCA) mechanism is proposed to integrate frame-level acoustic features with encoded reference linguistic features.CoCA also enables our model to perform mispronunciation classification with whole utterances.The proposed model allows system fusion between the streaming output and mispronunciation classification output for further performance enhancement.We evaluate CoCA-MDD on publicly available corpora.CoCA-MDD achieves F1 scores of 57.03% and 60.78% for streaming and fusion modes respectively on L2-ARCTIC.For phone-level pronunciation scoring, CoCA-MDD achieves 0.58 Pearson correlation coefficient (PCC) value on SpeechOcean762. Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung, Baohua Xu, Yasheng Wang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
INTERSPEECH | 8 |
| 2022 | A Principle Solution for Enroll-Test Mismatch in Speaker RecognitionabstractMismatch between enrollment and test conditions causes serious performance degradation on speaker recognition systems. This paper presents a statistics decomposition (SD) approach to solve this problem. This approach decomposes the PLDA score into three components that corresponding to enrollment, prediction and normalization respectively. Given that correct statistics are used in each component, the resultant score is theoretically optimal. A comprehensive experimental study was conducted on three datasets with different types of mismatch: (1) physical channel mismatch, (2) long-term speaker characteristics mismatch, (3) near-far recording mismatch. The results demonstrated that the proposed SD approach is highly effective, and outperforms the ad-hoc multi-condition training approach that is commonly adopted but not optimal in theory. Lantian Li, Dong Wang 0013, Jiawen Kang 0002, Renyu Wang, Zhendong Gao, Xiao Chen 0012 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2021 | GhostBERT: Generate More Features with Cheap Operations for BERTabstractZhiqi Huang, Lu Hou, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhiqi Huang 0001, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language ModelsabstractYichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | EditSpeech: A Text Based Speech Editing System Using Partial Inference and Bidirectional FusionabstractThis paper presents the design, implementation and evaluation of a speech editing system, named EditSpeech, which allows a user to perform deletion, insertion and replacement of words in a given speech utterance, without causing audible degradation in speech quality and naturalness. The EditSpeech system is developed upon a neural text-to-speech (NTTS) synthesis framework. Partial inference and bidirectional fusion are proposed to effectively incorporate the contextual information related to the edited region and achieve smooth transition at both left and right boundaries. Distortion introduced to the unmodified parts of the utterance is alleviated. The EditSpeech system is developed and evaluated on English and Chinese in multi-speaker scenarios. Objective and subjective evaluation demonstrate that EditSpeech outperforms a few baseline systems in terms of low spectral distortion and preferred speech quality. Audio samples are available online for demonstration11https://daxintan-cuhk.github.io/EditSpeech/. Daxin Tan, Liqun Deng, Yu Ting Yeung, Xin Jiang 0002, Xiao Chen 0012, Tan Lee |
ASRU | 5 |
| 2021 | DyLex: Incorporating Dynamic Lexicons into BERT for Sequence LabelingabstractBaojun Wang, Zhao Zhang, Kun Xu, Guang-Yuan Hao, Yuyang Zhang, Lifeng Shang, Linlin Li, Xiao Chen, Xin Jiang, Qun Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Baojun Wang, Guang-Yuan Hao, Lifeng Shang, Linlin Li 0001, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 8 |
| 2021 | Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
Yichun Yin, Lifeng Shang, Zhi Wang 0001, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ICANN (3) | 6 |
| 2021 | Fcl-Taco2: Towards Fast, Controllable and Lightweight Text-to-Speech SynthesisabstractSequence-to-sequence (seq2seq) learning has greatly improved text-to-speech (TTS) synthesis performance, but effective implementation on resource-restricted devices remains challenging as seq2seq models are usually computationally expensive and memory intensive. To achieve fast inference speed and small model size while maintain high-quality speech, we propose FCL-taco2, a Fast, Controllable and Lightweight (FCL) TTS model based on Tacotron2. FCL-taco2 adopts a novel semi-autoregressive (SAR) mode for phoneme level based parallel mel-spectrograms generation conditioned on prosody features, leading to faster inference speed and higher prosody controllability than Tacotron2. Besides, knowledge distillation (KD) is leveraged to compress a relatively large FCL-taco2 model to its small version with minor loss of speech quality. Experimental results on English (EN) and Chinese (CN) datasets show that the small version of FCL-taco2 achieves comparable performance with Tacotron2 in terms of speech quality, while it has a 4.8× smaller footprint with 17.7× and 18.5× faster inference speeds on average for EN and CN experiments respectively. Besides, execution on mobile devices shows that the proposed model can achieve faster than real-time speech synthesis. Our code and audio samples are released1. Disong Wang, Liqun Deng, Yang Zhang 0025, Nianzu Zheng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng |
ICASSP | 6 |
| 2021 | VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-Shot Voice ConversionabstractOne-shot voice conversion (VC), which performs conversion across arbitrary speakers with only a single target-speaker utterance for reference, can be effectively achieved by speech representation disentanglement.Existing work generally ignores the correlation between different speech representations during training, which causes leakage of content information into the speaker representation and thus degrades VC performance.To alleviate this issue, we employ vector quantization (VQ) for content encoding and introduce mutual information (MI) as the correlation metric during training, to achieve proper disentanglement of content, speaker and pitch representations, by reducing their inter-dependencies in an unsupervised manner.Experimental results reflect the superiority of the proposed method in learning effective disentangled speech representations for retaining source linguistic content and intonation variations, while capturing target speaker characteristics.In doing so, the proposed approach achieves higher speech naturalness and speaker similarity than current state-of-the-art one-shot VC systems.Our code, pre-trained models and demo are available at https://github.com/Wendison/VQMIVC. Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng |
Interspeech | 4 |
| 2021 | Unsupervised Domain Adaptation for Dysarthric Speech Detection via Domain Adversarial Training and Mutual Information MinimizationabstractDysarthric speech detection (DSD) systems aim to detect characteristics of the neuromotor disorder from speech.Such systems are particularly susceptible to domain mismatch where the training and testing data come from the source and target domains respectively, but the two domains may differ in terms of speech stimuli, disease etiology, etc.It is hard to acquire labelled data in the target domain, due to high costs of annotating sizeable datasets.This paper makes a first attempt to formulate cross-domain DSD as an unsupervised domain adaptation (UDA) problem.We use labelled source-domain data and unlabelled target-domain data, and propose a multi-task learning strategy, including dysarthria presence classification (DPC), domain adversarial training (DAT) and mutual information minimization (MIM), which aim to learn dysarthriadiscriminative and domain-invariant biomarker embeddings.Specifically, DPC helps biomarker embeddings capture critical indicators of dysarthria; DAT forces biomarker embeddings to be indistinguishable in source and target domains; and MIM further reduces the correlation between biomarker embeddings and domain-related cues.By treating the UASPEECH and TORGO corpora respectively as the source and target domains, experiments show that the incorporation of UDA attains absolute increases of 22.2% and 20.0% respectively in utterancelevel weighted average recall and speaker-level accuracy. Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng |
Interspeech | 4 |
| 2021 | Energy-Friendly Keyword Spotting System Using Add-Based Convolution
Wenchao Hu, Yu Ting Yeung, Xiao Chen 0012 |
Interspeech | 4 |
| 2021 | Improving task-agnostic BERT distillation with layer mapping search
Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Linlin Li 0001, Fang Wang 0001, Qun Liu 0001 |
Neurocomputing | 6 |
| 2020 | Dialog State Tracking with Reinforced Data AugmentationabstractNeural dialog state trackers are generally limited due to the lack of quantity and diversity of annotated training data. In this paper, we address this difficulty by proposing a reinforcement learning (RL) based framework for data augmentation that can generate high-quality data to improve the neural state tracker. Specifically, we introduce a novel contextual bandit generator to learn fine-grained augmentation policies that can generate new effective instances by choosing suitable replacements for specific context. Moreover, by alternately learning between the generator and the state tracker, we can keep refining the generative policies to generate more high-quality training data for neural state tracker. Experimental results on the WoZ and MultiWoZ (restaurant) datasets demonstrate that the proposed framework significantly improves the performance over the state-of-the-art models, especially with limited training data. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
AAAI | 4 |
| 2020 | TernaryBERT: Distillation-aware Ultra-low Bit BERTabstractTransformer-based pre-training models like BERT have achieved remarkable performance in many natural language processing tasks.However, these models are both computation and memory expensive, hindering their deployment to resource-constrained devices.In this work, we propose TernaryBERT, which ternarizes the weights in a fine-tuned BERT model.Specifically, we use both approximation-based and loss-aware ternarization methods and empirically investigate the ternarization granularity of different parts of BERT.Moreover, to reduce the accuracy degradation caused by the lower capacity of low bits, we leverage the knowledge distillation technique (Jiao et al., 2019) in the training process.Experiments on the GLUE benchmark and SQuAD show that our proposed TernaryBERT outperforms the other BERT quantization methods, and even achieves comparable performance as the fullprecision model while being 14.9x smaller. Wei Zhang 0196, Lu Hou 0002, Yichun Yin, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 5 |
| 2020 | A Robust Audio-Visual Speech Enhancement ModelabstractMost existing audio-visual speech enhancement (AVSE) methods work well in conditions with strong noise, however when applied to conditions with a medium SNR, serious performance degradations are often observed. These degradations can be partly attributed to the feature-fusion(early fusion etc.) architecture that tightly couples the audio information that is very strong and the visual information that is relatively weak. In this paper, we present a safe AVSE approach that can make the visual stream contribute to audio speech enhancment(ASE) safely in conditions of various SNRs by late fusion.The key novelty is two-fold: Firstly, we define power binary masks (PBMs) as a rough representation of speech signals. This rough representation admits the weakness of the visual information and so can be easily predicted from the visual stream. Secondly, we design a posterior augmentation architecture that integrate the visual-derived PBMs to the audio-derived masks via a gating network. By this architecture, the entire performance is lower-bounded by the audio-based component. Our experiments on the Grid dataset demonstrated that this new approach consistently outperforms the audio-based system in all noise conditions, confirming that it is a safe way to incorporate visual knowledge in speech enhancement. Wupeng Wang, Dong Wang 0013, Xiao Chen 0012, Fengyu Sun |
ICASSP | 4 |
| 2020 | An Investigation of Few-Shot Learning in Spoken Term Classificationabstract202402 bcch Yangbin Chen, Tom Ko, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qing Li 0001 |
INTERSPEECH | 4 |
| 2020 | Conv-Transformer Transducer: Low Latency, Low Frame Rate, Streamable End-to-End Speech RecognitionabstractTransformer has achieved competitive performance against state-of-the-art end-to-end models in automatic speech recognition (ASR), and requires significantly less training time than RNN-based models.The original Transformer, with encoderdecoder architecture, is only suitable for offline ASR.It relies on an attention mechanism to learn alignments, and encodes input audio bidirectionally.The high computation cost of Transformer decoding also limits its use in production streaming systems.To make Transformer suitable for streaming ASR, we explore Transducer framework as a streamable way to learn alignments.For audio encoding, we apply unidirectional Transformer with interleaved convolution layers.The interleaved convolution layers are used for modeling future context which is important to performance.To reduce computation cost, we gradually downsample acoustic input, also with the interleaved convolution layers.Moreover, we limit the length of history context in self-attention to maintain constant computation cost for each decoding step.We show that this architecture, named Conv-Transformer Transducer, achieves competitive performance on LibriSpeech dataset (3.6% WER on test-clean) without external language models.The performance is comparable to previously published streamable Transformer Transducer and strong hybrid streaming ASR systems, and is achieved with smaller look-ahead window (140 ms), fewer parameters and lower frame rate. Wenyong Huang, Wenchao Hu, Yu Ting Yeung, Xiao Chen 0012 |
INTERSPEECH | 4 |
| 2020 | DynaBERT: Dynamic BERT with Adaptive Width and DepthabstractThe pre-trained language models like BERT, though powerful in many natural language processing tasks, are both computation and memory expensive. To alleviate this problem, one approach is to compress them for specific tasks before deployment. However, recent works on BERT compression usually compress the large BERT model to a fixed smaller size, and can not fully satisfy the requirements of different edge devices with various hardware performances. In this paper, we propose a novel dynamic BERT model (abbreviated as DynaBERT), which can flexibly adjust the size and latency by selecting adaptive width and depth. The training process of DynaBERT includes first training a width-adaptive BERT and then allowing both adaptive width and depth, by distilling knowledge from the full-sized model to small sub-networks. Network rewiring is also used to keep the more important attention heads and neurons shared by more sub-networks. Comprehensive experiments under various efficiency constraints demonstrate that our proposed dynamic BERT (or RoBERTa) at its largest size has comparable performance as BERT-base (or RoBERTa-base), while at smaller widths and depths consistently outperforms existing BERT compression methods. Code is available at https://github.com/huawei-noah/Pretrained-Language-Model/tree/master/DynaBERT. Lu Hou 0002, Zhiqi Huang 0001, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
NeurIPS | 5 |
| 2019 | Modeling Semantic Compositionality with Sememe KnowledgeabstractSemantic compositionality (SC) refers to the phenomenon that the meaning of a complex linguistic unit can be composed of the meanings of its constituents.Most related works focus on using complicated compositionality functions to model SC while few works consider external knowledge in models.In this paper, we verify the effectiveness of sememes, the minimum semantic units of human languages, in modeling SC by a confirmatory experiment.Furthermore, we make the first attempt to incorporate sememe knowledge into SC models, and employ the sememeincorporated models in learning representations of multiword expressions, a typical task of SC.In experiments, we implement our models by incorporating knowledge from a famous sememe knowledge base HowNet and perform both intrinsic and extrinsic evaluations.Experimental results show that our models achieve significant performance boost as compared to the baseline methods without considering sememe knowledge.We further conduct quantitative analysis and case studies to demonstrate the effectiveness of applying sememe knowledge in modeling SC.All the code and data of this paper can be obtained on https: //github.com/thunlp/Sememe-SC. Fanchao Qi, Junjie Huang 0003, Chenghao Yang 0001, Zhiyuan Liu 0001, Xiao Chen 0012, Qun Liu 0001, Maosong Sun 0001 |
ACL (1) | 5 |