Xiangang Li

dblp:124/9046 · DBLP profile ↗
← Back
43ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-7810-1077ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
abstract
The Speaker Diarization and Recognition (SDR) task aims to predict ``who spoke when and what'' within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers.
Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang 0030, Chao-Hong Tan, Qian Chen 0003, Wen Wang 0001, Xiangang Li
AAAI9
2026 UniVocal: Unified Speech-Singing Code-Switching Synthesis
abstract
We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis-a task where transitions are autonomously driven by textual semantics, akin to seamless human language blending.Unlike single-mode generation or systems relying on switching-control tags, our proposed UniVocal implicitly infers vocal modes solely from text context.To achieve this, we employ a data-efficient two-stage curriculum learning strategy that progressively trains a competitive TTS system to acquire the desired SCS capability.Addressing data scarcity, we introduce a scalable pipeline to synthesize diverse code-switching data that is both semantically and acoustically natural, alongside a new multi-scenario benchmark, SCSBench.To address limitations of semantic tokenizers in capturing acoustic details, we also introduce refined cent token and Chainof-Thought (CoT) generation for planning prosody before content generation, effectively enhancing empathetic speech generation and singing melody.Experimental results demonstrate that UniVocal achieves state-of-the-art performance on SCSBench while maintaining competitive performance on regular speech and singing tasks.
Qian Chen 0003, Wen Wang 0001, Xiangang Li, Zhen-Hua Ling, Yang Ai
ACL (1)4
2026 GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling
abstract
Hao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang, Qian Chen, Lujia Bao, Xiangang Li, Zhen-Hua Ling. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang 0001, Qian Chen 0003, Lujia Bao, Xiangang Li, Zhen-Hua Ling
ACL (1)7
2025 Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning
abstract
Ruoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Ben He, Yingfei Sun, Xiangang Li, Le Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ruoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu 0001, Xianpei Han, Ben He 0001, Yingfei Sun, Xiangang Li, Le Sun 0001
ACL (1)9
2025 Can large language models independently complete tasks? A dynamic evaluation framework for multi-turn task planning and completion
Junlin Cui, Huijia Wu, Liuyu Xiang, Xiangang Li, Yaodong Yang 0001, Zhaofeng He 0001
Neurocomputing6
2024 OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
abstract
Nowadays, open-source large language models like LLaMA have emerged. Recent developments have incorporated supervised fine-tuning (SFT) and reinforcement learning fine-tuning (RLFT) to align these models with human goals. However, SFT methods treat all training data with mixed quality equally, while RLFT methods require high-quality pairwise or ranking-based preference data. In this study, we present a novel framework, named OpenChat, to advance open-source language models with mixed-quality data. Specifically, we consider the general SFT training data, consisting of a small amount of expert data mixed with a large proportion of sub-optimal data, without any preference labels. We propose the C(onditioned)-RLFT, which regards different data sources as coarse-grained reward labels and learns a class-conditioned policy to leverage complementary data quality information. Interestingly, the optimal policy in C-RLFT can be easily solved through single-stage, RL-free supervised learning, which is lightweight and avoids costly human preference labeling. Through extensive experiments on three standard benchmarks, our openchat-13b fine-tuned with C-RLFT achieves the highest average performance among all 13b open-source language models. Moreover, we use AGIEval to validate the model generalization performance, in which only openchat-13b surpasses the base model. Finally, we conduct a series of analyses to shed light on the effectiveness and robustness of OpenChat. Our code, data, and models are publicly available at https://github.com/imoneoi/openchat and https://huggingface.co/openchat.
Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, Yang Liu 0165
ICLR4
2024 An Effective Magnetic Anomaly Detection Using Orthonormal Basis of Magnetic Gradient Tensor Invariants
abstract
Magnetic anomaly detection (MAD) technology has been widely used in geological exploration, unexploded ordnance exploration, wreck salvage, and other fields because of its attractive advantages such as all-weather implementation, high concealment, and signal propagation in water and soil. The most classic method in MAD is orthonormal basis functions (OBF) detection which can effectively detect the magnetic target in a comparatively low signal-to-noise ratio (SNR) environment. In this paper, attracted by the excellent properties of magnetic gradient tensor (MGT) invariant, the I1-OBF and I2-OBF methods are developed. Considering strong dependence of the I2-OBF method on the angle between magnetic moment vector and displacement vector, and moderate detection performance of the I1-OBF method, an effective invariants-OBF method through adaptively integrating search energy of the I1-OBF and I2-OBF methods is proposed. Simulations and field experiments witnessed a great improvement in the SNR of the proposed method from the perspective of signal energy. It is proved that the proposed method can stably produce promising results even under the shaking-platform condition or the specific geometric relationship between the magnetic moment vector and displacement vector.
Youyu Yan, Shenggang Yan, Xiangang Li
IEEE Trans. Geosci. Remote. Sens.5
2023 Domain-Adapted Dependency Parsing for Cross-Domain Named Entity Recognition
abstract
In recent years, many researchers have leveraged structural information from dependency trees to improve Named Entity Recognition (NER). Most of their methods take dependency-tree labels as input features for NER model training. However, such dependency information is not inherently provided in most NER corpora, making the methods with low usability in practice. To effectively exploit the potential of word-dependency knowledge, motivated by the success of Multi-Task Learning on cross-domain NER, we investigate a novel NER learning method incorporating cross-domain Dependency Parsing (DP) as its auxiliary learning task. Then, considering the high consistency of word-dependency relations across domains, we present an unsupervised domain-adapted method to transfer word-dependency knowledge from high-resource domains to low-resource ones. With the help of cross-domain DP to bridge different domains, both useful cross-domain and cross-task knowledge can be learned by our model to considerably benefit cross-domain NER. To make better use of the cross-task knowledge between NER and DP, we unify both tasks in a shared network architecture for joint learning, using Maximum Mean Discrepancy(MMD). Finally, through extensive experiments, we show our proposed method can not only effectively take advantage of word-dependency knowledge, but also significantly outperform other Multi-Task Learning methods on cross-domain NER. Our code is open-source and available at https://github.com/xianghuisun/DADP.
Chenxiao Dou, Xianghui Sun, Yaoshu Wang, Yunjie Ji, Baochang Ma, Xiangang Li
AAAI6
2023 Semi-Supervised 2D Human Pose Estimation Driven by Position Inconsistency Pseudo Label Correction Module
abstract
In this paper, we delve into semi-supervised 2D human pose estimation. The previous method ignored two problems: (i) When conducting interactive training between large model and lightweight model, the pseudo label of lightweight model will be used to guide large models. (ii) The negative impact of noise pseudo labels on training. Moreover, the labels used for 2D human pose estimation are relatively complex: keypoint category and keypoint position. To solve the problems mentioned above, we propose a semi-supervised 2D human pose estimation framework driven by a position inconsistency pseudo label correction module (SSPCM). We introduce an additional auxiliary teacher and use the pseudo labels generated by the two teacher model in different periods to calculate the inconsistency score and remove outliers. Then, the two teacher models are updated through interactive training, and the student model is updated using the pseudo labels generated by two teachers. To further improve the performance of the student model, we use the semi-supervised Cut-Occlude based on pseudo keypoint perception to generate more hard and effective samples. In addition, we also proposed a new indoor overhead fisheye human keypoint dataset WEPDTOF-Pose. Extensive experiments demonstrate that our method outperforms the previous best semi-supervised 2D human pose estimation method. We will release the code and dataset at https://github.com/hlz0606/SSPCM
Linzhi Huang, Hongbo Tian, Xiangang Li, Weihong Deng, Jieping Ye
CVPR5
2022 Time Domain Adversarial Voice Conversion for ADD 2022
abstract
In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) system to convert source speech with arbitrary language content into target speaker’s fake speech. Then the converted speech generated from VC is post-processed in time-domain to improve the deception ability. The experimental results show that our system has adversarial ability against anti-spoofing detectors with a little compromise in audio quality and speaker similarity. This system ranks top in Track 3.1 in the ADD 2022, showing that our method could also gain good generalization ability against different detectors.
Cheng Wen 0004, Tingwei Guo, Xingjun Tan, Shuran Zhou, Chuandong Xie, Xiangang Li
ICASSP8
2022 Audio-Visual Wake Word Spotting System for MISP Challenge 2021
abstract
This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage both audio and video information to improve the environmental robustness of far-field wake word spotting. In the proposed system, firstly, we take advantage of speech enhancement algorithms such as beamforming and weighted prediction error (WPE) to address the multi-microphone conversational audio. Secondly, several data augmentation techniques are applied to simulate a more realistic far-field scenario. For the video information, the provided region of interest (ROI) is used to obtain visual representation. Then the multi-layer CNN is proposed to learn audio and visual representations, and these representations are fed into our two-branch attention-based net-work which can be employed for fusion, such as transformer and conformer. The focal loss is used to fine-tune the model and improve the performance significantly. Finally, multiple trained models are integrated by casting vote to achieve our final 0.091 score.
Yanguang Xu, Shuaijiang Zhao, Chaoyang Mei, Tingwei Guo, Shuran Zhou, Chuandong Xie, Xiangang Li
ICASSP10
2022 Audio Deepfake Detection System with Neural Stitching for ADD 2022
abstract
This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge[1]. The very same system was used for both two rounds of evaluation in Track 3.2 with similar training methodology. The first round of Track 3.2 data is generated from Text-to-Speech(TTS) or voice conversion (VC) algorithms, while the second round of data consists of generated fake audio from other participants in Track 3.1, aming to spoof our systems. Our systems uses a standard 34-layer ResNet [2], with multi-head attention pooling [3] to learn the discriminative embedding for fake audio and spoof detection. We further utilize neural stitching to boost the model’s generalization capability in order to perform equally well in different tasks, and more details will be explained in the following sessions. The experiments show that our proposed method outperforms all other systems with 10.1% equal error rate(EER) in Track 3.2.
Cheng Wen 0004, Shuran Zhou, Tingwei Guo, Xiangang Li
ICASSP6
2021 Didispeech: A Large Scale Mandarin Speech Corpus
abstract
This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is recorded in quiet environment and is suitable for various speech processing tasks, such as voice conversion, multi-speaker text-to-speech and automatic speech recognition. We conduct experiments with multiple speech tasks and evaluate the performance, showing that it is promising to use the corpus for both academic research and practical application. The corpus is available at https://outreach.didichuxing.com/research/opendata/.
Tingwei Guo, Cheng Wen 0004, Dongwei Jiang, Ne Luo, Ruixiong Zhang, Shuaijiang Zhao, Wubo Li, Xiangang Li
ICASSP11
2021 A Further Study of Unsupervised Pretraining for Transformer Based Speech Recognition
abstract
The construction of an effective good speech recognition system typically requires large amounts of transcribed data, which is expensive to collect. To overcome this problem, many unsupervised pretraining methods have been proposed. Among these methods, Masked Predictive Coding achieved significant improvements on various speech recognition datasets with BERT-like Masked Reconstruction loss and transformer backbone. However, many aspects of MPC have yet to be fully investigated. In this paper, we conduct a further study on MPC and focus on three important aspects: the effect of pretraining data speaking style, its extension on streaming model, and strategies for better transferring learned knowledge from pretraining stage to downstream tasks. The experimental results demonstrated that pretraining data with a matching speaking style is more useful on downstream recognition tasks. A unified training objective with APC and MPC provided an 8.46% relative error reduction on the streaming model trained on HKUST. Additionally, the combination of target data adaption and layerwise discriminative training facilitated the knowledge transfer of MPC, which realized 3.99% relative error reduction on AISHELL over a strong baseline.
Dongwei Jiang, Wubo Li, Ruixiong Zhang, Ne Luo, Xiangang Li
ICASSP9
2021 Transformer Based Unsupervised Pre-Training for Acoustic Representation Learning
abstract
Recently, a variety of acoustic tasks and related applications arised. For many acoustic tasks, the labeled data size may be limited. To handle this problem, we propose an unsupervised pre-training method using Transformer based encoder to learn a general and robust high-level representation for all acoustic tasks. Experiments have been conducted on three kinds of acoustic tasks: speech emotion recognition, sound event detection and speech translation. All the experiments have shown that pre-training using its own training data can significantly improve the performance. With a larger pre-training data combining MuST-C, Librispeech and ESC-US datasets, for speech emotion recognition, the UAR can further improve absolutely 4.3% on IEMOCAP dataset. For sound event detection, the F1 score can further improve absolutely 1.5% on DCASE2018 task5 development set and 2.1% on evaluation set. For speech translation, the BLEU score can further improve relatively 12.2% on En-De dataset and 8.4% on En-Fr dataset.
Ruixiong Zhang, Haiwei Wu, Wubo Li, Dongwei Jiang, Xiangang Li
ICASSP6
2021 GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10, 000 Hours of Transcribed Audio
abstract
This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training.Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc.A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription.For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h.For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%.The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality.Baseline systems are provided for popular speech recognition toolkits, namely Athena, ESPnet, Kaldi and Pika.
Guoguo Chen, Shuzhou Chai, Guan-Bo Wang, Jiayu Du, Weiqiang Zhang 0001, Chao Weng, Dan Su 0002, Daniel Povey, Jan Trmal, Mingjie Jin, Sanjeev Khudanpur, Shinji Watanabe 0001, Shuaijiang Zhao, Xiangang Li, Xuchen Yao, Zhao You, Zhiyong Yan
Interspeech16
2021 Speech SimCLR: Combining Contrastive and Reconstruction Objective for Self-Supervised Speech Representation Learning
abstract
Self-supervised visual pretraining has shown significant progress recently.Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semisupervised learning on ImageNet.The input feature representations for speech and visual tasks are both continuous, so it is natural to consider applying similar objective on speech representation learning.In this paper, we propose Speech SimCLR, a new self-supervised objective for speech representation learning.During training, Speech SimCLR applies augmentation on raw speech and its spectrogram.Its objective is the combination of contrastive loss that maximizes agreement between differently augmented samples in the latent space and reconstruction loss of input representation.The proposed method achieved competitive results on speech emotion recognition and speech recognition.
Dongwei Jiang, Wubo Li, Xiangang Li
Interspeech5
2021 Semantic Data Augmentation for End-to-End Mandarin Speech Recognition
abstract
End-to-end models have gradually become the preferred option for automatic speech recognition (ASR) applications.During the training of end-to-end ASR, data augmentation is a quite effective technique for regularizing the neural networks.This paper proposes a novel data augmentation technique based on semantic transposition of the transcriptions via syntax rules for end-to-end Mandarin ASR.Specifically, we first segment the transcriptions based on part-of-speech tags.Then transposition strategies, such as placing the object in front of the subject or swapping the subject and the object, are applied on the segmented sentences.Finally, the acoustic features corresponding to the transposed transcription are reassembled based on the audio-to-text forced-alignment produced by a pre-trained ASR system.The combination of original data and augmented one is used for training a new ASR system.The experiments are conducted on the Transformer[1] and Conformer[2] based ASR.The results show that the proposed method can give consistent performance gain to the system.Augmentation related issues, such as comparison of different strategies and ratios for data combination are also investigated.
Zhiyuan Tang, Hengxin Yin, Shuaijiang Zhao, Xiaoning Lei, Xiangang Li
Interspeech9
2020 DNN-based Mask Estimation Integrating Spectral and Spatial Features for Robust Beamforming
abstract
Spectral mask based beamforming has showed competitive performance on multi-channel speech enhancement in recent years. However, such methods apply mask estimation on each channel and ensemble the masks from multiple channels into one for speech and noise covariance estimation. Spectral-spatial mask estimation has not been well extended yet. In this paper, we propose a novel spectral-spatial mask based beamforming method for two-channel noisy signals, where spectral amplitude and cross-channel spatial features are integrated to improve mask estimation. Multi-channel masks are not merged in order to preserve channel characteristics for robust beamforming. Furthermore, this two-channel method is extended to six-channel scenario. Experiments on CHiME-3 evaluation confirm the superior performance of the proposed method over two spectral mask estimation approaches in terms of word error rates (WER) improvement.
Chengyun Deng, Yongtao Sha, Xiangang Li
ICASSP5
2020 Selective Attention Encoders by Syntactic Graph Convolutional Networks for Document Summarization
abstract
Abstractive text summarization is a challenging task, and one need to design a mechanism to effectively extract salient information from the source text and then generate a summary. A parsing process of the source text contains critical syntactic or semantic structures, which is useful to generate more accurate summary. However, modeling a parsing tree for text summarization is not trivial due to its non-linear structure and it is harder to deal with a document that includes multiple sentences and their parsing trees. In this paper, we propose to use a graph to connect the parsing trees from the sentences in a document and utilize the stacked graph convolutional networks (GCNs) to learn the syntactic representation for a document. The selective attention mechanism is used to extract salient information in semantic and structural aspect and generate an abstractive summary. We evaluate our approach on the CNN/Daily Mail text summarization dataset. The experimental results show that the proposed GCNs based selective attention approach outperforms the baselines and achieves the state-of-the-art performance on the dataset.
Haiyang Xu 0001, Baochang Ma, Junwen Chen 0005, Xiangang Li
ICASSP6
2020 Conv-TasSAN: Separative Adversarial Network Based on Conv-TasNet
Chengyun Deng, Shiqian Ma, Yongtao Sha, Xiangang Li
INTERSPEECH6
2020 TMT: A Transformer-Based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-Aware Dialog
abstract
Audio Visual Scene-aware Dialog (AVSD) is a task to generate responses when discussing about a given video. The previous state-of-the-art model shows superior performance for this task using Transformer-based architecture. However, there remain some limitations in learning better representation of modalities. Inspired by Neural Machine Translation (NMT), we propose the Transformer-based Modal Translator (TMT) to learn the representations of the source modal sequence by translating the source modal sequence to the related target modal sequence in a supervised manner. Based on Multimodal Transformer Networks (MTN), we apply TMT to video and dialog, proposing MTN-TMT for the video-grounded dialog system. On the AVSD track of the Dialog System Technology Challenge 7, MTN-TMT outperforms the MTN and other submission models in both Video and Text task and Text Only task. Compared with MTN, MTN-TMT improves all metrics, especially, achieving relative improvement up to 14.1% on CIDEr. Index Terms: multimodal learning, audio-visual scene-aware dialog, neural machine translation, multi-task learning
Wubo Li, Dongwei Jiang, Xiangang Li
INTERSPEECH4
2020 Generative Adversarial Network Based Acoustic Echo Cancellation
Chengyun Deng, Shiqian Ma, Yongtao Sha, Xiangang Li
INTERSPEECH6
2020 On Loss Functions and Recurrency Training for GAN-Based Speech Enhancement Systems
abstract
Recent work has shown that it is feasible to use generative adversarial networks (GANs) for speech enhancement, however, these approaches have not been compared to state-of-the-art (SOTA) non GAN-based approaches.Additionally, many loss functions have been proposed for GAN-based approaches, but they have not been adequately compared.In this study, we propose novel convolutional recurrent GAN (CRGAN) architectures for speech enhancement.Multiple loss functions are adopted to enable direct comparisons to other GAN-based systems.The benefits of including recurrent layers are also explored.Our results show that the proposed CRGAN model outperforms the SOTA GAN-based models using the same loss functions and it outperforms other non-GAN based systems, indicating the benefits of using a GAN for speech enhancement.Overall, the CRGAN model that combines an objective metric loss function with the mean squared error (MSE) provides the best performance over comparison approaches across many evaluation metrics.
Zhuohuang Zhang, Chengyun Deng, Yi Shen 0008, Donald S. Williamson, Yongtao Sha, Xiangang Li
INTERSPEECH8
2019 Replay Attack Detection Using Magnitude and Phase Information with Attention-based Adaptive Filters
abstract
Automatic Speech Verification (ASV) systems are highly vulnerable to spoofing attacks, and replay attack poses the greatest threat among various spoofing attacks. In this paper, we propose a novel multi-channel feature extraction method with attention-based adaptive filters (AAF). Original phase information, discarded by conventional feature extraction techniques after Fast Fourier Transform (FFT), is promising in distinguishing genuine from replay spoofed speech. Accordingly, phase and magnitude information are respectively extracted as phase channel and magnitude channel complementary features in our system. First, we make discriminative ability analysis on full frequency bands with F-ratio methods. Then attention-based adaptive filters are implemented to maximize capturing of high discriminative information on frequency bands, and the results on ASVspoof 2017 challenge indicate that our proposed approach achieved relative error reduction rates of 78.7% and 59.8% on development and evaluation dataset than the baseline method.
Meng Liu 0017, Longbiao Wang, Jianwu Dang 0001, Seiichi Nakagawa, Haotian Guan, Xiangang Li
ICASSP6
2019 NVSRN: A Neural Variational Scaling Reasoning Network for Initiative Response Generation
abstract
Open-domain multi-turn dialogue systems are booming in human-machine interactions, which encourage to chat actively and freely in an intelligent natural way. Previous generative conversational models usually employ a single and deterministic encoder-decoder framework to model the semantic consistency between the context and corresponding response. However, they neglect the various dialog patterns (we denote the regularity of topic shifting as dialog pattern) in the conversations, leading to uninformative, non-initiative yet plausible responses. Although the existing variational methods have improved the response diversity to some extent by introducing a global variability into the generative process, they fail to simulate the transfer between topics with directional information due to the weak interpretability of the Gaussian-distributed latent variables. In this paper, we propose a novel Neural Variational Scaling Reasoning Network (NVSRN) for initiative response generation. To this end, our approach has two core ingredients: neural dialog pattern reasoner (reasoner) and topic scaling mechanism. Specifically, inspired by the advantage of von Mises-Fisher (vMF) distribution modeling the directional data (e.g., the topic transfer state), we employ it as the latent space of the reasoner to explore the regularity of topic shifting, which is then used to reason the topic of response. Based on this, a topic scaling mechanism is designed to control the transfer degree of topic in the response generator. The experimental results on two large dialog datasets demonstrate that the proposed model outperforms state-of-the-art baselines. The human evaluation shows the proposed model can produce more informative and initiative responses actively.
Jinxin Chang, Ruifang He, Haiyang Xu 0001, Longbiao Wang, Xiangang Li, Jianwu Dang 0001
ICDM6
2019 Learning Syntactic and Dynamic Selective Encoding for Document Summarization
abstract
Text summarization aims to generate a headline or a short summary consisting of the major information of the source text. Recent studies employ the sequence-to-sequence framework to encode the input with a neural network and generate abstractive summary. However, most studies feed the encoder with the semantic word embedding but ignore the syntactic information of the text. Further, although previous studies proposed the selective gate to control the information flow from the encoder to the decoder, it is static during the decoding and cannot differentiate the information based on the decoder states. In this paper, we propose a novel neural architecture for document summarization. Our approach has the following contributions: first, we incorporate syntactic information such as constituency parsing trees into the encoding sequence to learn both the semantic and syntactic information from the document, resulting in more accurate summary; second, we propose a dynamic gate network to select the salient information based on the context of the decoder state, which is essential to document summarization. The proposed model has been evaluated on CNN/Daily Mail summarization datasets and the experimental results show that the proposed approach outperforms baseline approaches.
Haiyang Xu 0001, Yahao He, Junwen Chen 0005, Xiangang Li
IJCNN5
2019 Environment-Dependent Attention-Driven Recurrent Convolutional Neural Network for Robust Speech Enhancement
Meng Ge, Longbiao Wang, Jianwu Dang 0001, Xiangang Li
INTERSPEECH6
2019 Learning Alignment for Multimodal Emotion Recognition from Speech
abstract
Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways.For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech recognition techniques to generate text from speech and then apply natural language processing to analyze the sentiment.Further, emotion recognition will be beneficial from using audio-textual multimodal information, it is not trivial to build a system to learn from multimodality.One can build models for two input sources separately and combine them in a decision level, but this method ignores the interaction between speech and text in the temporal domain.In this paper, we propose to use an attention mechanism to learn the alignment between speech frames and text words, aiming to produce more accurate multimodal feature representations.The aligned multimodal features are fed into a sequential model for emotion recognition.We evaluate the approach on the IEMOCAP dataset and the experimental results show the proposed approach achieves the state-of-the-art performance on the dataset. 1
Haiyang Xu 0001, Hui Zhang 0016, Yiping Peng, Xiangang Li
INTERSPEECH6
2018 Implicit Discourse Relation Recognition using Neural Tensor Network with Interactive Attention and Sparse Learning
abstract
Implicit discourse relation recognition aims to understand and annotate the latent relations between two discourse arguments, such as temporal, comparison, etc. Most previous methods encode two discourse arguments separately, the ones considering pair specific clues ignore the bidirectional interactions between two arguments and the sparsity of pair patterns. In this paper, we propose a novel neural Tensor network framework with Interactive Attention and Sparse Learning (TIASL) for implicit discourse relation recognition. (1) We mine the most correlated word pairs from two discourse arguments to model pair specific clues, and integrate them as interactive attention into argument representations produced by the bidirectional long short-term memory network. Meanwhile, (2) the neural tensor network with sparse constraint is proposed to explore the deeper and the more important pair patterns so as to fully recognize discourse relations. The experimental results on PDTB show that our proposed TIASL framework is effective.
Fengyu Guo, Ruifang He, Di Jin 0001, Jianwu Dang 0001, Longbiao Wang, Xiangang Li
COLING6
2018 Interaction-Aware Topic Model for Microblog Conversations through Network Embedding and User Attention
abstract
Traditional topic models are insufficient for topic extraction in social media. The existing methods only consider text information or simultaneously model the posts and the static characteristics of social media. They ignore that one discusses diverse topics when dynamically interacting with different people. Moreover, people who talk about the same topic have different effects on the topic. In this paper, we propose an Interaction-Aware Topic Model (IATM) for microblog conversations by integrating network embedding and user attention. A conversation network linking users based on reposting and replying relationship is constructed to mine the dynamic user behaviours. We model dynamic interactions and user attention so as to learn interaction-aware edge embeddings with social context. Then they are incorporated into neural variational inference for generating the more consistent topics. The experiments on three real-world datasets show that our proposed model is effective.
Ruifang He, Di Jin 0001, Longbiao Wang, Jianwu Dang 0001, Xiangang Li
COLING6
2018 Speech Emotion Recognition by Combining Amplitude and Phase Information Using Convolutional Neural Network
Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Linjuan Zhang, Haotian Guan, Xiangang Li
INTERSPEECH6
2018 Multiple Phase Information Combination for Replay Attacks Detection
Dongbo Li, Longbiao Wang, Jianwu Dang 0001, Meng Liu 0017, Zeyan Oo, Seiichi Nakagawa, Haotian Guan, Xiangang Li
INTERSPEECH8
2017 Gram-CTC: Automatic Unit Selection and Target Decomposition for Sequence Labelling
abstract
Most existing sequence labelling models rely on a fixed decomposition of a target sequence into a sequence of basic units. These methods suffer from two major drawbacks: $1$) the set of basic units is fixed, such as the set of words, characters or phonemes in speech recognition, and $2$) the decomposition of target sequences is fixed. These drawbacks usually result in sub-optimal performance of modeling sequences. In this paper, we extend the popular CTC loss criterion to alleviate these limitations, and propose a new loss function called Gram-CTC. While preserving the advantages of CTC, Gram-CTC automatically learns the best set of basic units (grams), as well as the most suitable decomposition of target sequences. Unlike CTC, Gram-CTC allows the model to output variable number of characters at each time step, which enables the model to capture longer term dependency and improves the computational efficiency. We demonstrate that the proposed Gram-CTC improves CTC in terms of both performance and efficiency on the large vocabulary speech recognition task at multiple scales of data, and that with Gram-CTC we can outperform the state-of-the-art on a standard speech benchmark.
Hairong Liu, Zhenyao Zhu, Xiangang Li, Sanjeev Satheesh
ICML3
2016 Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu
ICML21
2015 Constructing long short-term memory based deep recurrent neural networks for large vocabulary speech recognition
abstract
Long short-term memory (LSTM) based acoustic modeling methods have recently been shown to give state-of-the-art performance on some speech recognition tasks. To achieve a further performance improvement, in this research, deep extensions on LSTM are investigated considering that deep hierarchical model has turned out to be more efficient than a shallow one. Motivated by previous research on constructing deep recurrent neural networks (RNNs), alternative deep LSTM architectures are proposed and empirically evaluated on a large vocabulary conversational telephone speech recognition task. Meanwhile, regarding to multi-GPU devices, the training process for LSTM networks is introduced and discussed. Experimental results demonstrate that the deep LSTM networks benefit from the depth and yield the state-of-the-art performance on this task.
Xiangang Li, Xihong Wu
ICASSP1
2015 Improving long short-term memory networks using maxout units for large vocabulary speech recognition
abstract
Long short-tem memory (LSTM) recurrent neural networks have been shown to give state-of-the-art performance on many speech recognition tasks. To achieve a further performance improvement, in this paper, maxout units are proposed to be integrated with the LSTM cells, considering those units have brought significant improvements to deep feed-forward neural networks. A novel architecture was constructed by replacing the input activation units (generally tanh) in the LSTM networks with maxout units. We implemented the LSTM network training on multi-GPU devices with truncated BPTT, and empirically evaluated the proposed designs on a large vocabulary Mandarin conversational telephone speech recognition task. The experimental results support our claim that the performance of LSTM based acoustic models can be further improved using the maxout units.
Xiangang Li, Xihong Wu
ICASSP1
2015 Modeling speaker variability using long short-term memory networks for speech recognition
Xiangang Li, Xihong Wu
INTERSPEECH1
2015 Long short-term memory based convolutional recurrent neural networks for large vocabulary speech recognition
abstract
Long short-term memory (LSTM) recurrent neural networks (RNNs) have been shown to give state-of-the-art performance on many speech recognition tasks, as they are able to provide the learned dynamically changing contextual window of all sequence history.On the other hand, the convolutional neural networks (CNNs) have brought significant improvements to deep feed-forward neural networks (FFNNs), as they are able to better reduce spectral variation in the input signal.In this paper, a network architecture called as convolutional recurrent neural network (CRNN) is proposed by combining the CNN and LSTM RNN.In the proposed CRNNs, each speech frame, without adjacent context frames, is organized as a number of local feature patches along the frequency axis, and then a LSTM network is performed on each feature patch along the time axis.We train and compare FFNNs, LSTM RNNs and the proposed LSTM CRNNs at various number of configurations.Experimental results show that the LSTM CRNNs can exceed stateof-the-art speech recognition performance.
Xiangang Li, Xihong Wu
INTERSPEECH1
2015 I-vector dependent feature space transformations for adaptive speech recognition
Xiangang Li, Xihong Wu
INTERSPEECH1
2015 A comparative study on selecting acoustic modeling units in deep neural networks based large vocabulary Chinese speech recognition
Xiangang Li, Zaihu Pang, Xihong Wu
Neurocomputing1
2014 Query-based composition for large-scale language model in LVCSR
abstract
This paper describes a query-based composition algorithm that can integrate an ARPA format language model in the unified WFST framework, which avoids the memory and time cost of converting the language models to WFST and optimizing the WFST of language models. The proposed algorithm is applied to on-the-fly one-pass decoder and rescoring decoder. Both modified decoder require less memory during decoding on different scale of language models. What's more, query-based on-the-fly one-pass decoder nearly has the same decoding speed as standard one and query-based rescoring decoder even use less time to rescore the lattice. Because of these advantages, large-scale language models can be applied by query-based composition algorithm to improve the performance of large vocabulary continuous speech recognition.
Xiangang Li, Xihong Wu
ICASSP3
2012 Probabilistic Speaker-Class based Acoustic Modeling for Large Vocabulary Continuous Speech Recognition
Xiangang Li, Zaihu Pang, Xihong Wu
INTERSPEECH1