Zhijie Yan

dblp:50/8761 · also Zhi-Jie Yan · DBLP profile ↗
← Back
50ranked-venue papers
8as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 27 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 A Composable Emulation Framework for Whitebox Switches
Congcong Miao, Xianneng Zou, Chuwen Zhang, Qihang Liu, Zhijie Yan, Yanke Zhang, Yong Jiang 0001, Qiao Xiang, Xin Jin 0008, Zili Meng, Ang Chen 0001
NSDI6
2025 OmniAudio: Generating Spatial Audio from 360-Degree Video
abstract
Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and perspective video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets are available at https://github.com/liuhuadai/OmniAudio. The project website is available at https://OmniAudio-360V2SA.github.io.
Huadai Liu, Tianyi Luo, Kaicheng Luo, Qikai Jiang, Peiwen Sun, Rongjie Huang 0001, Qian Chen 0003, Wen Wang 0001, Xiangtai Li, Shiliang Zhang, Zhijie Yan, Zhou Zhao 0001, Wei Xue 0002
ICML12
2025 IndVisSGG: VLM-based scene graph generation for industrial spatial intelligence
Zuoxu Wang, Zhijie Yan, Jihong Liu
Adv. Eng. Informatics2
2024 TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Bu Jin, Yupeng Zheng, Pengfei Li 0007, Weize Li 0001, Yuhang Zheng 0004, Sujie Hu, Zhijie Yan, Kun Zhan, Peng Jia 0007, Xiaoxiao Long, Hao Zhao 0002
ECCV (18)9
2024 Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving
abstract
Motion prediction is among the most fundamental tasks in autonomous driving. Traditional methods of motion forecasting primarily encode vector information of maps and historical trajectory data of traffic participants, lacking a comprehensive understanding of overall traffic semantics, which in turn affects the performance of prediction tasks. In this paper, we utilized Large Language Models (LLMs) to enhance the global traffic context understanding for motion prediction tasks. We first conducted systematic prompt engineering, visualizing complex traffic environments and historical trajectory information of traffic participants into image prompts— Transportation Context Map (TC-Map), accompanied by corresponding text prompts. Through this approach, we obtained rich traffic context information from the LLM. By integrating this information into the motion prediction model, we demonstrate that such context can enhance the accuracy of motion predictions. Furthermore, considering the cost associated with LLMs, we propose a cost-effective deployment strategy: enhancing the accuracy of motion prediction tasks at scale with 0.7% LLM-augmented datasets. Our research offers valuable insights into enhancing the understanding of traffic scenes of LLMs and the motion prediction performance of autonomous driving. The source code is available at https://github.com/AIR-DISCOVER/LLM-Augmented-MTR and https://aistudio.baidu.com/projectdetail/7809548.
Xiaoji Zheng, Lixiu Wu, Zhijie Yan, Yuanrong Tang, Hao Zhao 0002, Bokui Chen, Jiangtao Gong
IROS3
2023 The Second Multi-Channel Multi-Party Meeting Transcription Challenge (M2MeT 2.0): A Benchmark for Speaker-Attributed ASR
abstract
With the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle the complex task of speaker-attributed ASR (SAASR), which directly addresses the practical and challenging problem of “who spoke what at when” at typical meeting scenario. We particularly established two sub-tracks. The fixed training condition sub-track, where the training data is constrained to predetermined datasets, but participants can use any open-source pre-trained model. The open training condition sub-track, which allows for the use of all available data and models without limitation. In addition, we release a new 10-hour test set for challenge ranking. This paper provides an overview of the dataset, track settings, results, and analysis of submitted systems, as a benchmark to show the current state of speaker-attributed ASR.
Yuhao Liang, Mohan Shi, Fan Yu 0002, Yangze Li, Shiliang Zhang, Zhihao Du, Qian Chen 0003, Lei Xie 0001, Yanmin Qian, Jian Wu 0027, Zhuo Chen 0006, Kong-Aik Lee, Zhijie Yan, Hui Bu
ASRU13
2023 Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)
abstract
ICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SLP applications are critical to improve users’ efficiency in grasping important information in meetings. MUG includes five tracks, including topic segmentation, topic-level and session-level extractive summarization, topic title generation, keyphrase extraction, and action item detection. To facilitate MUG, we construct and release a large-scale meeting dataset, the AliMeeting4MUG Corpus. We review the dataset, track settings and baselines, and summarize the challenge results and major techniques used in the submissions.
Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001
ICASSP7
2023 MUG: A General Meeting Understanding and Generation Benchmark
abstract
Listening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR transcripts only partly speeds up seeking information. It has been observed that a range of NLP applications, such as keyphrase extraction, topic segmentation, and summarization, significantly improve users’ efficiency in grasping important information. The meeting scenario is among the most valuable scenarios for deploying these spoken language processing (SLP) capabilities. However, the lack of large-scale public meeting datasets annotated for these SLP tasks severely hinders their advancement. To prompt SLP advancement, we establish a large-scale general Meeting Understanding and Generation Benchmark (MUG) to benchmark the performance of a wide range of SLP tasks, including topic segmentation, topic-level and session-level extractive summarization and topic title generation, keyphrase extraction, and action item detection. To facilitate the MUG benchmark, we construct and release a large-scale meeting dataset for comprehensive long-form SLP development, the AliMeeting4MUG Corpus, which consists of 654 recorded Mandarin meeting sessions with diverse topic coverage, with manual annotations for SLP tasks on manual transcripts of meeting recordings. To the best of our knowledge, the AliMeeting4MUG Corpus is so far the largest meeting corpus in scale and facilitates most SLP tasks. In this paper, we provide a detailed introduction of this corpus, SLP tasks and evaluation methods, baseline systems and their performance1.
Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001
ICASSP7
2023 INT2: Interactive Trajectory Prediction at Intersections
abstract
Motion forecasting is an important component in autonomous driving systems. One of the most challenging problems in motion forecasting is interactive trajectory prediction, whose goal is to jointly forecasts the future trajectories of interacting agents. To this end, we present a large-scale interactive trajectory prediction dataset named INT2 for INTeractive trajectory prediction at INTersections. INT2 includes 612,000 scenes, each lasting 1 minute, containing up to 10,200 hours of data. The agent trajectories are auto-labeled by a high-performance offline temporal detection and fusion algorithm, whose quality is further inspected by human judges. Vectorized semantic maps and traffic light information are also included in INT2. Additionally, the dataset poses an interesting domain mismatch challenge. For each intersection, we treat rush-hour and non-rush-hour segments as different domains. We benchmark the best open-sourced interactive trajectory prediction method on INT2 and Waymo Open Motion, under in-domain and cross-domain settings. The dataset, code and models are publicly available at https://github.com/AIRDISCOVER/INT2.
Zhijie Yan, Pengfei Li 0007, Zheng Fu, Shaocong Xu, Yongliang Shi, Xiaoxue Chen, Yuhang Zheng 0004, Yang Li 0178, Tianyu Liu 0008, Chuxuan Li, Nairui Luo, Zuoxu Wang, Yifeng Shi, Zhengxiao Han, Jirui Yuan, Jiangtao Gong, Guyue Zhou, Hang Zhao 0021, Hao Zhao 0002
ICCV1
2023 Accurate and Reliable Confidence Estimation Based on Non-Autoregressive End-to-End Speech Recognition System
Xian Shi, Haoneng Luo, Zhifu Gao, Shiliang Zhang, Zhijie Yan
INTERSPEECH5
2023 MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for speech recognition
Xiaohuan Zhou, Jiaming Wang 0004, Zeyu Cui, Shiliang Zhang, Zhijie Yan, Jingren Zhou 0001, Chang Zhou 0005
INTERSPEECH5
2022 Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis
abstract
Recently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis.However, current models always treat overlapped speaker diarization as a multi-label classification problem, where speaker dependency and overlaps are not well considered.To overcome the disadvantages, we reformulate overlapped speaker diarization task as a single-label prediction problem via the proposed power set encoding (PSE).Through this formulation, speaker dependency and overlaps can be explicitly modeled.To fully leverage this formulation, we further propose the speaker overlap-aware neural diarization (SOND) model, which consists of a contextindependent (CI) scorer to model global speaker discriminability, a context-dependent scorer (CD) to model local discriminability, and a speaker combining network (SCN) to combine and reassign speaker activities.Experimental results show that using the proposed formulation can outperform the state-ofthe-art methods based on target speaker voice activity detection, and the performance can be further improved with SOND, resulting in a 6.30% relative diarization error reduction.
Zhihao Du, Shiliang Zhang, Zhijie Yan
EMNLP4
2022 Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-Speech
abstract
Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable errors, which hurts the prosody modeling; 2) different attributes of prosody (e.g., pitch, duration and energy) are dependent on each other and produce the natural prosody together; and 3) due to high variability of prosody and the limited amount of high-quality data for TTS training, the distribution of prosody cannot be fully shaped. To tackle these issues, we propose ProsoSpeech, which enhances the prosody using quantized latent vectors pre-trained on large-scale unpaired and low-quality text and speech data. Specifically, we first introduce a word-level prosody encoder, which quantizes the low-frequency band of the speech and compresses prosody at-tributes in the latent prosody vector (LPV). Then we introduce an LPV predictor, which predicts LPV given word sequence. We pre-train the LPV predictor on large-scale text and low-quality speech data and fine-tune it on the high-quality TTS dataset. Finally, our model can generate expressive speech conditioned on the predicted LPV. Experimental results show that ProsoSpeech can generate speech with richer prosody compared with baseline methods.
Yi Ren 0006, Zhiying Huang, Shiliang Zhang, Qian Chen 0003, Zhijie Yan, Zhou Zhao 0001
ICASSP6
2022 M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
abstract
Recent development of speech signal processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most challenging scenarios for the deployment of speech technologies. Speaker diarization and multi-speaker automatic speech recognition in meeting scenarios have attracted much attention recently. However, the lack of large public meeting data has been a major obstacle for advancement of the field. Therefore, we make available the AliMeeting corpus, which consists of 120 hours of recorded Mandarin meeting data, including far-field data collected by 8-channel microphone array as well as near-field data collected by headset microphone. Each meeting session is composed of 2-4 speakers with different speaker overlap ratio, recorded in meeting rooms with different size. Along with the dataset, we launch the ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) with two tracks, namely speaker diarization and multi-speaker ASR, aiming to provide a common testbed for meeting rich transcription and promote reproducible research in this field. In this paper we provide a detailed introduction of the AliMeeting dateset, challenge rules, evaluation methods and baseline systems.
Fan Yu 0002, Shiliang Zhang, Yihui Fu, Lei Xie 0001, Zhihao Du, Weilong Huang, Zhijie Yan, Bin Ma 0001, Hui Bu
ICASSP9
2022 Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
abstract
The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic speech recognition (ASR) (track 2). Along with the challenge, we released 120 hours of real-recorded Mandarin meeting speech data with manual annotation, including far-field data collected by 8-channel micro-phone array as well as near-field data collected by each participants’ headset microphone. We briefly describe the released dataset, track setups, baselines and summarize the challenge results and major techniques used in the submissions.
Fan Yu 0002, Shiliang Zhang, Yihui Fu, Zhihao Du, Weilong Huang, Lei Xie 0001, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong-Aik Lee, Zhijie Yan, Bin Ma 0001, Hui Bu
ICASSP13
2022 Surface Defect Detection and Classification Based on Fusing Multiple Computer Vision Techniques
Bingqing Shen, Chongyu Wang, Guoxin Hou, Zhijie Yan, Hongming Cai 0001
IEA/AIE6
2022 Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
abstract
Transformers have recently dominated the ASR field. Although able to yield good performance, they involve an autoregressive (AR) decoder to generate tokens one by one, which is computationally inefficient. To speed up inference, non-autoregressive (NAR) methods, e.g. single-step NAR, were designed, to enable parallel generation. However, due to an independence assumption within the output tokens, performance of single-step NAR is inferior to that of AR models, especially with a large-scale corpus. There are two challenges to improving single-step NAR: Firstly to accurately predict the number of output tokens and extract hidden variables; secondly, to enhance modeling of interdependence between output tokens. To tackle both challenges, we propose a fast and accurate parallel transformer, termed Paraformer. This utilizes a continuous integrate-and-fire based predictor to predict the number of tokens and generate hidden variables. A glancing language model sampler then generates semantic embeddings to enhance the NAR decoder's ability to model context interdependence. Finally, we design a strategy to generate negative samples for minimum word error rate training to further improve performance. Experiments using the AISHELL-1, AISHELL-2 benchmark, and an industrial-level 20,000 hour task demonstrate that the proposed Paraformer can attain comparable performance to the state-of-the-art AR transformer, with over 10x speedup.
Zhifu Gao, Shiliang Zhang, Ian McLoughlin 0001, Zhijie Yan
INTERSPEECH4
2021 A Real-Time Speaker Diarization System Based on Spatial Spectrum
abstract
In this paper we describe a speaker diarization system that enables localization and identification of all speakers present in a conversation or meeting. We propose a novel systematic approach to tackle several long-standing challenges in speaker diarization tasks: (1) to segment and separate overlapping speech from two speakers; (2) to estimate the number of speakers when participants may enter or leave the conversation at any time; (3) to provide accurate speaker identification on short text-independent utterances; (4) to track down speakers movement during the conversation; (5) to detect speaker change incidence real-time. First, a differential directional microphone array-based approach is exploited to capture the target speakers’ voice in far-field adverse environment. Second, an online speaker-location joint clustering approach is proposed to keep track of speaker location. Third, an instant speaker number detector is developed to trigger the mechanism that separates overlapped speech. The results suggest that our system effectively incorporates spatial information and achieves significant gains.
Weilong Huang, Xianliang Wang, Hongbin Suo, Jinwei Feng, Zhijie Yan
ICASSP6
2021 Investigation of Spatial-Acoustic Features for Overlapping Speech Detection in Multiparty Meetings
Shiliang Zhang, Weilong Huang, Hongbin Suo, Jinwei Feng, Zhijie Yan
Interspeech7
2020 Neural Zero-Inflated Quality Estimation Model for Automatic Speech Recognition System
abstract
The performances of automatic speech recognition (ASR) systems are usually evaluated by the metric word error rate (WER) when the manually transcribed data are provided, which are, however, expensively available in the real scenario.In addition, the empirical distribution of WER for most ASR systems usually tends to put a significant mass near zero, making it difficult to simulate with a single continuous distribution.In order to address the two issues of ASR quality estimation (QE), we propose a novel neural zero-inflated model to predict the WER of the ASR result without transcripts.We design a neural zeroinflated beta regression on top of a bidirectional transformer language model conditional on speech features (speech-BERT).We adopt the pre-training strategy of token level masked language modeling for speech-BERT as well, and further fine-tune with our zero-inflated layer for the mixture of discrete and continuous outputs.The experimental results show that our approach achieves better performance on WER prediction compared with strong baselines.
Kai Fan 0002, Bo Li 0121, Jiayi Wang 0010, Shiliang Zhang, Boxing Chen, Niyu Ge, Zhijie Yan
INTERSPEECH7
2020 Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
abstract
Recently, streaming end-to-end automatic speech recognition (E2E-ASR) has gained more and more attention.Many efforts have been paid to turn the non-streaming attention-based E2E-ASR system into streaming architecture.In this work, we propose a novel online E2E-ASR system by using Streaming Chunk-Aware Multihead Attention (SCAMA) and a latency control memory equipped self-attention network (LC-SAN-M).LC-SAN-M uses chunk-level input to control the latency of encoder.As to SCAMA, a jointly trained predictor is used to control the output of encoder when feeding to decoder, which enables decoder to generate output in streaming manner.Experimental results on the open 170-hour AISHELL-1 and an industrial-level 20000-hour Mandarin speech recognition tasks show that our approach can significantly outperform the MoChA-based baseline system under comparable setup.On the AISHELL-1 task, our proposed method achieves a character error rate (CER) of 7.39%, to the best of our knowledge, which is the best published performance for online ASR.
Shiliang Zhang, Zhifu Gao, Haoneng Luo, Zhijie Yan, Lei Xie 0001
INTERSPEECH6
2019 Investigation of Transformer Based Spelling Correction Model for CTC-Based End-to-End Mandarin Speech Recognition
Shiliang Zhang, Zhijie Yan
INTERSPEECH3
2018 Deep Feed-Forward Sequential Memory Networks for Speech Synthesis
abstract
The Bidirectional LSTM (BLSTM) RNN based speech synthesis system is among the best parametric Text-to-Speech (TTS) systems in terms of the naturalness of generated speech, especially the naturalness in prosody. However, the model complexity and inference cost of BLSTM prevents its usage in many runtime applications. Meanwhile, Deep Feed-forward Sequential Memory Networks (DFSMN) has shown its consistent out-performance over BLSTM in both word error rate (WER) and the runtime computation cost in speech recognition tasks. Since speech synthesis also requires to model long-term dependencies compared to speech recognition, in this paper, we investigate the Deep-FSMN (DFSMN) in speech synthesis. Both objective and subjective experiments show that, compared with BLSTM TTS method, the DFSMN system can generate synthesized speech with comparable speech quality while drastically reduce model complexity and speech generation time.
Mengxiao Bi, Shiliang Zhang, Zhijie Yan
ICASSP5
2018 Linear Networks Based Speaker Adaptation for Speech Synthesis
abstract
Speaker adaptation methods aim to create fair quality synthesis speech voice font for target speakers while only limited resources available. Recently, as deep neural networks based statistical parametric speech synthesis (SPSS) methods become dominant in SPSS TTS back-end modeling, speaker adaptation under the neural network based SPSS framework has also became an important task. In this paper, linear networks (LN) is inserted in multiple neural network layers and fine-tuned together with output layer for best speaker adaptation performance. When adaptation data is extremely small, the low-rank plus diagonal(LRPD) decomposition for LN is employed to make the adapted voice more stable. Speaker adaptation experiments are conducted under a range of adaptation utterances numbers. Moreover, speaker adaptation from 1) female to female, 2) male to female and 3) female to male are investigated. Objective measurement and subjective tests show that LN with LRPD decomposition performs most stable when adaptation data is extremely limited, and our best speaker adaptation (SA) model with only 200 adaptation utterances achieves comparable quality with speaker dependent (SD) model trained with 1000 utterances, in both naturalness and similarity to target speaker.
Zhiying Huang, Zhijie Yan
ICASSP4
2018 Deep-FSMN for Large Vocabulary Continuous Speech Recognition
abstract
In this paper, we present an improved feedforward sequential memory networks (FSMN) architecture, namely Deep-FSMN (DFSMN), by introducing skip connections between memory blocks in adjacent layers. These skip connections enable the information flow across different layers and thus alleviate the gradient vanishing problem when building very deep structure. As a result, DFSMN significantly benefits from these skip connections and deep structure. We have compared the performance of DFSMN to BLSTM both with and without lower frame rate (LFR) on several large speech recognition tasks, including English and Mandarin. Experimental results shown that DFSMN can consistently outperform BLSTM with dramatic gain, especially trained with LFR using CD-Phone as modeling units. In the 20000 hours Fisher (FSH) task, the proposed DFSMN can achieve a word error rate of 9.4% by purely using the cross-entropy criterion and decoding with a 3-gram language model, which achieves a 1.5% absolute improvement compared to the BLSTM. In a 20000 hours Mandarin recognition task, the LFR trained DFSMN can achieve more than 20% relative improvement compared to the LFR trained BLSTM. Moreover, we can easily design the lookahead filter order of the memory blocks in DFSMN to control the latency for real-time applications.
Shiliang Zhang, Zhijie Yan, Li-Rong Dai 0001
ICASSP3
2017 Improving latency-controlled BLSTM acoustic models for online speech recognition
abstract
Bidirectional long short-term memory (BLSTM) recurrent neural networks are powerful acoustic models in terms of recognition accuracy. When BLSTM acoustic models are used in decoding, the speech decoder needs to wait until the end of a whole sentence is reached, such that forward-propagation in the backward direction can then be performed. The nature of BLSTM acoustic models makes them inappropriate for real-time online speech recognition because of the latency issue. Recently, the context-sensitive-chunk BLSTM and latency-controlled BLSTM acoustic models have been proposed, both chop a whole sentence into several overlapping chunks. By appending several left and/or right contextual frames, forward-propagation of BLSTM can be down within a controlled time delay, while the recognition accuracy is maintained when comparing with conventional BLSTM models. In this paper, two improved versions of latency-controlled BLSTM acoustic models are presented. By using different types of neural network topology to initialize the BLSTM memory cell states, we aim at reducing the computational cost introduced by the contextual frames and enabling faster online recognition. Experimental results on a 320-hour Switchboard task have shown that the improved versions accelerate from 24% to 61% in decoding without significant loss in recognition accuracy.
Shaofei Xue, Zhijie Yan
ICASSP2
2015 A context-sensitive-chunk BPTT approach to training deep LSTM/BLSTM recurrent neural networks for offline handwriting recognition
abstract
We propose a context-sensitive-chunk based back-propagation through time (BPTT) approach to training deep (bidirectional) long short-term memory ((B)LSTM) recurrent neural networks (RNN) that splits each training sequence into chunks with appended contextual observations for character modeling of offline handwriting recognition. Using short context-sensitive chunks in both training and recognition brings following benefits: (1) the learned (B)LSTM will model mainly local character image dependency and the effect of long-range language model information reflected in training data is reduced; (2) mini-batch based training on GPU can be made more efficient; (3) low-latency BLSTM-based handwriting recognition is made possible by incurring only a delay of a short chunk rather than a whole sentence. Our approach is evaluated on IAM offline handwriting recognition benchmark task and performs better than the previous state-of-the-art BPTT-based approaches.
Kai Chen 0001, Zhijie Yan, Qiang Huo
ICDAR2
2015 Training deep bidirectional LSTM acoustic model for LVCSR by a context-sensitive-chunk BPTT approach
Kai Chen 0001, Zhijie Yan, Qiang Huo
INTERSPEECH2
2014 An Unsupervised Adaptation Approach to Leveraging Feedback Loop Data by Using i-Vector for Data Clustering and Selection
abstract
We present a study of using unsupervised adaptation approaches to improve speech recognition accuracy of a deployed speech service by leveraging large-scale untranscribed speech data collected from a feedback loop (FBL). For a regular user with lots of adaptation utterances, conventional CMLLR-based adaptation can be used for personalization directly. For a casual user with a few adaptation utterances, we propose to use CMLLR-based adaptation by augmenting his / her adaptation utterances with utterances acoustically close to the user, which are selected from the FBL data by an i-vector based approach. For a new user, we propose to perform a CMLLR-based recognition of an unknown utterance by selecting a set of CMLLR transforms from the most similar cluster, which are pre-trained by using the utterances from the corresponding cluster generated by an i-vector based utterance clustering method from the FBL data. The effectiveness of the above approaches are confirmed by our experiments on a short message dictation task on smart phones.
Jian Xu 0006, Zhijie Yan, Qiang Huo
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Tied-state based discriminative training of context-expanded region-dependent feature transforms for LVCSR
abstract
We present a new discriminative feature transform approach to large vocabulary continuous speech recognition (LVCSR) using Gaussian mixture density hidden Markov models (GMM-HMMs) for acoustic modeling. The feature transform is formulated with a set of context-expanded region-dependent linear transforms (RDLTs) utilizing both long-span features and contextual weight expansion. The RDLTs are estimated by lattice-free, tied-state based discriminative training using maximum mutual information (MMI) criterion, while the GMM-HMMs are trained by conventional lattice-based, boosted MMI training. Compared with two baseline systems, which use RDLTs with either long-span features or weight expansion only and are trained using the conventional lattice-based discriminative training for both RDLTs and HMMs, the proposed approach achieves a relative word error rate reduction of 10% and 6% respectively on Switchboard-1 conversational telephone speech transcription task.
Zhijie Yan, Qiang Huo, Jian Xu 0006, Yu Zhang 0007
ICASSP1
2013 A scalable approach to using DNN-derived features in GMM-HMM based acoustic modeling for LVCSR
abstract
We present a new scalable approach to using deep neural network (DNN) derived features in Gaussian mixture density hidden Markov model (GMM-HMM) based acoustic modeling for large vocabulary continuous speech recognition (LVCSR). The DNN-based feature extractor is trained from a subset of training data to mitigate the scalability issue of DNN training, while GMM-HMMs are trained by using state-of-the-art scalable training methods and tools to leverage the whole training set. In a benchmark evaluation, we used 309-hour Switchboard-I (SWB) training data to train a DNN first, which achieves a word error rate (WER) of 15.4% on NIST-2000 Hub5 evaluation set by a traditional DNN-HMM based approach. When the same DNN is used as a feature extractor and 2,000-hour “SWB+Fisher” training data is used to train the GMM-HMMs, our DNN-GMM-HMM approach achieves a WER of 13.8%. If per-conversation-side based unsupervised adaptation is performed, a WER of 13.1% can be achieved.
Zhijie Yan, Qiang Huo, Jian Xu 0006
INTERSPEECH1
2013 A Unified Trajectory Tiling Approach to High Quality Speech Rendering
abstract
It is technically challenging to make a machine talk as naturally as a human so as to facilitate “frictionless” interactions between machine and human. We propose a trajectory tiling-based approach to high-quality speech rendering, where speech parameter trajectories, extracted from natural, processed, or synthesized speech, are used to guide the search for the best sequence of waveform “tiles” stored in a pre-recorded speech database. We test the proposed unified algorithm in both Text-To-Speech (TTS) synthesis and cross-lingual voice transformation applications. Experimental results show that the proposed trajectory tiling approach can render speech which is both natural and highly intelligible. The perceived high quality of rendered speech is also confirmed in both objective and subjective evaluations.
Yao Qian, Frank K. Soong, Zhijie Yan
IEEE Trans. Speech Audio Process.3
2012 A study of discriminative feature extraction for i-vector based acoustic sniffing in IVN acoustic model training
abstract
Recently, we proposed an i-vector approach to acoustic sniffing for irrelevant variability normalization based acoustic model training in large vocabulary continuous speech recognition (LVCSR). Its effectiveness has been confirmed by experimental results on Switchboard- 1 conversational telephone speech transcription task. In this paper, we study several discriminative feature extraction approaches in i-vector space to improve both recognition accuracy and run-time efficiency. New experimental results are reported on a much larger scale LVCSR task with about 2000 hours training data.
Yu Zhang 0007, Jian Xu 0006, Zhijie Yan, Qiang Huo
ICASSP3
2012 Tip tap tones: mobile microtraining of mandarin sounds
abstract
Learning a second language is hard, especially when the learner's brain must be retrained to identify sounds not present in his or her native language. It also requires regular practice, but many learners struggle to find the time and motivation. Our solution is to break down the challenge of mastering a foreign sound system into minute-long episodes of "microtraining" delivered through mobile gaming. We present the example of Tip Tap Tones - a mobile game with the purpose of helping learners acquire the tonal sound system of Mandarin Chinese. In a 3-week, 12-user study of this system, we found that an average of 71 minutes' gameplay significantly improved tone identification by around 25%, regardless of whether the underlying sounds had been used to train tone perception. Overall, results suggest that mobile microtraining is an efficient, effective, and enjoyable way to master the sounds of Mandarin Chinese, with applications to other languages and domains.
Darren Edge, Kai-Yin Cheng, Michael Whitney, Yao Qian, Zhijie Yan, Frank K. Soong
Mobile HCI5
2011 Speaker characterization using spectral subband energy ratio based on Harmonic plus Noise Model
abstract
This paper proposes a feature extraction for speaker characterization by exploring the relationship between the two distinct components of the speech signal, one is harmonics accounting for the periodicity of the signal and the other is modulated noise accounting for the turbulences of the glottal airflow. The harmonic and noise parts of the speech signal are decomposed based on the Harmonic plus Noise Model approach. We estimate the spectral subband energy ratios (SSERs) as the speaker characteristic features, which are expected to reflect the interaction property of the vocal tract and glottal airflow of individual speakers for speaker verification. The speaker verification experiments based on a GMM-UBM system have shown the efficiency of the SSER features, reducing the error equal rate by 27.2% by combining with the conventional MFCC features.
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo
ICASSP2
2011 A study of an irrelevant variability normalization based discriminative training approach for LVCSR
abstract
This paper presents a discriminative training (DT) approach to irrelevant variability normalization (IVN) based training of feature transforms and hidden Markov models for large vocabulary continuous speech recognition. A speaker-clustering based method is used for acoustic sniffing and maximum mutual information (MMI) is used as a training criterion. Combined with unsupervised adaptation of feature transforms, the IVN-based DT approach achieves a 14.5% relative word error rate reduction over an MMI-trained baseline system on a Switchboard-1 conversational telephone speech transcription task.
Yu Zhang 0007, Jian Xu 0006, Zhijie Yan, Qiang Huo
ICASSP3
2011 Improvements in Speaker Characterization Using Spectral Subband Energy Based on Harmonic plus Noise Model
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo
INTERSPEECH2
2011 An i-vector Based Approach to Acoustic Sniffing for Irrelevant Variability Normalization Based Acoustic Model Training and Speech Recognition
Jian Xu 0006, Yu Zhang 0007, Zhijie Yan, Qiang Huo
INTERSPEECH3
2011 An i-vector Based Approach to Training Data Clustering for Improved Speech Recognition
abstract
We present a new approach to clustering training data for improved speech recognition. Given a training corpus, a so-called i-vector is extracted from each training utterance. A hierarchical divisive clustering algorithm is then used to cluster the training i-vectors into multiple clusters. For each cluster, an acoustic model (AM) is trained accordingly. Such trained multiple AMs can then be used in recognition stage to improve recognition accuracy. The proposed approach is very efficient therefore can deal with very large scale training corpus on current mainstream computing platforms. We report experimental results on a voice search task with 7,500 hours of speech training data.
Yu Zhang 0007, Jian Xu 0006, Zhijie Yan, Qiang Huo
INTERSPEECH3
2010 RIch-context Unit Selection (RUS) approach to high quality TTS
abstract
This paper presents a Rich-context Unit Selection (RUS) approach to high quality speech synthesis. Based upon our previous work on rich context modeling, we use the corresponding parametric HMMs to represent waveform units and form a “sausage-like” lattice. A prune-and-search procedure is proposed, in which Kullback-Leibler divergence is adopted to select potential candidate units, and normalized cross-correlation is used as the final objective measure to search for the optimal unit path. The maximum cross-correlation criterion provides the optimal concatenation between successive units, in terms of spectral similarity, phase continuity and best connecting timing instants. Subjectively, both preference and MOS tests were conducted to compare RUS with our current Weight-table based Unit Selection (WUS) synthesis. Experimental results show that the voice quality of synthesized speech is significantly improved by RUS over the conventional WUS.
Zhijie Yan, Yao Qian, Frank K. Soong
ICASSP1
2010 Improved modeling for F0 generation and V/U decision in HMM-based TTS
abstract
The HMM-based TTS can produce a highly intelligible and decent quality voice. However, sometimes the synthesized speech exhibits perceptibly annoying glitches due to F0 extraction errors in the training data and voiced/unvoiced swapping errors in F0 generation. In the conventional MSD based F0 modeling [10], the dual but incompatible two probabilistic spaces, the continuous probability density for voiced observations or the discrete probability for unvoiced observations, prevent us from using likelihood based frame occupancy to alleviate the deteriorating effect of F0 extraction errors in training a more robust model for synthesis. In this paper, we propose a new approach to improved modeling the piece-wise continuous F0 trajectory and v/u decision for HMM-based TTS. Voicing strength, characterized by the normalized correlation coefficient magnitude calculated in F0 feature extraction, is used as an additional feature in F0 modeling and for v/u decision. Experimental results show the new approach to F0 modeling and generation outperforms MSD-HMM method and a newly proposed GTD-HMM method [9] significantly. The improvements are both objectively measurable and subjectively perceivable.
Frank K. Soong, Yao Qian, Zhijie Yan, Jielin Pan, Yonghong Yan 0002
ICASSP4
2010 Cross-validation based decision tree clustering for HMM-based TTS
abstract
In HMM-based speech synthesis, we usually use complex, context dependent models to characterize prosodically and linguistically rich speech units. It is therefore difficult to prepare training data which can cover all combinatorial possibilities of contexts. A common approach to cope with this insufficient training data problem is to build a clustered tree via the MDL criterion. However, an MDL-based tree still tends to be inadequate in its power to predict unseen data. In this paper, we adopt the cross-validation principle to build such a decision tree to minimize the generation error of unseen contexts. An efficient training algorithm is implemented by exploiting the sufficient statistics. Experimental results show that the proposed method can achieve better speech synthesis results, both objectively and subjectively, than the baseline results of the MDL-based decision tree.
Yu Zhang 0007, Zhijie Yan, Frank K. Soong
ICASSP2
2010 A perceptual study of acceleration parameters in HMM-based TTS
Zhijie Yan, Frank K. Soong
INTERSPEECH2
2010 An HMM trajectory tiling (HTT) approach to high quality TTS
Yao Qian, Zhijie Yan, Yi-Jian Wu, Frank K. Soong, Xin Zhuang, Shengyi Kong
INTERSPEECH2
2009 A trust region based optimization for maximum mutual information estimation of HMMS in speech recognition
abstract
In this paper, we present a new optimization method for MMIE-based discriminative training of HMMs in speech recognition. In our method, the MMIE training of Gaussian mixture HMMs is formulated as a so-called trust region problem, where a quadratic objective function is minimized under a spherical constraint, so that an efficient global optimization method for the trust region problem can be used to solve the MMIE training problem of HMMs. Experimental results on the WSJ0 Nov'92 evaluation task demonstrate that the trust region based optimization significantly outperforms the conventional EBWmethod in terms of optimization convergence behavior as well as speech recognition performance. It has been observed that the trust region method achieves up to 23.3% relative recognition error reduction over a well-trained MLE system while the EBW method gives only 13.3% relative error reduction.
Zhijie Yan, Cong Liu 0006, Yu Hu 0003, Hui Jiang 0001
ICASSP1
2009 Rich context modeling for high quality HMM-based TTS
abstract
This paper presents a rich context modeling approach to high quality HMM-based speech synthesis. We first analyze the over-smoothing problem in conventional decision tree tyingbased HMM, and then propose to model the training speech tokens with rich context models. Special training procedure is adopted for reliable estimation of the rich context model parameters. In synthesis, a search algorithm following a contextbased pre-selection is performed to determine the optimal rich context model sequence which generates natural and crisp output speech. Experimental results show that spectral envelopes synthesized by the rich context models are with crisper formant structures and evolve with richer details than those obtained by the conventional models. The speech quality improvement is also perceived by listeners in a subjective preference test, in which 76% of the sentences synthesized using rich context modeling are preferred. Index Terms: HMM-based TTS, rich context modeling
Zhijie Yan, Yao Qian, Frank K. Soong
INTERSPEECH1
2008 Minimum word classification error training of HMMS for automatic speech recognition
abstract
This paper presents a novel discriminative training criterion, minimum word classification error (MWCE). By localizing conventional string-level MCE loss function to word-level, a more direct measure of empirical word classification error is approximated and minimized. Because the word-level criterion better matches performance evaluation criteria such as WER, an improved word recognition performance can be achieved. We evaluated and compared MWCE criterion in a unified DT framework, with other commonly-used criteria including MCE, MMI, MWE, and MPE. Experiments on TIMIT and WS JO evaluation tasks suggest that word-level MWCE criterion can achieve consistently better results than string-level MCE. MWCE even outperforms other substring-level criteria on the above two tasks, including MWE and MPE.
Zhijie Yan, Yu Hu 0003, Renhua Wang
ICASSP1
2008 Soft margin estimation with various separation levels for LVCSR
abstract
We continue our previous work on soft margin estimation (SME) to large vocabulary continuous speech recognition (LVCSR) in two new aspects. The first is to formulate SME with different unit separation. SME methods focusing on string-, word-, and phone-level separation are defined. The second is to compare SME with all the popular conventional discriminative training (DT) methods, including maximum mutual information estimation (MMIE), minimum classification error (MCE), and minimum word/phone error (MWE/MPE). Tested on the 5k-word Wall Street Journal task, all the SME methods achieves a relative word error rate (WER) reduction from 17% to 25% over our baseline. Among them, phone-level SME obtains the best performance. Its performance is slightly better than that of MPE, and much better than those of other conventional DT methods. With the comprehensive comparison with conventional DT methods, SME demonstrates its success on LVCSR tasks.
Jinyu Li 0001, Zhijie Yan, Chin-Hui Lee 0001, Renhua Wang
INTERSPEECH2
2007 A study on soft margin estimation for LVCSR
abstract
We extend our previous work on soft margin estimation (SME) to large vocabulary continuous speech recognition in two aspects. The first is to use the extended Baum-Welch method to replace the conventional generalized probabilistic descent algorithm for optimization. The second is to compare SME with minimum classification error (MCE) training with the same implementation details in order to show that it is indeed the margin component in the objective function with margin-based utterance and frame selection that contributes to the success of SME. Tested on the 5 k-word Wall Street Journal task, all the SME methods work better than MCE. The best SME approach achieves a relative word error rate reduction of about 19% over our best baseline performance. This enhancement can only be demonstrated because of our use of margin-based objective function and the extended Baum-Welch parameter optimization method.
Jinyu Li 0001, Zhijie Yan, Chin-Hui Lee 0001, Renhua Wang
ASRU2
2007 Word Graph Based Feature Enhancement for Noisy Speech Recognition
abstract
This paper presents a word graph based feature enhancement method for robust speech recognition in noise. The approach uses signal processing based speech enhancement as a starting point, and then performs Wiener filtering to remove residual noise. During the process, a decoded word graph is used to directly guide the feature enhancement with respect to the HMM for recognition, so that the enhanced feature can match the clean speech model better in the acoustic space. The proposed word graph based feature enhancement method was tested on the Aurora 2 database. Experimental results show that an improved recognition performance can be obtained comparing with conventional signal processing based and GMM based feature enhancement methods. With signal processing based weighted noise estimation and GMM based method, the relative error rate reductions are 35.44% and 42.58%, respectively. The proposed word graph based method improves the performance further, and a relative error rate reduction of 57.89% is obtained.
Zhijie Yan, Frank K. Soong, Renhua Wang
ICASSP (4)1