Si Wei

dblp:06/7720 · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021
YearPublicationVenuePosition
2026 CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling
abstract
The mismatch between the growing demand for psychological counseling and the limited availability of services has motivated research into the application of Large Language Models (LLMs) in this domain. Consequently, there is a need for a robust and unified benchmark to assess the counseling competence of various LLMs. Existing works, however, are limited by unprofessional client simulation, static question-and-answer evaluation formats, and unidimensional metrics. These limitations hinder their effectiveness in assessing a model's comprehensive ability to handle diverse and complex clients. To address this gap, we introduce CARE-Bench, a dynamic and interactive automated benchmark. It is built upon diverse client profiles derived from real-world counseling cases and simulated according to expert guidelines. CARE-Bench provides a multidimensional performance evaluation grounded in established psychological scales. Using CARE-Bench, we evaluate several general-purpose LLMs and specialized counseling models, revealing their current limitations. In collaboration with psychologists, we conduct a detailed analysis of the reasons for LLMs' failures when interacting with clients of different types, which provides directions for developing more comprehensive, universal, and effective counseling models.
Bichen Wang, Yixin Sun 0002, Hao Yang 0066, Si Wei, Shijin Wang 0001, Bing Qin 0001
AAAI7
2026 WeaveRec: An LLM-Based Cross-Domain Sequential Recommendation Framework with Model Merging
abstract
Cross-Domain Sequential Recommendation (CDSR) seeks to improve user preference modeling by transferring knowledge from multiple domains. Despite the progress made in CDSR, most existing methods rely on overlapping users or items to establish cross-domain correlations-a requirement that rarely holds in real-world settings. The advent of large language models (LLM) and model-merging techniques appears to overcome this limitation by unifying multi-domain data without explicit overlaps. Yet, our empirical study shows that naively training an LLM on combined domains—or simply merging several domain-specific LLMs—often degrades performance relative to a model trained solely on the target domain.
Min Hou 0004, Le Wu 0001, Chenyi He, Hao Liu 0078, Zhi Li 0057, Xin Li 0064, Si Wei
WWW8
2026 Mitigating Fine-tuning Bias: A Parameter-Efficient Debiasing Framework for Large Language Models
Kun Zhang 0015, Le Wu 0001, Hao Liu 0078, Hefei Xu, Xin Li 0064, Si Wei
WWW7
2025 How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
abstract
Pre-trained language models represented by the Transformer have been proven to possess strong base capabilities, and the representative self-attention mechanism in the Transformer has become a classic in sequence modeling architectures. Different from the work of proposing sequence modeling architecture to improve the efficiency of attention mechanism, this work focuses on the impact of sequence modeling architectures on base capabilities. Specifically, our concern is: How exactly do sequence modeling architectures affect the base capabilities of pre-trained language models? In this work, we first point out that the mixed domain pre-training setting commonly adopted in existing architecture design works fails to adequately reveal the differences in base capabilities among various architectures. To address this, we propose a limited domain pre-training setting with out-of-distribution testing, which successfully uncovers significant differences in base capabilities among architectures at an early stage. Next, we analyze the base capabilities of stateful sequence modeling architectures, and find that they exhibit significant degradation in base capabilities compared to the Transformer. Then, through a series of architecture component analysis, we summarize a key architecture design principle: A sequence modeling architecture need possess full-sequence arbitrary selection capability to avoid degradation in base capabilities. Finally, we empirically validate this principle using an extremely simple Top-1 element selection architecture and further generalize it to a more practical Top-1 chunk selection architecture. Experimental results demonstrate our proposed sequence modeling architecture design principle and suggest that our work can serve as a valuable reference for future architecture improvements and novel designs.
Si Wei, Shijin Wang 0001, Bing Qin 0001, Ting Liu 0001
NeurIPS3
2025 Large language models meet text-centric multimodal sentiment analysis: a survey
Hao Yang 0066, Yang Wu 0010, Shilong Wang 0003, Zongyang Ma, Wanxiang Che, Shijin Wang 0001, Si Wei, Bing Qin 0001
Sci. China Inf. Sci.10
2025 From unimodal to multimodal: a framework for generating high-quality multimodal emotional chit-chat dialogue
Hao Yang 0066, Yang Wu 0010, Jianhua Yuan, Wanxiang Che, Shijin Wang 0001, Si Wei, Bing Qin 0001
Sci. China Inf. Sci.8
2025 Optimizing low-rank adaptation with decomposed matrices and adaptive rank allocation
Dacao Zhang, Fan Yang 0063, Kunwu Zhang, Xin Li 0082, Si Wei, Richang Hong, Meng Wang 0001
Frontiers Comput. Sci.5
2024 Maths: Multimodal Transformer-Based Human-Readable Solver
abstract
Multimodal mathematical reasoning has gained increasing attention in recent times. However, previous effective methods have not tried to reason in the form of natural language. In this paper, we introduce a model named MATHS (MultimodAl Transformer-based Human-readable Solver) for visual arithmetic and geometry problems in multimodal mathematical reasoning tasks. Drawing inspiration from Multimodal Large Language Models (MLLMs), our approach involves generating problem-solving processes expressed in natural language, in order to leverage the inherent reasoning capabilities embedded within language models. To address the challenge of precise calculations for language models, our work proposes a Math-Constrained Generation (MCG) method to impose hard constraints on generated outputs. Extensive experiments demonstrate our model excels in visual arithmetic task, and achieves results that are either better or comparable to existing methods in geometry problems. Code is available at https://github.com/ycpNotFound/MATHS.
Yicheng Pan 0004, Jiefeng Ma, Pengfei Hu 0006, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001, Dan Liu 0008, Si Wei
ICME9
2020 Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based Chatbots
abstract
In this paper, we study the problem of employing pre-trained language models for multi-turn response selection in retrieval-based chatbots. A new model, named Speaker-Aware BERT (SA-BERT), is proposed in order to make the model aware of the speaker change information, which is an important and intrinsic property of multi-turn dialogues. Furthermore, a speaker-aware disentanglement strategy is proposed to tackle the entangled dialogues. This strategy selects a small number of most important utterances as the filtered context according to the speakers' information in them. Finally, domain adaptation is performed to incorporate the in-domain knowledge into pre-trained language models. Experiments on five public datasets show that our proposed model outperforms the present models on all metrics by large margins and achieves new state-of-the-art performances for multi-turn response selection.
Jia-Chen Gu, Tianda Li, Quan Liu 0003, Zhen-Hua Ling, Zhiming Su, Si Wei, Xiaodan Zhu 0001
CIKM6
2020 A Tree-Structured Decoder for Image-to-Markup Generation
abstract
Recent encoder-decoder approaches typically employ string decoders to convert images into serialized strings for image-to-markup. However, for tree-structured representational markup, string representations can hardly cope with the structural complexity. In this work, we first show via a set of toy problems that string decoders struggle to decode tree structures, especially as structural complexity increases, we then propose a tree-structured decoder that specifically aims at generating a tree-structured markup. Our decoders works sequentially, where at each step a child node and its parent node are simultaneously generated to form a sub-tree. This sub-tree is consequently used to construct the final tree structure in a recurrent manner. Key to the success of our tree decoder is twofold, (i) it strictly respects the parent-child relationship of trees, and (ii) it explicitly outputs trees as oppose to a linear string. Evaluated on both math formula recognition and chemical formula recognition, the proposed tree decoder is shown to greatly outperform strong string decoder baselines.
Jianshu Zhang 0001, Jun Du 0002, Yongxin Yang, Yi-Zhe Song, Si Wei, Li-Rong Dai 0001
ICML5
2020 End-to-End Transition-Based Online Dialogue Disentanglement
abstract
Dialogue disentanglement aims to separate intermingled messages into detached sessions. The existing research focuses on two-step architectures, in which a model first retrieves the relationships between two messages and then divides the message stream into separate clusters. Almost all existing work puts significant efforts on selecting features for message-pair classification and clustering, while ignoring the semantic coherence within each session. In this paper, we introduce the first end-to- end transition-based model for online dialogue disentanglement. Our model captures the sequential information of each session as the online algorithm proceeds on processing a dialogue. The coherence in a session is hence modeled when messages are sequentially added into their best-matching sessions. Meanwhile, the research field still lacks data for studying end-to-end dialogue disentanglement, so we construct a large-scale dataset by extracting coherent dialogues from online movie scripts. We evaluate our model on both the dataset we developed and the publicly available Ubuntu IRC dataset [Kummerfeld et al., 2019]. The results show that our model significantly outperforms the existing algorithms. Further experiments demonstrate that our model better captures the sequential semantics and obtains more coherent disentangled sessions.
Hui Liu 0033, Jia-Chen Gu, Quan Liu 0003, Si Wei, Xiaodan Zhu 0001
IJCAI5
2018 Exercise-Enhanced Sequential Modeling for Student Performance Prediction
abstract
In online education systems, for offering proactive services to students (e.g., personalized exercise recommendation), a crucial demand is to predict student performance (e.g., scores) on future exercising activities. Existing prediction methods mainly exploit the historical exercising records of students, where each exercise is usually represented as the manually labeled knowledge concepts, and the richer information contained in the text description of exercises is still underexplored. In this paper, we propose a novel Exercise-Enhanced Recurrent Neural Network (EERNN) framework for student performance prediction by taking full advantage of both student exercising records and the text of each exercise. Specifically, for modeling the student exercising process, we first design a bidirectional LSTM to learn each exercise representation from its text description without any expertise and information loss. Then, we propose a new LSTM architecture to trace student states (i.e., knowledge states) in their sequential exercising process with the combination of exercise representations. For making final predictions, we design two strategies under EERNN, i.e., EERNNM with Markov property and EERNNA with Attention mechanism. Extensive experiments on large-scale real-world data clearly demonstrate the effectiveness of EERNN framework. Moreover, by incorporating the exercise correlations, EERNN can well deal with the cold start problems from both student and exercise perspectives.
Yu Su 0002, Qingwen Liu 0002, Qi Liu 0003, Zhenya Huang, Yu Yin 0002, Enhong Chen, Chris Ding, Si Wei
AAAI8
2018 Neural Natural Language Inference Models Enhanced with External Knowledge
abstract
Modeling natural language inference is a very challenging task.With the availability of large annotated data, it has recently become feasible to train complex models such as neural-network-based inference models, which have shown to achieve the state-of-the-art performance.Although there exist relatively large annotated data, can machines learn all knowledge needed to perform natural language inference (NLI) from these data?If not, how can neural-network-based NLI models benefit from external knowledge and how to build NLI models to leverage it?In this paper, we enrich the state-of-the-art neural natural language inference models with external knowledge.We demonstrate that the proposed models improve neural NLI models to achieve the state-of-the-art performance on the SNLI and MultiNLI datasets.
Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Diana Inkpen, Si Wei
ACL (1)5
2017 Question Difficulty Prediction for READING Problems in Standard Tests
abstract
Standard tests aim to evaluate the performance of examinees using different tests with consistent difficulties. Thus, a critical demand is to predict the difficulty of each test question before the test is conducted. Existing studies are usually based on the judgments of education experts (e.g., teachers), which may be subjective and labor intensive. In this paper, we propose a novel Test-aware Attention-based Convolutional Neural Network (TACNN) framework to automatically solve this Question Difficulty Prediction (QDP) task for READING problems (a typical problem style in English tests) in standard tests. Specifically, given the abundant historical test logs and text materials of questions, we first design a CNN-based architecture to extract sentence representations for the questions. Then, we utilize an attention strategy to qualify the difficulty contribution of each sentence to questions. Considering the incomparability of question difficulties in different tests, we propose a test-dependent pairwise strategy for training TACNN and generating the difficulty prediction value. Extensive experiments on a real-world dataset not only show the effectiveness of TACNN, but also give interpretable insights to track the attention information for questions.
Zhenya Huang, Qi Liu 0003, Enhong Chen, Hongke Zhao, Mingyong Gao, Si Wei, Yu Su 0002
AAAI6
2017 Enhanced LSTM for Natural Language Inference
abstract
Reasoning and inference are central to human and artificial intelligence.Modeling inference in human language is very challenging.With the availability of large annotated data (Bowman et al., 2015), it has recently become feasible to train neural network based inference models, which have shown to be very effective.In this paper, we present a new state-of-the-art result, achieving the accuracy of 88.6% on the Stanford Natural Language Inference Dataset.Unlike the previous top models that use very complicated network architectures, we first demonstrate that carefully designing sequential inference models based on chain LSTMs can outperform all previous models.Based on this, we further show that by explicitly considering recursive architectures in both local inference modeling and inference composition, we achieve additional improvement.Particularly, incorporating syntactic parsing information contributes to our best result-it further improves the performance even when added to the already very strong model.
Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Si Wei, Hui Jiang 0001, Diana Inkpen
ACL (1)4
2017 Attention-over-Attention Neural Networks for Reading Comprehension
abstract
Cloze-style queries are representative problems in reading comprehension. Over the past few months, we have seen much progress that utilizing neural network approach to solve Cloze-style questions. In this paper, we present a novel model called attention-over-attention reader for the Cloze-style reading comprehension task. Our model aims to place another attention mechanism over the document-level attention, and induces "attended attention" for final predictions. Unlike the previous works, our neural network model requires less pre-defined hyper-parameters and uses an elegant architecture for modeling. Experimental results show that the proposed attention-over-attention model significantly outperforms various state-of-the-art systems by a large margin in public datasets, such as CNN and Children's Book Test datasets.
Yiming Cui 0001, Zhipeng Chen 0001, Si Wei, Shijin Wang 0001, Ting Liu 0001
ACL (1)3
2017 The iFLYTEK system for blizzard machine learning challenge 2017-ES1
abstract
This paper introduces the speech synthesis system submitted by IFLYTEK for the Blizzard Machine Learning Challenge 2017-ES1. Linguistic and acoustic features from a 4hour corpus were released for this task. Participants are expected to build a speech synthesis system on the given linguist and acoustic features without using any external data. Our system is composed of a long short term memory (LSTM) recurrent neural network (RNN)-based acoustic model and a generative adversarial network (GAN)-based post-filter for mel-cepstra. Two approaches to build GAN-based post-filter are implemented and compared in our experiments. The first one is to predict the residuals of mel-cepstra given the mel-cepstra predicted by the LSTM-based acoustic model. However, this method leads to unstable synthetic speech sounds in our experiments, which may be due to the poor quality of analysis-synthesis speech using the natural acoustic features given by this corpus. The other approach is to ignore the detailed components of natural mel-cepstra by dimension reduction using principal component analysis (PCA) and then recover them back using GAN given the main PCA components. At synthesis time, mel-cepstra predicted by the RNN acoustic model are first projected to the main PCA components, which are then sent to the GAN for detail recovering. Finally, the second approach is used in the final submitted system. The evaluation results show the effectiveness of our submitted system.
Li-Juan Liu, Chuang Ding, Ya-Jun Hu, Zhen-Hua Ling, Yuan Jiang 0006, Si Wei
ASRU7
2017 Cause-Effect Knowledge Acquisition and Neural Association Model for Solving A Set of Winograd Schema Problems
abstract
This paper focuses on the investigations in Winograd Schema (WS), a challenging problem which has been proposed for measuring progress in commonsense reasoning.Due to the lack of commonsense knowledge and training data, very little work has been found on the WS problems in recent years.Actually, there is no shortcut to solve this problem except to collect more commonsense knowledge and design suitable models.Therefore, this paper addresses a set of WS problems by proposing a knowledge acquisition method and a general neural association model.To avoid the sparseness issue, the knowledge we aim to collect is the cause-effect relationships between thousands of commonly used words.The knowledge acquisition method supports us to extract hundreds of thousands of cause-effect pairs from large text corpus automatically.Meanwhile, a neural association model (NAM) is proposed to encode the association relationships between any two discrete events.Based on the extracted knowledge and the NAM models, in this paper, we successfully build a system for solving WS problems from scratch and achieve 70.0% accuracy.Most importantly, this paper provides a flexible framework to solve WS problems based on event association and neural network methods.
Quan Liu 0003, Hui Jiang 0001, Andrew Evdokimov, Zhen-Hua Ling, Xiaodan Zhu 0001, Si Wei, Yu Hu 0003
IJCAI6
2017 Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition
Jianshu Zhang 0001, Jun Du 0002, Shiliang Zhang, Dan Liu 0008, Yulong Hu, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
Pattern Recognit.7
2017 Nonrecurrent Neural Structure for Long-Term Dependence
abstract
In this paper, we propose a novel neural network structure, namely feedforward sequential memory networks (FSMN), to model long-term dependence in time series without using recurrent feedback. The proposed FSMN is a standard fully connected feedforward neural network equipped with some learnable memory blocks in its hidden layers. The memory blocks use a tapped-delay line structure to encode the long context information into a fixed-size representation as short-term memory mechanism which are somehow similar to the time-delay neural networks layers. We have evaluated the FSMNs in several standard benchmark tasks, including speech recognition and language modeling. Experimental results have shown that FSMNs outperform the conventional recurrent neural networks (RNN) while can be learned much more reliably and faster in modeling sequential signals like speech or language. Moreover, we also propose a compact feedforward sequential memory networks (cFSMN) by combining FSMN with low-rank matrix factorization and make a slight modification to the encoding method used in FSMNs in order to further simplify the network architecture. On the speech recognition Switchboard task, the proposed cFSMN structures can reduce the model size by 60% and speed up the learning by more than seven times while the model can still significantly outperform the popular bidirectional LSTMs for both frame-level cross-entropy criterion-based training and MMI-based sequence training.
Shiliang Zhang, Cong Liu 0006, Hui Jiang 0001, Si Wei, Li-Rong Dai 0001, Yu Hu 0003
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Distraction-Based Neural Networks for Modeling Document
Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Si Wei, Hui Jiang 0001
IJCAI4
2016 Future Context Attention for Unidirectional LSTM Based Acoustic Model
Shiliang Zhang, Si Wei, Li-Rong Dai 0001
INTERSPEECH3
2016 Compact Feedforward Sequential Memory Networks for Large Vocabulary Continuous Speech Recognition
Shiliang Zhang, Hui Jiang 0001, Shifu Xiong, Si Wei, Li-Rong Dai 0001
INTERSPEECH4
2015 Revisiting Word Embedding for Contrasting Meaning
abstract
Zhigang Chen, Wei Lin, Qian Chen, Xiaoping Chen, Si Wei, Hui Jiang, Xiaodan Zhu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Zhigang Chen 0003, Qian Chen 0003, Si Wei, Hui Jiang 0001, Xiaodan Zhu 0001
ACL (1)5
2015 Learning Semantic Word Embeddings based on Ordinal Knowledge Constraints
abstract
Quan Liu, Hui Jiang, Si Wei, Zhen-Hua Ling, Yu Hu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Quan Liu 0003, Hui Jiang 0001, Si Wei, Zhen-Hua Ling, Yu Hu 0003
ACL (1)3
2015 Multi-task deep neural network acoustic models with model adaptation using discriminative speaker identity for whisper recognition
abstract
This paper presents a study on large vocabulary continuous whisper automatic recognition (wLVCSR). wLVCSR provides the ability to use ASR equipment in public places without concern for disturbing others or leaking private information. However the task of wLVCSR is much more challenging than normal LVCSR due to the absence of pitch which not only causes the signal to noise ratio (SNR) of whispers to be much lower than normal speech but also leads to flatness and formant shifts in whisper spectra. Furthermore, the amount of whisper data available for training is much less than for normal speech. In this paper, multi-task deep neural network (DNN) acoustic models are deployed to solve these problems. Moreover, model adaptation is performed on the multi-task DNN to normalize speaker and environmental variability in whispers based on discriminative speaker identity information. On a Mandarin whisper dictation task, with 55 hours of whisper data, the proposed SI multi-task DNN model can achieve 56.7% character error rate (CER) improvement over a baseline Gaussian Mixture Model (GMM), discriminatively trained only using the whisper data. Besides, the CER of the proposed model for normal speech can reach 15.2%, which is close to the performance of a state-of-the-art DNN trained with one thousand hours of speech data. From this baseline, the model-adapted DNN gains a further 10.9% CER reduction over the generic model.
Ian McLoughlin 0001, Cong Liu 0006, Shaofei Xue, Si Wei
ICASSP5
2015 Writer adaptive feature extraction based on convolutional neural networks for online handwritten Chinese character recognition
abstract
This paper presents a novel approach to writer adaptation based on convolutional neural network (CNN) as a feature extractor and improved discriminative linear regression for online handwritten Chinese character recognition. First, the proposed recognizer consisting of CNN-based feature extractor and prototype-based classifier can achieve comparable performance with the state-of-the-art CNN-based classifier while it could be designed more compact and efficient as a practical solution. Second, the writer adaption is performed via a linear transformation of the extracted feature from CNN. The transformation parameters are optimized with a so-called sample separation margin based minimum classification error criterion, which can be further improved by using more synthesized adaptation data and a simple regularization method. The experiments on the data collected from user inputs of Smartphones with a vocabulary of 20,936 characters demonstrate that our writer adaptation approach can yield significant improvements of recognition accuracy over a high-performance baseline system and also outperform a state-of-the-art approach based on style transfer mapping especially with increased adaptation data.
Jun Du 0002, Jian-Fang Zhai, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
ICDAR5
2015 Rectified linear neural networks with tied-scalar regularization for LVCSR
Shiliang Zhang, Hui Jiang 0001, Si Wei, Li-Rong Dai 0001
INTERSPEECH3
2014 Lattice based optimization of bottleneck feature extractor with linear transformation
abstract
This paper proposes a lattice-based sequential discriminative training method to extract more discriminative bottleneck features. In our method, the bottleneck neural network is first trained with cross entropy criteria, and then only the weights of bottleneck layer are retrained with sequential criteria. If the outputs of the layer before bottleneck are treated as the raw features, the new method is an equivalent to a linear feature transformation algorithm. This linearity makes the optimization much easier than updating the whole neural network. Just like the fMPE and RDLT, the neural network is retrained with batch mode gradient descent, making the training to be easily implemented in parallel. Meanwhile, batch mode optimization can naturally deal with the indirect gradient to make the optimization more precise. Experimental results on a Mandarin transcription task and the Switchboard task have shown the effectiveness of the proposed method with the CER decreases from 12.2% to 11.3% and the WER from 16.1% to 15.0%, respectively.
Diyuan Liu, Si Wei, Wu Guo, Yebo Bao, Shifu Xiong, Li-Rong Dai 0001
ICASSP2
2014 Writer Adaptation Using Bottleneck Features and Discriminative Linear Regression for Online Handwritten Chinese Character Recognition
abstract
This paper presents a novel approach to writer adaptation using bottleneck features and discriminative linear regression for the recognition of online handwritten Chinese characters. First, bottleneck features extracted from a bottleneck layer of a deep neural network representing a nonlinear and discriminative transformation of the input features are verified to be much more effective in adaptation of writing styles than the conventional features after linear discriminant analysis transformation. Second, discriminative linear regression via a so-called sample separation margin based minimum classification error criterion is adopted for writer adaptation. The experiments on an in-house developed online Chinese handwriting corpus with a vocabulary of 15,167 characters and testing data collected from user inputs of Smartphones show that our proposed approach can achieve very significant improvements of recognition accuracy compared with a state-of-the-art adaptation approach for writer adaptation.
Jun Du 0002, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
ICFHR4
2014 A Study of Designing Compact Classifiers Using Deep Neural Networks for Online Handwritten Chinese Character Recognition
abstract
This paper presents a study of designing compact classifiers using deep neural networks for recognition of online handwritten Chinese characters. Two schemes are investigated based on practical considerations. First, deep neural networks are adopted purely as a classifier with a state-of-the-art feature extractor of online handwritten Chinese characters. Second, the so-called bottleneck features extracted from a bottleneck layer of deep neural networks are fed to the prototype-based classifier. The experiments on an in-house developed online Chinese handwriting corpus with a vocabulary of 15,167 characters show that compared with prototype-based classifier widely developed on the mobile device, deep neural network based classifier can yield significant improvements of recognition accuracy with acceptably increased footprint and latency while the bottleneck-feature approach can bring a more compact classifier with an observable performance gain.
Jun Du 0002, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
ICPR4
2014 Task-aware deep bottleneck features for spoken language identification
abstract
Recently, deep bottleneck features (DBF) extracted from a deep neural network (DNN) containing a narrow bottleneck lay-er, have been applied for language identification (LID), and yield significant performance improvement over state-of-the-art methods on NIST LRE 2009. However, the DNN is trained us-ing a large corpus of specific language which is not directly related to the LID task. More recently, lattice based discrimi-native training methods for extracting more targeted DBF were proposed for ASR. Inspired by this, this paper proposes to tune the post-trained DNN parameters using an LID-specific train-ing corpus, which may make the resulting DBF, termed a Dis-criminative DBF (D2BF), more discriminative and task-aware. Specifically, the maximum mutual information (MMI) criteri-on, with gradient descent, is applied to update the DNN param-eters of the bottleneck layer in an iterative fashion. We evaluate the performance of the proposed D2BF using different back-end models, including GMM-MMI and ivector, over the most con-fused 6-languages selected from NIST LRE 2009. The results show that the proposed D2BF is more appropriate and effective than the original DBF. Index Terms: language identification, deep bottleneck feature, deep neural network, discriminative training, Gaussian mixture model, maximum mutual information 1.
Bing Jiang, Yan Song 0001, Si Wei, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH3
2010 Automatic error detection for unit selection speech synthesis using log likelihood ratio based SVM classifier
abstract
This paper proposes a method to detect the errors in synthetic speech of a unit selection speech synthesis system automatically using log likelihood ratio and support vector machine (SVM). For SVM training, a set of synthetic speech are firstly generated by a given speech synthesis system and their synthetic errors are labeled by manually annotating the segments that sound unnatural. Then, two context-dependent acoustic models are trained using the natural and unnatural segments of labeled synthetic speech respectively. The log likelihood ratio of acoustic features between these two models is adopted to train the SVM classifier for error detection. Experimental results show the proposed method is effective in detecting the errors of pitch contour within a word for a Mandarin speech synthesis system. The proposed SVM method using log likelihood ratio between context-dependent acoustic models outperforms the SVM classifier trained on acoustic features directly.
Heng Lu 0002, Zhen-Hua Ling, Si Wei, Li-Rong Dai 0001, Renhua Wang
INTERSPEECH3
2009 A new method for mispronunciation detection using Support Vector Machine based on Pronunciation Space Models
Si Wei, Yu Hu 0003, Renhua Wang
Speech Commun.1
2007 CDF-Matching for Automatic Tone Error Detection in Mandarin Call System
abstract
This paper introduces a tone pronunciation error detection algorithm for Mandarin CALL system. HMM is introduced to build tone model. FO after CDF-matching normalization is used as the feature of tone model. Tone error detection algorithm is based on the posterior probability calculated from HMM tone models. Comparing to the other normalization methods, CDF-matching can get better tone error detection performance with CC's increasing from 0.76 to 0.79, which is close to that between human evaluators (0.83).
Si Wei, Hai-Kun Wang, Qing-Sheng Liu, Renhua Wang
ICASSP (4)1
2006 Automatic Mandarin pronunciation scoring for native learners with dialect accent
Si Wei, Qing-Sheng Liu, Yu Hu 0003, Renhua Wang
INTERSPEECH1