EDBT 2026 Demo / reviewers in the wild / expert
Yerbolat Khassanov
dblp:142/3161
· DBLP profile ↗
18ranked-venue papers
9as first author
10since 2021 · last 2024
0000-0001-9422-6833ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Extending Multilingual ASR to New Languages Using Supplementary Encoder and Decoder ComponentsabstractExtending multilingual automatic speech recognition (mASR) systems to new languages poses challenges, particularly when training data for existing languages is limited or unavailable. To tackle this issue, we suggest utilizing supplementary encoder and decoder components. Specifically, we propose appending and fine-tuning a distinct decoder designed for new languages, while preserving the parameters of existing languages to minimize disruption to their performance. Furthermore, we advocate attaching an additional encoder component to enhance acoustic representation learning for new languages, resulting in substantial improvements in word error rate performance. Our experimental findings demonstrate the effectiveness of the proposed methods for the task of extending language support within mASR systems. Yerbolat Khassanov, Tianfeng Chen, Tze Yuang Chong, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001 |
ICASSP | 1 |
| 2024 | Dual-Pipeline with Low-Rank Adaptation for New Language Integration in Multilingual ASR
Yerbolat Khassanov, Tianfeng Chen, Tze Yuang Chong |
INTERSPEECH | 1 |
| 2023 | Knowledge Distillation Approach for Efficient Internal Language Model Estimation
Haihua Xu 0001, Yerbolat Khassanov, Lu Lu 0015, Zejun Ma 0001, Ji Wu 0002 |
INTERSPEECH | 3 |
| 2023 | Random Utterance Concatenation Based Data Augmentation for Improving Short-video Speech RecognitionabstractOne of limitations in end-to-end automatic speech recognition (ASR) framework is its performance would be compromised if train-test utterance lengths are mismatched.In this paper, we propose an on-the-fly random utterance concatenation (RUC) based data augmentation method to alleviate train-test utterance length mismatch issue for short-video ASR task.Specifically, we are motivated by observations that our human-transcribed training utterances tend to be much shorter for short-video spontaneous speech (∼3 seconds on average), while our test utterance generated from voice activity detection front-end is much longer (∼10 seconds on average).Such a mismatch can lead to suboptimal performance.Empirically, it's observed the proposed RUC method significantly improves long utterance recognition without performance drop on short one.Overall, it achieves 5.72% word error rate reduction on average for 15 languages and improved robustness to various utterance length. Yist Y. Lin, Haihua Xu 0001, Van Tung Pham, Yerbolat Khassanov, Tze Yuang Chong, Lu Lu 0015, Zejun Ma 0001 |
INTERSPEECH | 5 |
| 2023 | Multilingual Text-to-Speech Synthesis for Turkic Languages Using TransliterationabstractThis work aims to build a multilingual text-to-speech (TTS) synthesis system for ten lower-resourced Turkic languages: Azerbaijani, Bashkir, Kazakh, Kyrgyz, Sakha, Tatar, Turkish, Turkmen, Uyghur, and Uzbek. We specifically target the zero-shot learning scenario, where a TTS model trained using the data of one language is applied to synthesise speech for other, unseen languages. An end-to-end TTS system based on the Tacotron 2 architecture was trained using only the available data of the Kazakh language. To generate speech for the other Turkic languages, we first mapped the letters of the Turkic alphabets onto the symbols of the International Phonetic Alphabet (IPA), which were then converted to the Kazakh alphabet letters. To demon strate the feasibility of the proposed approach, we evaluated the multilingual Turkic TTS model subjectively and obtained promising results. To enable replication of the experiments, we make our code and dataset publicly available in our GitHub repository. Rustem Yeshpanov, Saida Mussakhojayeva, Yerbolat Khassanov |
INTERSPEECH | 3 |
| 2022 | KSC2: An Industrial-Scale Open-Source Kazakh Speech CorpusabstractWe present the first industrial-scale open-source Kazakh speech corpus for automatic speech recognition research and development. Our corpus subsumes two previously presented corpora: 1) Kazakh speech corpus (KSC) and 2) Kazakh text-to-speech 2 (KazakhTTS2). We also provide additional data from other sources, including television news, television and radio programs, parliament speeches, and podcasts. Our corpus, which we have named KSC2, contains over a thousand hours of high-quality transcribed data, which is triple the size of KSC. KSC2 was manually transcribed with the help of native Kazakh speakers and validated via preliminary speech recognition experiments on various evaluation sets. Moreover, it contains utterances with Kazakh-Russian code-switching, a conversational practice common among Kazakh speakers. We believe that our corpus will facilitate speech processing research for Kazakh, which is widely considered an under-resourced language. To ensure the reproducibility of experiments, we share the KSC2 corpus, training recipes, and pretrained models. Saida Mussakhojayeva, Yerbolat Khassanov, Huseyin Atakan Varol |
INTERSPEECH | 2 |
| 2022 | KazakhTTS2: Extending the Open-Source Kazakh TTS Corpus With More Data, Speakers, and TopicsabstractWe present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two to five (three females and two males), and the topic coverage has been diversified with the help of new sources, including a book and Wikipedia articles. This corpus is necessary for building high-quality TTS systems for Kazakh, a Central Asian agglutinative language from the Turkic family, which presents several linguistic challenges. We describe the corpus construction process and provide the details of the training and evaluation procedures for the TTS system. Our experimental results indicate that the constructed corpus is sufficient to build robust TTS models for real-world applications, with a subjective mean opinion score ranging from 3.6 to 4.2 for all the five speakers. We believe that our corpus will facilitate speech and language research for Kazakh and other Turkic languages, which are widely considered to be low-resource due to the limited availability of free linguistic data. The constructed corpus, code, and pretrained models are publicly available in our GitHub repository. Saida Mussakhojayeva, Yerbolat Khassanov, Huseyin Atakan Varol |
LREC | 2 |
| 2022 | KazNERD: Kazakh Named Entity Recognition DatasetabstractWe present the development of a dataset for Kazakh named entity recognition. The dataset was built as there is a clear need for publicly available annotated corpora in Kazakh, as well as annotation guidelines containing straightforward—but rigorous—rules and examples. The dataset annotation, based on the IOB2 scheme, was carried out on television news text by two native Kazakh speakers under the supervision of the first author. The resulting dataset contains 112,702 sentences and 136,333 annotations for 25 entity classes. State-of-the-art machine learning models to automatise Kazakh named entity recognition were also built, with the best-performing model achieving an exact match F1-score of 97.22% on the test set. The annotated dataset, guidelines, and codes used to train the models are freely available for download under the CC BY 4.0 licence from https://github.com/IS2AI/KazNERD. Rustem Yeshpanov, Yerbolat Khassanov, Huseyin Atakan Varol |
LREC | 2 |
| 2021 | A Crowdsourced Open-Source Kazakh Speech Corpus and Initial Speech Recognition BaselineabstractYerbolat Khassanov, Saida Mussakhojayeva, Almas Mirzakhmetov, Alen Adiyev, Mukhamet Nurpeiissov, Huseyin Atakan Varol. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Yerbolat Khassanov, Saida Mussakhojayeva, Almas Mirzakhmetov, Alen Adiyev, Mukhamet Nurpeiissov, Huseyin Atakan Varol |
EACL | 1 |
| 2021 | KazakhTTS: An Open-Source Kazakh Text-to-Speech Synthesis DatasetabstractThis paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two professional speakers (female and male). It is the first publicly available large-scale dataset developed to promote Kazakh text-to-speech (TTS) applications in both academia and industry. In this paper, we share our experience by describing the dataset development procedures and faced challenges, and discuss important future directions. To demonstrate the reliability of our dataset, we built baseline end-to-end TTS models and evaluated them using the subjective mean opinion score (MOS) measure. Evaluation results show that the best TTS models trained on our dataset achieve MOS above 4 for both speakers, which makes them applicable for practical use. The dataset, training recipe, and pretrained TTS models are freely available. Saida Mussakhojayeva, Aigerim Janaliyeva, Almas Mirzakhmetov, Yerbolat Khassanov, Huseyin Atakan Varol |
Interspeech | 4 |
| 2020 | Independent Language Modeling Architecture for End-To-End ASRabstractThe attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network. However, in a vanilla E2E ASR architecture, the decoder sub-network (subnet), which incorporates the role of the language model (LM), is conditioned on the encoder output. This means that the acoustic encoder and the language model are entangled that doesn’t allow language model to be trained separately from external text data. To address this problem, in this work, we propose a new architecture that separates the decoder subnet from the encoder output. In this way, the decoupled subnet becomes an independently trainable LM subnet, which can easily be updated using the external text data. We study two strategies for updating the new architecture. Experimental results show that, 1) the independent LM architecture benefits from external text data, achieving 9.3% and 22.8% relative character and word error rate reduction on Mandarin HKUST and English NSC datasets respectively; 2) the proposed architecture works well with external LM and can be generalized to different amount of labelled data. Van Tung Pham, Haihua Xu 0001, Yerbolat Khassanov, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 3 |
| 2019 | Constrained Output Embeddings for End-to-End Code-Switching Speech Recognition with Only Monolingual DataabstractThe lack of code-switch training data is one of the major concerns in the development of end-to-end code-switching automatic speech recognition (ASR) models. In this work, we propose a method to train an improved end-to-end code-switching ASR using only monolingual data. Our method encourages the distributions of output token embeddings of monolingual languages to be similar, and hence, promotes the ASR model to easily code-switch between languages. Specifically, we propose to use Jensen-Shannon divergence and cosine distance based constraints. The former will enforce output embeddings of monolingual languages to possess similar distributions, while the later simply brings the centroids of two distributions to be close to each other. Experimental results demonstrate high effectiveness of the proposed method, yielding up to 4.5% absolute mixed error rate improvement on Mandarin-English code-switching ASR task. Yerbolat Khassanov, Haihua Xu 0001, Van Tung Pham, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001 |
INTERSPEECH | 1 |
| 2019 | Enriching Rare Word Representations in Neural Language Models by Embedding Matrix AugmentationabstractThe neural language models (NLM) achieve strong generalization capability by learning the dense representation of words and using them to estimate probability distribution function. However, learning the representation of rare words is a challenging problem causing the NLM to produce unreliable probability estimates. To address this problem, we propose a method to enrich representations of rare words in pre-trained NLM and consequently improve its probability estimation performance. The proposed method augments the word embedding matrices of pre-trained NLM while keeping other parameters unchanged. Specifically, our method updates the embedding vectors of rare words using embedding vectors of other semantically and syntactically similar words. To evaluate the proposed method, we enrich the rare street names in the pre-trained NLM and use it to rescore 100-best hypotheses output from the Singapore English speech recognition system. The enriched NLM reduces the word error rate by 6% relative and improves the recognition accuracy of the rare words by 16% absolute as compared to the baseline NLM. Yerbolat Khassanov, Zhiping Zeng, Van Tung Pham, Haihua Xu 0001, Chng Eng Siong |
INTERSPEECH | 1 |
| 2019 | On the End-to-End Solution to Mandarin-English Code-Switching Speech RecognitionabstractCode-switching (CS) refers to a linguistic phenomenon where a speaker uses different languages in an utterance or between alternating utterances.In this work, we study end-to-end (E2E) approaches to the Mandarin-English code-switching speech recognition task.We first examine the effectiveness of using data augmentation and byte-pair encoding (BPE) subword units.More importantly, we propose a multitask learning recipe, where a language identification task is explicitly learned in addition to the E2E speech recognition task.Furthermore, we introduce an efficient word vocabulary expansion method for language modeling to alleviate data sparsity issues under the code-switching scenario.Experimental results on the SEAME data, a Mandarin-English code-switching corpus, demonstrate the effectiveness of the proposed methods. Zhiping Zeng, Yerbolat Khassanov, Van Tung Pham, Haihua Xu 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2018 | Unsupervised and Efficient Vocabulary Expansion for Recurrent Neural Network Language Models in ASRabstractIn automatic speech recognition (ASR) systems, recurrent neural network language models (RNNLM) are used to rescore a word lattice or N-best hypotheses list. Due to the expensive training, the RNNLM's vocabulary set accommodates only small shortlist of most frequent words. This leads to suboptimal performance if an input speech contains many out-of-shortlist (OOS) words. An effective solution is to increase the shortlist size and retrain the entire network which is highly inefficient. Therefore, we propose an efficient method to expand the shortlist set of a pretrained RNNLM without incurring expensive retraining and using additional training data. Our method exploits the structure of RNNLM which can be decoupled into three parts: input projection layer, middle layers, and output projection layer. Specifically, our method expands the word embedding matrices in projection layers and keeps the middle layers unchanged. In this approach, the functionality of the pretrained RNNLM will be correctly maintained as long as OOS words are properly modeled in two embedding spaces. We propose to model the OOS words by borrowing linguistic knowledge from appropriate in-shortlist words. Additionally, we propose to generate the list of OOS words to expand vocabulary in unsupervised manner by automatically extracting them from ASR output. Yerbolat Khassanov, Chng Eng Siong |
INTERSPEECH | 1 |
| 2017 | Unsupervised Language Model Adaptation by Data Selection for Speech Recognition
Yerbolat Khassanov, Tze Yuang Chong, Benjamin Bigot, Chng Eng Siong |
ACIIDS (1) | 1 |
| 2014 | Inertial motion capture based reference trajectory generation for a mobile manipulatorabstractThis paper presents a human-robot interaction system integrating a mobile manipulator with a full-body inertial motion capture suit. The framework aims to provide an intuitive, effective and easy-to-deploy teleoperation interface. User control intent is acquired by the motion capture system in real-time, processed and relayed to the mobile manipulator wirelessly. Specifically, body center of mass and the right hand kinematic data of the user are used to generate position and orientation references for the robot base and manipulator, respectively. The left arm is employed to provide high-level user commands such as "Manipulator On/Off", "Base On/Off" and "Manipulator Pause/Resume". The efficacy of the presented system was demonstrated in real-time teleoperation experiments using KUKA youBot mobile manipulator accomplishing pick-and-place tasks. Yerbolat Khassanov, Nursultan Imanberdiyev, Huseyin Atakan Varol |
HRI | 1 |
| 2014 | Real-time gesture recognition for the high-level teleoperation interface of a mobile manipulatorabstractThis paper describes an inertial motion capture based arm gesture recognition system for the high-level control of a mobile manipulator. Left arm kinematic data of the user is acquired by an inertial motion capture system (Xsens MVN) in real-time and processed to extract supervisory user interface commands such as "Manipulator On/Off", "Base On/Off" and "Operation Pause/Resume" for a mobile manipulator system (KUKA youBot). Principal Component Analysis and Linear Discriminant Analysis are employed for dimension reduction and classification of the user kinematic data, respectively. The classification accuracy for the six class gesture recognition problem is 95.6 percent. In order to increase the reliability of the gesture recognition framework in real-time operation, a consensus voting scheme involving the last ten classification results is implemented. During the five-minute long teleoperation experiment, a total of 25 high-level commands were recognized correctly by the consensus voting enhanced gesture recognizer. The experimental subject stated that the user interface was easy to learn and did not require extensive mental effort to operate. Yerbolat Khassanov, Nursultan Imanberdiyev, Huseyin Atakan Varol |
HRI | 1 |