VLDB 2026 Research / reviewers in the wild / expert
Ekapol Chuangsuwanich
dblp:26/10649
· DBLP profile ↗
37ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0001-6104-4857ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?abstractWuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng, Jann Railey Montalan, Thura Aung, Jian Gang Ngui, Yosephine Susanto, William Chandra Tjhi, Panuthep Tasawong, Erik Cambria, Ekapol Chuangsuwanich, Sarana Nutanong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng, Jann Railey Montalan, Thura Aung, Jian Gang Ngui, Yosephine Susanto, William-Chandra Tjhi, Panuthep Tasawong, Erik Cambria, Ekapol Chuangsuwanich, Sarana Nutanong |
ACL (1) | 11 |
| 2026 | NPC: Automated Tool for Detecting and Explaining ChatGPT-Generated Programs
Pachanitha Saeheng, Napat Boongaree, Chutweeraya Sriwilailak, Chaiyong Ragkhitwetsagul, Teeradaj Racharak, Ekapol Chuangsuwanich |
ICAART (5) | 6 |
| 2025 | Explainable Depression Detection using Masked Hard Instance Mining
Patawee Prakrankamanant, Shinji Watanabe 0001, Ekapol Chuangsuwanich |
INTERSPEECH | 3 |
| 2025 | Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
Thanapat Trachu, Thanathai Lertpetchpun, Ekapol Chuangsuwanich |
INTERSPEECH | 3 |
| 2025 | Thai Speech Spoofing Detection Dataset with Variations in Speaking Styles
Ticho Urai, Pachara Boonsarngsuk, Ekapol Chuangsuwanich |
INTERSPEECH | 3 |
| 2025 | Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations
Supawich Sitdhipol, Waritwong Sukprasongdee, Ekapol Chuangsuwanich, Rina Tse |
SMC | 3 |
| 2024 | An Empirical Study of Multilingual Reasoning Distillation for Question AnsweringabstractPatomporn Payoungkhamdee, Peerat Limkonchotiwat, Jinheon Baek, Potsawee Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Patomporn Payoungkhamdee, Peerat Limkonchotiwat, Jinheon Baek, P. P. Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong |
EMNLP | 6 |
| 2024 | Efficient Overshadowed Entity Disambiguation by Mitigating Shortcut LearningabstractPanuthep Tasawong, Peerat Limkonchotiwat, Potsawee Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Panuthep Tasawong, Peerat Limkonchotiwat, P. P. Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong |
EMNLP | 5 |
| 2024 | Attribute-Aware Amplification of Facial Feature Sequences for Facial Emotion RecognitionabstractMany works have been proposed to automatically recognize facial human emotions, yet distinguishing subtle emotions still remains a challenge. However, just amplifying these facial movements as a whole does not accurately reflect the actual expressions as action intensity is different for each facial part. We propose an attributionaware amplification of facial descriptors that considers gender, facial components, and emotions to alleviate this issue. The amplifications are also continuously updated in an evolutionary manner during training using Population Based Augmentation (PBA). We found that applying different amplifications for each facial region/unit is crucial for the success of our method, outperforming the competing approaches on both the RAVDESS and THAI-SER datasets. Tagon Sompong, Chawan Piansaddhayanon, Ekapol Chuangsuwanich |
ICASSP | 3 |
| 2024 | Thunder : Unified Regression-Diffusion Speech Enhancement with a Single Reverse Step using Brownian Bridge
Thanapat Trachu, Chawan Piansaddhayanon, Ekapol Chuangsuwanich |
INTERSPEECH | 3 |
| 2024 | MHCSeqNet2 - improved peptide-class I MHC binding prediction for alleles with low dataabstractMOTIVATION: The binding of a peptide antigen to a Class I major histocompatibility complex (MHC) protein is part of a key process that lets the immune system recognize an infected cell or a cancer cell. This mechanism enabled the development of peptide-based vaccines that can activate the patient's immune response to treat cancers. Hence, the ability of accurately predict peptide-MHC binding is an essential component for prioritizing the best peptides for each patient. However, peptide-MHC binding experimental data for many MHC alleles are still lacking, which limited the accuracy of existing prediction models. RESULTS: In this study, we presented an improved version of MHCSeqNet that utilized sub-word-level peptide features, a 3D structure embedding for MHC alleles, and an expanded training dataset to achieve better generalizability on MHC alleles with small amounts of data. Visualization of MHC allele embeddings confirms that the model was able to group alleles with similar binding specificity, including those with no peptide ligand in the training dataset. Furthermore, an external evaluation suggests that MHCSeqNet2 can improve the prioritization of T cell epitopes for MHC alleles with small amount of training data. AVAILABILITY AND IMPLEMENTATION: The source code and installation instruction for MHCSeqNet2 are available at https://github.com/cmb-chula/MHCSeqNet2. Patiphan Wongklaew, Sira Sriswasdi, Ekapol Chuangsuwanich |
Bioinform. | 3 |
| 2023 | Thai-Dialect: Low Resource Thai Dialectal Speech to Text CorporaabstractWe release a speech-to-text benchmark dataset containing 10 Thai dialects that cover different regions of Thailand. Our corpora consists of the standard dialect, Thai-central (THA); the northern dialects (Khummuang (NOD), Nan (KHB) and Yno (YNO)); the northeastern dialects (Korat (TTS), Khmer (KXM) and Laos (TTS)); and the southern dialects (Krabi (SOU), Pattani (MFA) and Phangnga (SOU)). All transcriptions are based on the Thai writing standard. We constructed baseline models by fine-tuning from self-supervised pre-trained models. Results show that multilingual/multidialectal systems outperform monolingual ones, and different dialect combinations can affect the performance of multilingual/multidialectal training. Artit Suwanbandit, Jaturong Chitiyaphol, Sutthinan Chuenchom, Kanyarat Kwiecien, Husen Sawal, Ruslan Uthai, Orathai Sangpetch, Ekapol Chuangsuwanich |
ASRU | 8 |
| 2023 | Zero-guidance Segmentation Using Zero Segment LabelsabstractThe joint visual-language model CLIP has enabled new and exciting applications, such as open-vocabulary segmentation, which can locate any segment given an arbitrary text query. In our research, we ask whether it is possible to discover semantic segments without any user guidance in the form of text queries or predefined classes, and label them using natural language automatically? We propose a novel problem zero-guidance segmentation and the first baseline that leverages two pre-trained generalist models, DINO and CLIP, to solve this problem without any fine-tuning or segmentation dataset. The general idea is to first segment an image into small over-segments, encode them into CLIP’s visual-language space, translate them into text labels, and merge semantically similar segments together. The key challenge, however, is how to encode a visual segment into a segment-specific embedding that balances global and local context information, both useful for recognition. Our main contribution is a novel attention-masking technique that balances the two contexts by analyzing the attention layers inside CLIP. We also introduce several metrics for the evaluation of this new task. With CLIP’s innate knowledge, our method can precisely locate the Mona Lisa painting among a museum crowd (Figure 1). More results are available at https://zero-guide-seg.github.io/. Pitchaporn Rewatbowornwong, Nattanat Chatthee, Ekapol Chuangsuwanich, Supasorn Suwajanakorn |
ICCV | 3 |
| 2023 | Instance-based Temporal Normalization for Speaker Verification
Thanathai Lertpetchpun, Ekapol Chuangsuwanich |
INTERSPEECH | 2 |
| 2023 | Word-level Confidence Estimation for CTC Models
Burin Naowarat, Thananchai Kongthaworn, Ekapol Chuangsuwanich |
INTERSPEECH | 3 |
| 2023 | Crowdsourced Data Validation for ASR Training
Wannaphong Phatthiyaphaibun, Chompakorn Chaksangchaichot, Thanawin Rakthanmanon, Ekapol Chuangsuwanich, Sarana Nutanong |
INTERSPEECH | 4 |
| 2023 | Thai Dialect Corpus and Transfer-based Curriculum Learning Investigation for Dialect Automatic Speech Recognition
Artit Suwanbandit, Burin Naowarat, Orathai Sangpetch, Ekapol Chuangsuwanich |
INTERSPEECH | 4 |
| 2023 | ReCasNet: Improving consistency within the two-stage mitosis detection framework
Chawan Piansaddhayanon, Sakun Santisukwongchote, Shanop Shuangshoti, Qingyi Tao, Sira Sriswasdi, Ekapol Chuangsuwanich |
Artif. Intell. Medicine | 6 |
| 2023 | An Efficient Self-Supervised Cross-View Training For Sentence EmbeddingabstractAbstract Self-supervised sentence representation learning is the task of constructing an embedding space for sentences without relying on human annotation efforts. One straightforward approach is to finetune a pretrained language model (PLM) with a representation learning method such as contrastive learning. While this approach achieves impressive performance on larger PLMs, the performance rapidly degrades as the number of parameters decreases. In this paper, we propose a framework called Self-supervised Cross-View Training (SCT) to narrow the performance gap between large and small PLMs. To evaluate the effectiveness of SCT, we compare it to 5 baseline and state-of-the-art competitors on seven Semantic Textual Similarity (STS) benchmarks using 5 PLMs with the number of parameters ranging from 4M to 340M. The experimental results show that STC outperforms the competitors for PLMs with less than 100M parameters in 18 of 21 cases.1 Peerat Limkonchotiwat, Wuttikorn Ponwitayarat, Lalita Lowphansirikul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | Mitigating Spurious Correlation in Natural Language Understanding with Counterfactual InferenceabstractCan Udomcharoenchaikit, Wuttikorn Ponwitayarat, Patomporn Payoungkhamdee, Kanruethai Masuk, Weerayut Buaphet, Ekapol Chuangsuwanich, Sarana Nutanong. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Can Udomcharoenchaikit, Wuttikorn Ponwitayarat, Patomporn Payoungkhamdee, Kanruethai Masuk, Weerayut Buaphet, Ekapol Chuangsuwanich, Sarana Nutanong |
EMNLP | 6 |
| 2021 | Reducing Spelling Inconsistencies in Code-Switching ASR Using Contextualized CTC LossabstractCode-Switching (CS) remains a challenge for Automatic Speech Recognition (ASR), especially character-based models. With the combined choice of characters from multiple languages, the out-come from character-based models suffers from phoneme duplication, resulting in language-inconsistent spellings. We propose Contextualized Connectionist Temporal Classification (CCTC) loss to encourage spelling consistencies of a character-based non-autoregressive ASR which allows for faster inference. The model trained by CCTC loss is aware of contexts since it learns to predict both center and surrounding letters in a multi-task manner. In contrast to existing CTC-based approaches, CCTC loss does not require frame-level alignments, since the context ground truth is obtained from the model’s estimated path. Compared to the same model trained with regular CTC loss, our method consistently improved the ASR performance on both CS and monolingual corpora. Burin Naowarat, Thananchai Kongthaworn, Korrawe Karunratanakul, Sheng Hui Wu, Ekapol Chuangsuwanich |
ICASSP | 5 |
| 2021 | Spectral and Latent Speech Representation Distortion for TTS Evaluation
Thananchai Kongthaworn, Burin Naowarat, Ekapol Chuangsuwanich |
Interspeech | 3 |
| 2021 | Set Prediction in the Latent SpaceabstractSet prediction tasks require the matching between predicted set and ground truth set in order to propagate the gradient signal. Recent works have performed this matching in the original feature space thus requiring predefined distance functions. We propose a method for learning the distance function by performing the matching in the latent space learned from encoding networks. This method enables the use of teacher forcing which was not possible previously since matching in the feature space must be computed after the entire output sequence is generated. Nonetheless, a naive implementation of latent set prediction might not converge due to permutation instability. To address this problem, we provide sufficient conditions for permutation stability which begets an algorithm to improve the overall model convergence. Experiments on several set prediction tasks, including image captioning and object detection, demonstrate the effectiveness of our method. Konpat Preechakul, Chawan Piansaddhayanon, Burin Naowarat, Tirasan Khandhawit, Sira Sriswasdi, Ekapol Chuangsuwanich |
NeurIPS | 6 |
| 2021 | MetaSleepLearner: A Pilot Study on Fast Adaptation of Bio-Signals-Based Sleep Stage Classifier to New Individual Subject Using Meta-LearningabstractIdentifying bio-signals based-sleep stages requires time-consuming and tedious labor of skilled clinicians. Deep learning approaches have been introduced in order to challenge the automatic sleep stage classification conundrum. However, the difficulties can be posed in replacing the clinicians with the automatic system due to the differences in many aspects found in individual bio-signals, causing the inconsistency in the performance of the model on every incoming individual. Thus, we aim to explore the feasibility of using a novel approach, capable of assisting the clinicians and lessening the workload. We propose the transfer learning framework, entitled MetaSleepLearner, based on Model Agnostic Meta-Learning (MAML), in order to transfer the acquired sleep staging knowledge from a large dataset to new individual subjects (source code is available at https://github.com/IoBT-VISTEC/MetaSleepLearner). The framework was demonstrated to require the labelling of only a few sleep epochs by the clinicians and allow the remainder to be handled by the system. Layer-wise Relevance Propagation (LRP) was also applied to understand the learning course of our approach. In all acquired datasets, in comparison to the conventional approach, MetaSleepLearner achieved a range of 5.4% to 17.7% improvement with statistical difference in the mean of both approaches. The illustration of the model interpretation after the adaptation to each subject also confirmed that the performance was directed towards reasonable learning. MetaSleepLearner outperformed the conventional approaches as a result from the fine-tuning using the recordings of both healthy subjects and patients. This is the first work that investigated a non-conventional pre-training method, MAML, resulting in a possibility for human-machine collaboration in sleep stage classification and easing the burden of the clinicians in labelling the sleep stages through only several epochs rather than an entire recording. Nannapas Banluesombatkul, Pichayoot Ouppaphan, Pitshaporn Leelaarporn, Payongkit Lakhan, Busarakum Chaitusaney, Nattapong Jaimchariyatam, Ekapol Chuangsuwanich, Wei Chen 0015, Huy Phan, Nat Dilokthanakul, Theerawit Wilaiprasitporn |
IEEE J. Biomed. Health Informatics | 7 |
| 2020 | Domain Adaptation of Thai Word Segmentation Models using Stacked EnsembleabstractPeerat Limkonchotiwat, Wannaphong Phatthiyaphaibun, Raheem Sarwar, Ekapol Chuangsuwanich, Sarana Nutanong. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Peerat Limkonchotiwat, Wannaphong Phatthiyaphaibun, Raheem Sarwar, Ekapol Chuangsuwanich, Sarana Nutanong |
EMNLP (1) | 4 |
| 2020 | A Comparative Study of Pretrained Language Models for Automated Essay Scoring with Adversarial InputsabstractAutomated Essay Scoring (AES) is a task that deals with grading written essays automatically without human intervention. This study compares the performance of three AES models which utilize different text embedding methods, namely Global Vectors for Word Representation (GloVe), Embeddings from Language Models (ELMo), and Bidirectional Encoder Representations from Transformers (BERT). We used two evaluation metrics: Quadratic Weighted Kappa (QWK) and a novel "robustness", which quantifies the models' ability to detect adversarial essays created by modifying normal essays to cause them to be less coherent. We found that: (1) the BERT-based model achieved the greatest robustness, followed by the GloVe-based and ELMo-based models, respectively, and (2) fine-tuning the embeddings improves QWK but lowers robustness. These findings could be informative on how to choose, and whether to fine-tune, an appropriate model based on how much the AES program places emphasis on proper grading of adversarial essays. Phakawat Wangkriangkri, Chanissara Viboonlarp, Attapol Rutherford, Ekapol Chuangsuwanich |
TENCON | 4 |
| 2019 | MHCSeqNet: a deep neural network model for universal MHC binding predictionabstractBACKGROUND: Immunotherapy is an emerging approach in cancer treatment that activates the host immune system to destroy cancer cells expressing unique peptide signatures (neoepitopes). Administrations of cancer-specific neoepitopes in the form of synthetic peptide vaccine have been proven effective in both mouse models and human patients. Because only a tiny fraction of cancer-specific neoepitopes actually elicits immune response, selection of potent, immunogenic neoepitopes remains a challenging step in cancer vaccine development. A basic approach for immunogenicity prediction is based on the premise that effective neoepitope should bind with the Major Histocompatibility Complex (MHC) with high affinity. RESULTS: In this study, we developed MHCSeqNet, an open-source deep learning model, which not only outperforms state-of-the-art predictors on both MHC binding affinity and MHC ligand peptidome datasets but also exhibits promising generalization to unseen MHC class I alleles. MHCSeqNet employed neural network architectures developed for natural language processing to model amino acid sequence representations of MHC allele and epitope peptide as sentences with amino acids as individual words. This consideration allows MHCSeqNet to accept new MHC alleles as well as peptides of any length. CONCLUSIONS: The improved performance and the flexibility offered by MHCSeqNet should make it a valuable tool for screening effective neoepitopes in cancer vaccine development. Poomarin Phloyphisut, Natapol Pornputtapong, Sira Sriswasdi, Ekapol Chuangsuwanich |
BMC Bioinform. | 4 |
| 2018 | Towards Asynchronous Motor Imagery-Based Brain-Computer Interfaces: a joint training scheme using deep learningabstractIn this paper, the deep learning (DL) approach is applied to a joint training scheme for asynchronous motor imagery-based Brain-Computer Interface (BCI). The proposed DL approach is a cascade of one-dimensional convolutional neural networks and fully-connected neural networks (CNN-FC). The focus is mainly on three types of brain responses: non-imagery EEG (background EEG), (pure imagery) EEG, and EEG during the transitional period between background EEG and pure imagery (transitional imagery). The study of transitional imagery signals should provide greater insight into real-world scenarios. It may be inferred that pure imagery and transitional EEG are high and low power EEG imagery, respectively. Moreover, the results from the CNN-FC are compared to the conventional approach for motor imagery-BCI, namely the common spatial pattern (CSP) for feature extraction and support vector machine (SVM) for classification (CSP-SVM). Under a joint training scheme, pure and transitional imagery are treated as the same class, while background EEG is another class. Ten-fold cross-validation is used to evaluate whether the joint training scheme significantly improves the performance task of classifying pure and transitional imagery signals from background EEG. Using sparse of just a few electrode channels (Cz, C3and C4), mean accuracy reaches 71.52% and 70.27% for CNN-FC and CSP-SVM, respectively. On the other hand, mean accuracy without the joint training scheme achieve only 62.68% and 52.41% for CNN-FC and CSP-SVM, respectively. Patcharin Cheng, Phairot Autthasan, Boriwat Pijarana, Ekapol Chuangsuwanich, Theerawit Wilaiprasitporn |
TENCON | 4 |
| 2016 | Multilingual data selection for training stacked bottleneck featuresabstractDeep Neural Networks (DNNs) trained on multilingual data have proven useful for improving speech recognition in languages with limited resources. In this framework, data from rich resource languages are pooled together to train a single system and then adapted to a new language. However, data from a rich language that are similar to the target language are generally more helpful. We explore methods of training bottleneck features by using data that are more similar to the target language. Our experiments on speech recognition and keyword spotting tasks with IARPA-Babel languages show that our proposed methods outperform typical multilingual DNNs. Ekapol Chuangsuwanich, Yu Zhang 0033, James R. Glass |
ICASSP | 1 |
| 2016 | Prediction-adaptation-correction recurrent neural networks for low-resource language speech recognitionabstractIn this paper, we investigate the use of prediction-adaptation-correction recurrent neural networks (PAC-RNNs) for low-resource speech recognition. A PAC-RNN is comprised of a pair of neural networks in which a correction network uses auxiliary information given by a prediction network to help estimate the state probability. The information from the correction network is also used by the prediction network in a recurrent loop. Our model outperforms other state-of-the-art neural networks (DNNs, LSTMs) on IARPA-Babel tasks. Moreover, transfer learning from a language that is similar to the target language can help improve performance further. Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass, Dong Yu 0001 |
ICASSP | 2 |
| 2014 | Extracting deep neural network bottleneck features using low-rank matrix factorizationabstractIn this paper, we investigate the use of deep neural networks (DNNs) to generate a stacked bottleneck (SBN) feature representation for low-resource speech recognition. We examine different SBN extraction architectures, and incorporate low-rank matrix factorization in the final weight layer. Experiments on several low-resource languages demonstrate the effectiveness of the SBN configurations when compared to state-of-the-art hybrid DNN approaches. Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass |
ICASSP | 2 |
| 2014 | Language ID-based training of multilingual stacked bottleneck featuresabstractIn this paper, we explore multilingual feature-level data sharing via Deep Neural Network (DNN) stacked bottleneck features. Given a set of available source languages, we apply language identification to pick the language most similar to the target language, for more efficient use of multilingual resources. Our experiments with IARPA-Babel languages show that bottleneck features trained on the most similar source language perform better than those trained on all available source languages. Further analysis suggests that only data similar to the target language is useful for multilingual training. Index Terms: Multilingual, Bottleneck features, DNN Anne Cutler, Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass |
INTERSPEECH | 3 |
| 2014 | Graph-based re-ranking using acoustic feature similarity between search results for spoken term detection on low-resource languagesabstractAcoustic feature similarity between search results has been shown to be very helpful for the task of spoken term detection (STD). A graph-based re-ranking approach for STD has been proposed based on the concept that search results, which are acoustically similar to other results with higher confidence scores, should have higher scores themselves. In this approach, the similarity between all search results of a given term are considered as a graph, and the confidence scores of the search results propagate through this graph. Since this approach can improve STD results without any additional labelled data, it is especially suitable for STD on languages with limited amounts of annotated data. However, its performance has not been widely studied on benchmark corpora. In this paper, we investigate the effectiveness of the graph-based reranking approach on limited language data from the IARPA Babel program. Experiments on the low-resource languages, Assamese, Bengali and Lao, show that graph-based re-ranking improves STD systems using fuzzy matching, and lattices based on different kinds of units including words, subwords, and hybrids. Index Terms: Random Walk, Spoken Term Detection Hung-yi Lee, Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass |
INTERSPEECH | 3 |
| 2012 | Handling uncertain observations in unsupervised topic-mixture language model adaptationabstractWe propose an extension to the recent approaches in topic-mixture modeling such as Latent Dirichlet Allocation and Topic Tracking Model for the purpose of unsupervised adaptation in speech recognition. Instead of using the 1-best input given by the speech recognizer, the proposed model takes confusion network as an input to alleviate recognition errors. We incorporate a selection variable which helps reweight the recognition output, thus creating a more accurate latent topic estimate. Compared to adapting based on just one recognition hypothesis, the proposed model show WER improvements on two different tasks. Ekapol Chuangsuwanich, Shinji Watanabe 0001, Takaaki Hori, Tomoharu Iwata, James R. Glass |
ICASSP | 1 |
| 2011 | Robust Voice Activity Detector for Real World Applications Using Harmonicity and Modulation FrequencyabstractThe task of robustly detecting distant speech in low SNR environments for automatic speech recognition is examined using a two-stage approach based on two distinguishing features of speech, namely harmonicity and modulation frequency (MF). A modified metric for harmonicity is used as a gating function to a set of parallel classifiers that incorporate MFs computed on different frequency bands. Performance is evaluated on both the frame-level discriminative power and also the system level ASR results on a real-world robotic forklift task. Compared to other previously proposed features such as relative spectral entropy, and classification strategies involving MFs, the combined approach shows good generalization across different kinds of dynamic noise conditions, and obtains a significant improvement on the false alarm rate at low speech miss rate settings. The overall ASR results also improved significantly compared to the ESTI AMR-VAD2, while reducing the number of false alarms by a factor of two. Index Terms: voice activity detection, modulation frequency, harmonicity, human-robot interaction. Ekapol Chuangsuwanich, James R. Glass |
INTERSPEECH | 1 |
| 2011 | Exploiting Intra-Conversation Variability for Speaker DiarizationabstractIn this paper, we propose a new approach to speaker diariza-tion based on the Total Variability approach to speaker verifica-tion. Drawing on previous work done in applying factor anal-ysis priors to the diarization problem, we arrive at a simplified approach that exploits intra-conversation variability in the To-tal Variability space through the use of Principal Component Analysis (PCA). Using our proposed methods, we demonstrate the ability to achieve state-of-the-art performance (0.9 % DER) in the diarization of summed-channel telephone data from the NIST 2008 SRE. Stephen H. Shum, Najim Dehak, Ekapol Chuangsuwanich, Douglas A. Reynolds, James R. Glass |
INTERSPEECH | 3 |
| 2010 | Spoken command of large mobile robots in outdoor environmentsabstractWe describe a speech system for commanding robots in human-occupied outdoor military supply depots. To operate in such environments, the robots must be as easy to interact with as are humans, i.e. they must reliably understand ordinary spoken instructions, such as orders to move supplies, as well as commands and warnings, spoken or shouted from distances of tens of meters. These design goals preclude close-talking microphones and “push-to-talk” buttons that are typically used to isolate commands from the sounds of vehicles, machinery and non-relevant speech. We used multiple microphones to provide omnidirectional coverage. A novel voice activity detector was developed to detect speech and select the appropriate microphone to listen to. Finally, we developed a recognizer model that could successfully recognize commands when heard amidst other speech within a noisy environment. When evaluated on speech data in the field, this system performed significantly better than a more computationally intensive baseline system, reducing the effective false alarm rate by a factor of 40, while maintaining the same level of precision. Ekapol Chuangsuwanich, D. Scott Cyphers, James R. Glass, Seth J. Teller |
SLT | 1 |