Minhwa Chung

dblp:53/5393 · DBLP profile ↗
← Back
38ranked-venue papers
1as first author
17since 2021 · last 2025
0009-0006-1844-9633ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 A Domain-Specific Multilingual Speech Translation Corpus via Simultaneous Interpretation
abstract
This paper presents a novel multilingual speech translation corpus for complex, domain-specific content in Korean, English, Spanish, and Japanese. The corpus contains 4,000 hours of parallel speech, including 1,000 hours of Korean audio with simultaneous sight interpretations in the other three languages by 294 professionals (242 interpreters and 52 Korean voice actors). It also includes transcriptions, translations, and annotations for all languages. The Dewey Decimal Classification was adapted to balance knowledge representation, and speech tasks were conducted in a controlled studio environment to ensure data consistency. Translation, transcription, and annotation workflows were managed through a custom-built platform. The corpus captures nuanced contexts, cultural sensitivities, and domain-specific terminology, addressing linguistic challenges like structural differences between SOV (Korean, Japanese) and SVO languages (English, Spanish). Preliminary evaluations indicate its potential to enhance end-to-end speech translation models, support cross-lingual transfer learning, and tackle real-time translation issues.
Gary Geunbae Lee, Hung Soon Kim, Sunhee Kim, Minhwa Chung
ICASSP5
2025 Exploring Acoustic Foundations in Speech Production Assessment Models for Children with Cochlear Implants
abstract
Although substantial research has been conducted on automatic speech assessment models leveraging speech representations derived from self-supervised learning models, the underlying mechanisms remain relatively underexplored. This study investigates the acoustic foundations of automatic speech production skill assessment models for children with cochlear implants, which helps enhance model performance and elucidate the basis of assessment outcomes. We analyze the statistical differences in acoustic characteristics as a function of speech scores for articulation and prosody. Using a general probing approach, models are trained with layer-wise embeddings from wav2vec2.0 and probed through simple regression models. The probing model performance is interpreted as an indicator of the information encoded within the speech representations. Experimental results demonstrate that the assessment models capture distinct acoustic features depending on the target of assessment, shedding light on the acoustic basis of the results and revealing the strengths and limitations of the models.
Sunhee Kim, Minhwa Chung
ICASSP3
2025 Speech-Based Automatic Chronic Kidney Disease Diagnosis via Transformer Fusion of Glottal and Spectrogram Features
abstract
Chronic kidney disease (CKD) is a global health concern characterized by a gradual and irreversible decline in kidney function. Early diagnosis and timely intervention are crucial, yet current methods rely primarily on invasive blood and urine tests. Since CKD affects the respiratory system and alters speech production, vocal characteristics may serve as biomarkers for disease detection. This study proposes a deep learning-based approach that integrates spectrogram and glottal features for CKD diagnosis. Spectrograms capture broad acoustic characteristics, whereas glottal features, known to be influenced by CKD, provide complementary phonatory information. To effectively fuse these features, we employ a transformer-like architecture. The proposed method achieves an accuracy and a macro F1 score of 0.96, demonstrating its potential as an objective, non-invasive diagnostic tool. In addition, we analyze attention weights and gradient-based saliency maps to enhance model interpretability.
Jihyun Mun, Minhwa Chung, Sunhee Kim
INTERSPEECH2
2025 A Cascaded Multimodal Framework for Automatic Social Communication Severity Assessment in Children with Autism Spectrum Disorder
abstract
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by deficits in social communication, affecting both language use and speech patterns. Since assessment relies on behavioral observations rather than standardized medical tests, developing an objective evaluation method is essential. Recognizing that ASD impacts both language and speech production, this study proposes a cascaded multimodal framework for ASD severity assessment. The framework processes raw audio, generates transcriptions via automatic speech recognition, and extracts linguistic and acoustic features using speech-language foundation models. Given the atypical suprasegmental and segmental speech characteristics in ASD, two speech foundation models are employed. A co-attention mechanism then integrates these representations to estimate severity. Achieving a Spearman's correlation of 0.5629 with human ratings, the proposed approach offers a scalable, fully automated ASD assessment tool.
Jihyun Mun, Sunhee Kim, Minhwa Chung
INTERSPEECH3
2025 Multilingual Speech Assessment Using Cross-Attention and Multitask Learning
abstract
Automatic speech assessment plays a vital role in language learning by providing essential feedback on pronunciation, fluency, and overall speaking ability. However, developing effective multilingual speech assessment systems poses significant challenges with the complexity of modeling multiple languages and limited availability of labeled data, especially for languages other than English. In this study, we propose a multilingual speech assessment system for three languages-English, German, and French, which are produced by Korean learners. Enhanced by cross-attention and multitask learning mechanisms, our model utilizes pre-trained models to capture both language-specific and cross-linguistic features, predicting overall speaking proficiency scores directly from raw speech audio. Experimental results demonstrate that our proposed method, especially with wav2vec 2.0, presents superior performance on both seen and unseen data compared to monolingual models.
Sehyun Oh, Minhwa Chung, Sunhee Kim
INTERSPEECH2
2025 Multimodal and Multitask Learning for Predicting Multiple Scores in L2 English Speech
abstract
This study presents a novel multimodal and multitask learning model for predicting five proficiency scores of L2 English speeches. The proposed approach integrates speech and text embeddings using multimodal transformer blocks with crossmodal attention to refine features dynamically between modalities, capturing complementary information. A joint loss function, combining MSE and a Trait-Aware (TA) loss, enhances the model by leveraging relationships among proficiency traits. Experiments with different combinations of four embeddings (MFCCs, GloVe, wav2vec 2.0, and BERT) revealed that the proposed model with wav2vec 2.0 and BERT embeddings achieved the best performance, with a mean PCC of 0.734 and a standard deviation of 0.0129 across five criteria. This approach significantly outperforms unimodal and baseline multimodal models, demonstrating the potential of advanced multimodal architectures and task-aware optimization in automated speech assessment systems.
Sehyun Oh, Sunhee Kim, Minhwa Chung
INTERSPEECH3
2024 Constructing Korean Learners' L2 Speech Corpus of Seven Languages for Automatic Pronunciation Assessment
abstract
Multilingual L2 speech corpora for developing automatic speech assessment are currently available, but they lack comprehensive annotations of L2 speech from non-native speakers of various languages. This study introduces the methodology of designing a Korean learners’ L2 speech corpus of seven languages: English, Japanese, Chinese, French, German, Spanish, and Russian. We describe the development of reading scripts, reading tasks, scoring criteria, and expert evaluation methods in detail. Our corpus contains 1,200 hours of L2 speech data from Korean learners (400 hours for English, 200 hours each for Japanese and Chinese, 100 hours each for French, German, Spanish, and Russian). The corpus is annotated with spelling and pronunciation transcription, expert pronunciation assessment scores (accuracy of pronunciation and fluency of prosody), and metadata such as gender, age, self-reported language proficiency, and pronunciation error types. We also propose a practical verification method and a reliability threshold to ensure the reliability and objectivity of large-scale subjective evaluation data.
Sunhee Kim, Minhwa Chung
LREC/COLING3
2024 Speech Corpus for Korean Children with Autism Spectrum Disorder: Towards Automatic Assessment Systems
abstract
Despite the growing demand for digital therapeutics for children with Autism Spectrum Disorder (ASD), there is currently no speech corpus available for Korean children with ASD. This paper introduces a speech corpus specifically designed for Korean children with ASD, aiming to advance speech technologies such as pronunciation and severity evaluation. Speech recordings from speech and language evaluation sessions were transcribed, and annotated for articulatory and linguistic characteristics. Three speech and language pathologists rated these recordings for social communication severity (SCS) and pronunciation proficiency (PP) using a 3-point Likert scale. The total number of participants will be 300 for children with ASD and 50 for typically developing (TD) children. The paper also analyzes acoustic and linguistic features extracted from speech data collected and completed for annotation from 73 children with ASD and 9 TD children to investigate the characteristics of children with ASD and identify significant features that correlate with the clinical scores. The results reveal some speech and linguistic characteristics in children with ASD that differ from those in TD children or another subgroup of ASD categorized by clinical scores, demonstrating the potential for developing automatic assessment systems for SCS and PP.
Jihyun Mun, Sunhee Kim, Minhwa Chung
LREC/COLING4
2024 Automatic Speech Recognition and Assessment Systems Incorporated into Digital Therapeutics for Children with Autism Spectrum Disorder
Jihyun Mun, Sunhee Kim, HyunJu Park, Suvin Yang, HyunDon Kim, SeungJae Noh, WonBin Kim, Minhwa Chung
ICCHP (2)9
2024 Automatic Assessment of Speech Production Skills for Children with Cochlear Implants Using Wav2Vec2.0 Acoustic Embeddings
abstract
This study introduces an automatic assessment model for speech production skills of children with cochlear implants (CIs) to support home-based speech therapy. The model employs acoustic embeddings from self-supervised models and considers speech traits of both normal hearing (NH) adults and children, which is a novel method for evaluating speech of children with disorders. It combines phoneme embeddings and two acoustic embeddings from Wav2Vec2.0 models, each trained on the speech of NH adults and children, via multi-head attention. Using a speech corpus of Korean-speaking children with CIs, our model outperforms single-embedding methods in a Pearson correlation coefficient between predicted and expert-rated scores, with a relative improvement of 51%. The results highlight the effectiveness of Wav2Vec2.0 acoustic embeddings and the importance of incorporating both of typical speech patterns of NH adults and children in assessing speech production skills in children with CIs.
Sunhee Kim, Minhwa Chung
INTERSPEECH3
2024 Developing an End-to-End Framework for Predicting the Social Communication Severity Scores of Children with Autism Spectrum Disorder
abstract
Autism Spectrum Disorder (ASD) is a lifelong condition that significantly influencing an individual's communication abilities and their social interactions. Early diagnosis and intervention are critical due to the profound impact of ASD's characteristic behaviors on foundational developmental stages. However, limitations of standardized diagnostic tools necessitate the development of objective and precise diagnostic methodologies. This paper proposes an end-to-end framework for automatically predicting the social communication severity of children with ASD from raw speech data. This framework incorporates an automatic speech recognition model, fine-tuned with speech data from children with ASD, followed by the application of fine-tuned pre-trained language models to generate a final prediction score. Achieving a Pearson Correlation Coefficient of 0.6566 with human-rated scores, the proposed method showcases its potential as an accessible and objective tool for the assessment of ASD.
Jihyun Mun, Sunhee Kim, Minhwa Chung
INTERSPEECH3
2023 Automatic Severity Classification of Dysarthric Speech by Using Self-Supervised Model with Multi-Task Learning
abstract
Automatic assessment of dysarthric speech is essential for sustained treatments and rehabilitation. However, obtaining atypical speech is challenging, often leading to data scarcity issues. To tackle the problem, we propose a novel automatic severity assessment method for dysarthric speech, using the self-supervised model in conjunction with multi-task learning. Wav2vec 2.0 XLS-R is jointly trained for two different tasks: severity classification and auxiliary automatic speech recognition (ASR). For the baseline experiments, we employ hand-crafted acoustic features and machine learning classifiers such as SVM, MLP, and XGBoost. Explored on the Korean dysarthric speech QoLT database, our model out-performs the traditional baseline methods, with a relative percentage increase of 1.25% for F1-score. In addition, the proposed model surpasses the model trained without ASR head, achieving 10.61% relative percentage improvements. Furthermore, we present how multi-task learning affects the severity classification performance by analyzing the latent representations and regularization effect.
Eunjung Yeo, Kwanghee Choi, Sunhee Kim, Minhwa Chung
ICASSP4
2023 An Analysis of Glottal Features of Chronic Kidney Disease Speech and Its Application to CKD Detection
Jihyun Mun, Sunhee Kim, Myeong-Ju Kim, Jiwon Ryu, Sejoong Kim, Minhwa Chung
INTERSPEECH6
2023 A Joint Model for Pronunciation Assessment and Mispronunciation Detection and Diagnosis with Multi-task Learning
abstract
Empirical studies report a strong correlation between pronunciation proficiency scores and phonetic errors in non-native speech assessments of human evaluators. However, the existing system of computer-assisted pronunciation training (CAPT) regards automatic pronunciation assessment (APA) and mis-pronunciation detection and diagnosis (MDD) as independent and focuses on individual performance improvement. Motivated by the correlation between two tasks, we propose a novel architecture that jointly tackles APA and MDD using CTC and cross-entropy criteria with a multi-task learning scheme to benefit both tasks. To leverage additional knowledge transfer, Wav2Vec2-robust finetuned on TIMIT is used for the joint optimization. The integrated model significantly outperforms single-task learning, with a mean of 0.057 PCC increase for APA and 0.004 F1 increase for MDD on Speechocean762, which reveals that proficiency scores and phonetic errors are correlated for both human and model assessments.
Hyungshin Ryu, Sunhee Kim, Minhwa Chung
INTERSPEECH3
2023 Speech Intelligibility Assessment of Dysarthric Speech by using Goodness of Pronunciation with Uncertainty Quantification
abstract
This paper proposes an improved Goodness of Pronunciation (GoP) that utilizes Uncertainty Quantification (UQ) for automatic speech intelligibility assessment for dysarthric speech. Current GoP methods rely heavily on neural network-driven overconfident predictions, which is unsuitable for assessing dysarthric speech due to its significant acoustic differences from healthy speech. To alleviate the problem, UQ techniques were used on GoP by 1) normalizing the phoneme prediction (entropy, margin, maxlogit, logit-margin) and 2) modifying the scoring function (scaling, prior normalization). As a result, prior-normalized maxlogit GoP achieves the best performance, with a relative increase of 5.66%, 3.91%, and 23.65% compared to the baseline GoP for English, Korean, and Tamil, respectively. Furthermore, phoneme analysis is conducted to identify which phoneme scores significantly correlate with intelligibility scores in each language.
Eunjung Yeo, Kwanghee Choi, Sunhee Kim, Minhwa Chung
INTERSPEECH4
2022 A Study on the Phonetic Inventory Development of Children with Cochlear Implants for 5 Years after Implantation
Sunhee Kim, Minhwa Chung
INTERSPEECH3
2021 Automatic Severity Classification of Korean Dysarthric Speech Using Phoneme-Level Pronunciation Features
Eunjung Yeo, Sunhee Kim, Minhwa Chung
Interspeech3
2020 Dysarthria Detection and Severity Assessment Using Rhythm-Based Metrics
Abner Hernandez, Eunjung Yeo, Sunhee Kim, Minhwa Chung
INTERSPEECH4
2019 Self-Imitating Feedback Generation Using GAN for Computer-Assisted Pronunciation Training
abstract
Self-imitating feedback is an effective and learner-friendly method for non-native learners in Computer-Assisted Pronunciation Training. Acoustic characteristics in native utterances are extracted and transplanted onto learner's own speech input, and given back to the learner as a corrective feedback. Previous works focused on speech conversion using prosodic transplantation techniques based on PSOLA algorithm. Motivated by the visual differences found in spectrograms of native and non-native speeches, we investigated applying GAN to generate self-imitating feedback by utilizing generator's ability through adversarial training. Because this mapping is highly under-constrained, we also adopt cycle consistency loss to encourage the output to preserve the global structure, which is shared by native and non-native utterances. Trained on 97,200 spectrogram images of short utterances produced by native and non-native speakers of Korean, the generator is able to successfully transform the non-native spectrogram input to a spectrogram with properties of self-imitating feedback. Furthermore, the transformed spectrogram shows segmental corrections that cannot be obtained by prosodic transplantation. Perceptual test comparing the self-imitating and correcting abilities of our method with the baseline PSOLA method shows that the generative approach with cycle consistency loss is promising.
Seung-Hee Yang, Minhwa Chung
INTERSPEECH2
2016 Optimizing Vocabulary Modeling for Dysarthric Speech Recognition
Minsoo Na, Minhwa Chung
ICCHP (2)2
2014 Pronunciation Variants Prediction Method to Detect Mispronunciations by Korean Learners of English
abstract
This article presents an approach to nonnative pronunciation variants modeling and prediction. The pronunciation variants prediction method was developed by generalized transformation-based error-driven learning (GTBL). The modified goodness of pronunciation (GOP) score was applied to effective mispronunciation detection using logistic regression machine learning under the pronunciation variants prediction. English-read speech data uttered by Korean-speaking learners of English were collected, then pronunciation variation knowledge was extracted from the differences between the canonical phonemes and the actual phonemes of the speech data. With this knowledge, an error-driven learning approach was designed that automatically learns phoneme variation rules from phoneme-level transcriptions. The learned rules generate an extended recognition network to detect mispronunciations. Three different mispronunciation detection methods were tested including our logistic regression machine learning method with modified GOP scores and mispronunciation preference features; all three methods yielded significant improvement in predictions of pronunciation variants, and our logistic regression method showed the best performance.
Jeesoo Bang, Gary Geunbae Lee, Minhwa Chung
ACM Trans. Asian Lang. Inf. Process.4
2012 Developing a Voice User Interface with Improved Usability for People with Dysarthria
Yumi Hwang, Daejin Shin, Chang-Yeal Yang, Seung-Yeun Lee, Byunggoo Kong, Jio Chung, Sunhee Kim, Minhwa Chung
ICCHP (2)9
2012 Comparing transcription agreement on non-native English speech corpus between native and non-native annotators
Hyuksu Ryu, Sunhee Kim, Minhwa Chung
INTERSPEECH3
2012 Dysarthric Speech Database for Development of QoLT Software Technology
Dae-Lim Choi, Bong-Wan Kim, Yeon-Whoa Kim, Yongnam Um, Minhwa Chung
LREC6
2012 Korean Children's Spoken English Corpus and an Analysis of its Pronunciation Variability
Hyejin Hong, Sunhee Kim, Minhwa Chung
LREC3
2011 A Corpus-Based Study of English Pronunciation Variations
Sunhee Kim, Kyuwhan Lee, Minhwa Chung
INTERSPEECH3
2010 Effects of Korean learners' consonant cluster reduction strategies on English speech recognition performance
Hyejin Hong, Minhwa Chung
INTERSPEECH3
2009 Improving phone recognition performance via phonetically-motivated units
Hyejin Hong, Minhwa Chung
INTERSPEECH2
2008 Effects of allophones on the performance of Korean speech recognition
Hyejin Hong, Sunhee Kim, Minhwa Chung
INTERSPEECH3
2005 Improved semi-dynamic network decoding using WFSTs
Dong-Hoon Ahn, Su-Byeong Oh, Minhwa Chung
INTERSPEECH3
2005 Automatic generation of domain-dependent pronunciation lexicon with data-driven rules and rule adaptation
Je Hun Jeon, Minhwa Chung
INTERSPEECH2
2004 Pronunciation lexicon modeling and design for Korean large vocabulary continuous speech recognition
abstract
In this paper, we describe a pronunciation lexicon model which is especially useful for constructing morpheme-based pronunciation lexicon to improve the performance of a Korean LVCSR. There are a lot of pronunciation variations occurring at morpheme boundaries in continuous speech. For modeling of cross-morpheme pronunciation variations, we usually used a context-dependent multiple pronunciation lexicon with possible multiple phonetic transcriptions for each word. Since phonemic context together with morphological category and morpheme boundary information affect Korean pronunciation variations, we have distinguished phonological rules that can be applied to phonemes in withinmorpheme and cross-morpheme. However, pronunciation variations in morpheme boundaries are increasing the lexicon size; we have designed the optimized pronunciation lexicon which is decreasing the confusability and increasing pronunciation coverage. The results of Korean Broadcast News Transcription experiments show that a reduction of 18% in pronunciation lexicon size and an absolute reduction of 0.27% in WER from the same lexical entries were achieved by building a proposed pronunciation lexicon.
Kyong-Nim Lee, Minhwa Chung
INTERSPEECH2
2003 Modeling cross-morpheme pronunciation variations for korean large vocabulary continuous speech recognition
Kyong-Nim Lee, Minhwa Chung
INTERSPEECH2
2003 Morpheme-based lexical modeling for korean broadcast news transcription
Young-Hee Park, Dong-Hoon Ahn, Minhwa Chung
INTERSPEECH3
2002 Compact subnetwork-based large vocabulary continuous speech recognition
Dong-Hoon Ahn, Minhwa Chung
INTERSPEECH2
2001 A one pass semi-dynamic network decoder based on language model network
Dong-Hoon Ahn, Minhwa Chung
INTERSPEECH2
1998 Automatic generation of Korean pronunciation variants by multistage applications of phonological rules
abstract
Phonetic transcriptions are often manually encoded in a pronunciation lexicon. This process is time consuming and requires linguistic expertise. Moreover, it is very difficult to maintain consistency. To handle these problems, we present a model that produces Korean pronunciation variants based on morphophonological analysis. By analyzing phonological variations frequently found in spoken Korean, we have derived about 800 phonemic contexts that would trigger the applications of the corresponding phonemic and allophonic rules. In generating pronunciation variants, morphological analysis is preceded to handle variations of phonological words. According to the morphological category, a set of finite state automata tables reflecting phonemic context is looked up to generate pronunciation variants. Our experiments show that the proposed model produces mostly correct pronunciation variants of phonological words consisting of several morphemes.
Je Hun Jeon, Sunhwa Cha, Minhwa Chung, Jun Park, Kyuwoong Hwang
ICSLP3
1995 Parallel Natural Language Processing on a Semantic Network Array Processor
abstract
This paper presents a parallel natural language processing system implemented on a marker-passing parallel AI computer, the Semantic Network Array Processor (SNAP). Our system uses a memory-based parsing approach in which parsing is viewed as a memory search process. Linguistic information is stored as phrasal patterns in a semantic network knowledge base distributed over the memory of the parallel computer. Parsing is performed by recognizing and linking phrasal patterns that reflect a sentence interpretation. This is achieved by propagating markers over the distributed network. We have developed a system capable of processing newswire articles from a particular domain. The paper presents the structure of the system, the memory-based parsing method used, and the performance results obtained.>
Minhwa Chung, Dan I. Moldovan
IEEE Trans. Knowl. Data Eng.1