VLDB 2026 Research / reviewers in the wild / expert
Sunhee Kim
dblp:02/9234
· DBLP profile ↗
33ranked-venue papers
7as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Domain-Specific Multilingual Speech Translation Corpus via Simultaneous InterpretationabstractThis paper presents a novel multilingual speech translation corpus for complex, domain-specific content in Korean, English, Spanish, and Japanese. The corpus contains 4,000 hours of parallel speech, including 1,000 hours of Korean audio with simultaneous sight interpretations in the other three languages by 294 professionals (242 interpreters and 52 Korean voice actors). It also includes transcriptions, translations, and annotations for all languages. The Dewey Decimal Classification was adapted to balance knowledge representation, and speech tasks were conducted in a controlled studio environment to ensure data consistency. Translation, transcription, and annotation workflows were managed through a custom-built platform. The corpus captures nuanced contexts, cultural sensitivities, and domain-specific terminology, addressing linguistic challenges like structural differences between SOV (Korean, Japanese) and SVO languages (English, Spanish). Preliminary evaluations indicate its potential to enhance end-to-end speech translation models, support cross-lingual transfer learning, and tackle real-time translation issues. Gary Geunbae Lee, Hung Soon Kim, Sunhee Kim, Minhwa Chung |
ICASSP | 4 |
| 2025 | Exploring Acoustic Foundations in Speech Production Assessment Models for Children with Cochlear ImplantsabstractAlthough substantial research has been conducted on automatic speech assessment models leveraging speech representations derived from self-supervised learning models, the underlying mechanisms remain relatively underexplored. This study investigates the acoustic foundations of automatic speech production skill assessment models for children with cochlear implants, which helps enhance model performance and elucidate the basis of assessment outcomes. We analyze the statistical differences in acoustic characteristics as a function of speech scores for articulation and prosody. Using a general probing approach, models are trained with layer-wise embeddings from wav2vec2.0 and probed through simple regression models. The probing model performance is interpreted as an indicator of the information encoded within the speech representations. Experimental results demonstrate that the assessment models capture distinct acoustic features depending on the target of assessment, shedding light on the acoustic basis of the results and revealing the strengths and limitations of the models. Sunhee Kim, Minhwa Chung |
ICASSP | 2 |
| 2025 | Speech-Based Automatic Chronic Kidney Disease Diagnosis via Transformer Fusion of Glottal and Spectrogram FeaturesabstractChronic kidney disease (CKD) is a global health concern characterized by a gradual and irreversible decline in kidney function. Early diagnosis and timely intervention are crucial, yet current methods rely primarily on invasive blood and urine tests. Since CKD affects the respiratory system and alters speech production, vocal characteristics may serve as biomarkers for disease detection. This study proposes a deep learning-based approach that integrates spectrogram and glottal features for CKD diagnosis. Spectrograms capture broad acoustic characteristics, whereas glottal features, known to be influenced by CKD, provide complementary phonatory information. To effectively fuse these features, we employ a transformer-like architecture. The proposed method achieves an accuracy and a macro F1 score of 0.96, demonstrating its potential as an objective, non-invasive diagnostic tool. In addition, we analyze attention weights and gradient-based saliency maps to enhance model interpretability. Jihyun Mun, Minhwa Chung, Sunhee Kim |
INTERSPEECH | 3 |
| 2025 | A Cascaded Multimodal Framework for Automatic Social Communication Severity Assessment in Children with Autism Spectrum DisorderabstractAutism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by deficits in social communication, affecting both language use and speech patterns. Since assessment relies on behavioral observations rather than standardized medical tests, developing an objective evaluation method is essential. Recognizing that ASD impacts both language and speech production, this study proposes a cascaded multimodal framework for ASD severity assessment. The framework processes raw audio, generates transcriptions via automatic speech recognition, and extracts linguistic and acoustic features using speech-language foundation models. Given the atypical suprasegmental and segmental speech characteristics in ASD, two speech foundation models are employed. A co-attention mechanism then integrates these representations to estimate severity. Achieving a Spearman's correlation of 0.5629 with human ratings, the proposed approach offers a scalable, fully automated ASD assessment tool. Jihyun Mun, Sunhee Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2025 | Multilingual Speech Assessment Using Cross-Attention and Multitask LearningabstractAutomatic speech assessment plays a vital role in language learning by providing essential feedback on pronunciation, fluency, and overall speaking ability. However, developing effective multilingual speech assessment systems poses significant challenges with the complexity of modeling multiple languages and limited availability of labeled data, especially for languages other than English. In this study, we propose a multilingual speech assessment system for three languages-English, German, and French, which are produced by Korean learners. Enhanced by cross-attention and multitask learning mechanisms, our model utilizes pre-trained models to capture both language-specific and cross-linguistic features, predicting overall speaking proficiency scores directly from raw speech audio. Experimental results demonstrate that our proposed method, especially with wav2vec 2.0, presents superior performance on both seen and unseen data compared to monolingual models. Sehyun Oh, Minhwa Chung, Sunhee Kim |
INTERSPEECH | 3 |
| 2025 | Multimodal and Multitask Learning for Predicting Multiple Scores in L2 English SpeechabstractThis study presents a novel multimodal and multitask learning model for predicting five proficiency scores of L2 English speeches. The proposed approach integrates speech and text embeddings using multimodal transformer blocks with crossmodal attention to refine features dynamically between modalities, capturing complementary information. A joint loss function, combining MSE and a Trait-Aware (TA) loss, enhances the model by leveraging relationships among proficiency traits. Experiments with different combinations of four embeddings (MFCCs, GloVe, wav2vec 2.0, and BERT) revealed that the proposed model with wav2vec 2.0 and BERT embeddings achieved the best performance, with a mean PCC of 0.734 and a standard deviation of 0.0129 across five criteria. This approach significantly outperforms unimodal and baseline multimodal models, demonstrating the potential of advanced multimodal architectures and task-aware optimization in automated speech assessment systems. Sehyun Oh, Sunhee Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2024 | Constructing Korean Learners' L2 Speech Corpus of Seven Languages for Automatic Pronunciation AssessmentabstractMultilingual L2 speech corpora for developing automatic speech assessment are currently available, but they lack comprehensive annotations of L2 speech from non-native speakers of various languages. This study introduces the methodology of designing a Korean learners’ L2 speech corpus of seven languages: English, Japanese, Chinese, French, German, Spanish, and Russian. We describe the development of reading scripts, reading tasks, scoring criteria, and expert evaluation methods in detail. Our corpus contains 1,200 hours of L2 speech data from Korean learners (400 hours for English, 200 hours each for Japanese and Chinese, 100 hours each for French, German, Spanish, and Russian). The corpus is annotated with spelling and pronunciation transcription, expert pronunciation assessment scores (accuracy of pronunciation and fluency of prosody), and metadata such as gender, age, self-reported language proficiency, and pronunciation error types. We also propose a practical verification method and a reliability threshold to ensure the reliability and objectivity of large-scale subjective evaluation data. Sunhee Kim, Minhwa Chung |
LREC/COLING | 2 |
| 2024 | Speech Corpus for Korean Children with Autism Spectrum Disorder: Towards Automatic Assessment SystemsabstractDespite the growing demand for digital therapeutics for children with Autism Spectrum Disorder (ASD), there is currently no speech corpus available for Korean children with ASD. This paper introduces a speech corpus specifically designed for Korean children with ASD, aiming to advance speech technologies such as pronunciation and severity evaluation. Speech recordings from speech and language evaluation sessions were transcribed, and annotated for articulatory and linguistic characteristics. Three speech and language pathologists rated these recordings for social communication severity (SCS) and pronunciation proficiency (PP) using a 3-point Likert scale. The total number of participants will be 300 for children with ASD and 50 for typically developing (TD) children. The paper also analyzes acoustic and linguistic features extracted from speech data collected and completed for annotation from 73 children with ASD and 9 TD children to investigate the characteristics of children with ASD and identify significant features that correlate with the clinical scores. The results reveal some speech and linguistic characteristics in children with ASD that differ from those in TD children or another subgroup of ASD categorized by clinical scores, demonstrating the potential for developing automatic assessment systems for SCS and PP. Jihyun Mun, Sunhee Kim, Minhwa Chung |
LREC/COLING | 3 |
| 2024 | Automatic Speech Recognition and Assessment Systems Incorporated into Digital Therapeutics for Children with Autism Spectrum Disorder
Jihyun Mun, Sunhee Kim, HyunJu Park, Suvin Yang, HyunDon Kim, SeungJae Noh, WonBin Kim, Minhwa Chung |
ICCHP (2) | 3 |
| 2024 | Automatic Assessment of Speech Production Skills for Children with Cochlear Implants Using Wav2Vec2.0 Acoustic EmbeddingsabstractThis study introduces an automatic assessment model for speech production skills of children with cochlear implants (CIs) to support home-based speech therapy. The model employs acoustic embeddings from self-supervised models and considers speech traits of both normal hearing (NH) adults and children, which is a novel method for evaluating speech of children with disorders. It combines phoneme embeddings and two acoustic embeddings from Wav2Vec2.0 models, each trained on the speech of NH adults and children, via multi-head attention. Using a speech corpus of Korean-speaking children with CIs, our model outperforms single-embedding methods in a Pearson correlation coefficient between predicted and expert-rated scores, with a relative improvement of 51%. The results highlight the effectiveness of Wav2Vec2.0 acoustic embeddings and the importance of incorporating both of typical speech patterns of NH adults and children in assessing speech production skills in children with CIs. Sunhee Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2024 | Developing an End-to-End Framework for Predicting the Social Communication Severity Scores of Children with Autism Spectrum DisorderabstractAutism Spectrum Disorder (ASD) is a lifelong condition that significantly influencing an individual's communication abilities and their social interactions. Early diagnosis and intervention are critical due to the profound impact of ASD's characteristic behaviors on foundational developmental stages. However, limitations of standardized diagnostic tools necessitate the development of objective and precise diagnostic methodologies. This paper proposes an end-to-end framework for automatically predicting the social communication severity of children with ASD from raw speech data. This framework incorporates an automatic speech recognition model, fine-tuned with speech data from children with ASD, followed by the application of fine-tuned pre-trained language models to generate a final prediction score. Achieving a Pearson Correlation Coefficient of 0.6566 with human-rated scores, the proposed method showcases its potential as an accessible and objective tool for the assessment of ASD. Jihyun Mun, Sunhee Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2023 | A computational method of identifying false positives in the variant callingabstractWe studied the strain-specific effect in genetic variant calling. To this end, we used two major strains of the rice genome, Indica and Japonica, and called the variant with models that are different in the composition of samples from the two strains. We found that the more the samples differed in their strains from the reference, the more variants were predicted. We used confusion matrices in machine learning methods to compare the performance of different variant calling models. We found that a significant portion of predicted variants are potential false positive variants. We then proposed a method to identify the false positives. The proposed method involves calling true variants from the purebred samples and the reference of the same strain. We demonstrated the validity of the proposed method on the different variant calling models. Sunhee Kim, Sang-Ho Chu, Chang-Yong Lee |
IEEE Big Data | 1 |
| 2023 | Automatic Severity Classification of Dysarthric Speech by Using Self-Supervised Model with Multi-Task LearningabstractAutomatic assessment of dysarthric speech is essential for sustained treatments and rehabilitation. However, obtaining atypical speech is challenging, often leading to data scarcity issues. To tackle the problem, we propose a novel automatic severity assessment method for dysarthric speech, using the self-supervised model in conjunction with multi-task learning. Wav2vec 2.0 XLS-R is jointly trained for two different tasks: severity classification and auxiliary automatic speech recognition (ASR). For the baseline experiments, we employ hand-crafted acoustic features and machine learning classifiers such as SVM, MLP, and XGBoost. Explored on the Korean dysarthric speech QoLT database, our model out-performs the traditional baseline methods, with a relative percentage increase of 1.25% for F1-score. In addition, the proposed model surpasses the model trained without ASR head, achieving 10.61% relative percentage improvements. Furthermore, we present how multi-task learning affects the severity classification performance by analyzing the latent representations and regularization effect. Eunjung Yeo, Kwanghee Choi, Sunhee Kim, Minhwa Chung |
ICASSP | 3 |
| 2023 | An Analysis of Glottal Features of Chronic Kidney Disease Speech and Its Application to CKD Detection
Jihyun Mun, Sunhee Kim, Myeong-Ju Kim, Jiwon Ryu, Sejoong Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2023 | Identifying Stable Sections for Formant Frequency Extraction of French Nasal Vowels Based on Difference ThresholdsabstractFormant frequencies of a vowel are generally extracted from midpoints or central sections on the time axis. Nasal vowels present a challenge for obtaining stable formant frequencies, as the midpoint often falls in an anti-formant section where vocal energy is lost through the nasal cavity. This study proposes a stable section for extracting nasal vowel formant frequencies using difference thresholds, which identify a vowel as being distinct when F1 is above 60 Hz and/or F2 is above 200 Hz. For the experiment, 481 disyllabic words (232 nasal vowels and 294 oral counterparts) are selected from an online French-Korean Dictionary. Each vowel is divided into 10 intervals, and the stable section is identified as one or more continuous intervals with lower frequencies than the difference thresholds. The results show that the stable section for nasal vowels is identified in 20%similar to 50% of the vowels, while the stable section for oral vowels is identified in 20%similar to 80% of the vowels. Hye-Sook Park, Sunhee Kim |
INTERSPEECH | 2 |
| 2023 | A Joint Model for Pronunciation Assessment and Mispronunciation Detection and Diagnosis with Multi-task LearningabstractEmpirical studies report a strong correlation between pronunciation proficiency scores and phonetic errors in non-native speech assessments of human evaluators. However, the existing system of computer-assisted pronunciation training (CAPT) regards automatic pronunciation assessment (APA) and mis-pronunciation detection and diagnosis (MDD) as independent and focuses on individual performance improvement. Motivated by the correlation between two tasks, we propose a novel architecture that jointly tackles APA and MDD using CTC and cross-entropy criteria with a multi-task learning scheme to benefit both tasks. To leverage additional knowledge transfer, Wav2Vec2-robust finetuned on TIMIT is used for the joint optimization. The integrated model significantly outperforms single-task learning, with a mean of 0.057 PCC increase for APA and 0.004 F1 increase for MDD on Speechocean762, which reveals that proficiency scores and phonetic errors are correlated for both human and model assessments. Hyungshin Ryu, Sunhee Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2023 | Speech Intelligibility Assessment of Dysarthric Speech by using Goodness of Pronunciation with Uncertainty QuantificationabstractThis paper proposes an improved Goodness of Pronunciation (GoP) that utilizes Uncertainty Quantification (UQ) for automatic speech intelligibility assessment for dysarthric speech. Current GoP methods rely heavily on neural network-driven overconfident predictions, which is unsuitable for assessing dysarthric speech due to its significant acoustic differences from healthy speech. To alleviate the problem, UQ techniques were used on GoP by 1) normalizing the phoneme prediction (entropy, margin, maxlogit, logit-margin) and 2) modifying the scoring function (scaling, prior normalization). As a result, prior-normalized maxlogit GoP achieves the best performance, with a relative increase of 5.66%, 3.91%, and 23.65% compared to the baseline GoP for English, Korean, and Tamil, respectively. Furthermore, phoneme analysis is conducted to identify which phoneme scores significantly correlate with intelligibility scores in each language. Eunjung Yeo, Kwanghee Choi, Sunhee Kim, Minhwa Chung |
INTERSPEECH | 3 |
| 2022 | Computational method of database construction for genetic variant callingabstractIn this study, we examined the impact of the variant database in recalibration and developed a database-generation model that gathers potential candidates directly from resequencing genome data. Based on human genome data, we optimize the hyper-parameters in the model and evaluate the performance improvements both in terms of recalibration and variant calling. To test whether our pseudo-database approach is applicable to species other than human, we constructed pseudo-databases for sheep, rice, and chickpea, and compared its performance with dbSNP. Consistently, we find that our pseudo-database provides improved recalibration and error rates. More importantly, the use of pseudo-databases led to the identification of additional genetic variants. Therefore, the reanalysis with our pseudo-databases approach effectively recalibrates the base quality scores and consequently uncovers hidden genetic variations in published resequencing data. Sunhee Kim, Chang-Yong Lee |
IEEE Big Data | 1 |
| 2022 | A Study on the Phonetic Inventory Development of Children with Cochlear Implants for 5 Years after Implantation
Sunhee Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2021 | Automatic Severity Classification of Korean Dysarthric Speech Using Phoneme-Level Pronunciation Features
Eunjung Yeo, Sunhee Kim, Minhwa Chung |
Interspeech | 2 |
| 2020 | Dysarthria Detection and Severity Assessment Using Rhythm-Based Metrics
Abner Hernandez, Eunjung Yeo, Sunhee Kim, Minhwa Chung |
INTERSPEECH | 3 |
| 2019 | Exploring Art with a Voice Controlled Multimodal Guide for Blind PeopleabstractThere is an increasing concern to improve the accessibility of artworks for blind people. Much of the effort has been focused on helping the visually impaired people to access the exhibition facilities, but the works of art hosted there are still difficult to experience for them. Particularly, the appreciation of visual artworks is hindered as blind visitors are not allowed to touch them in order to conserve their aesthetics and value. In this work we explore our findings using a prototype of a voice interactive multimodal guide designed to improve the accessibility of visual works of arts, such as paintings, for the blind people. The prototype identifies tactile gestures and voice commands that trigger audio descriptions and sounds while a person explores a 2.5D tactile representation of the artwork placed on the top surface of the prototype. Our preliminary findings include the results of eight user tests and Likert-type surveys. Jorge David Iranzo Bartolomé, Luis Cavazos Quero, Sunhee Kim, Myung-Yong Um, Jun-Dong Cho |
TEI | 3 |
| 2018 | An Interactive Multimodal Guide to Improve Art Accessibility for Blind PeopleabstractThe development of 3D printing technology has improved the engagement of the visually impaired people when experiencing two-dimensional visual artworks. However, it is still difficult to explore, experience and get a clear understanding. We introduce an interactive multimodal guide in which a 3D printed 2.5D representation of a painting can be explored by touch. Touching determined features in the representation triggers localized verbal, audio, wind, and light/heat feedback events that convey spatial and semantic information. In this work we present a working prototype developed through three sessions using a participatory design approach. Luis Cavazos Quero, Jorge David Iranzo Bartolomé, Seonggu Lee, En Han, Sunhee Kim, Jun-Dong Cho |
ASSETS | 5 |
| 2012 | Developing a Voice User Interface with Improved Usability for People with Dysarthria
Yumi Hwang, Daejin Shin, Chang-Yeal Yang, Seung-Yeun Lee, Byunggoo Kong, Jio Chung, Sunhee Kim, Minhwa Chung |
ICCHP (2) | 8 |
| 2012 | Comparing transcription agreement on non-native English speech corpus between native and non-native annotators
Hyuksu Ryu, Sunhee Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2012 | Korean Children's Spoken English Corpus and an Analysis of its Pronunciation Variability
Hyejin Hong, Sunhee Kim, Minhwa Chung |
LREC | 2 |
| 2012 | Multiple-Region Segmentation Without Supervision by Adaptive Global Maximum ClusteringabstractIn this paper, we propose a new method of segmenting an image into several sets of pixels with similar intensity values called regions. A multiple-region segmentation problem is unstable because the result considerably depends on the number of regions given a priori. Therefore, one of the most important tasks in solving the problem is automatically finding the number of regions. The method we propose is able to find the reasonable number of distinct regions not only for clean images but also for noisy ones. Our method is made up of two procedures. First, we develop the adaptive global maximum clustering. In this procedure, we deal with an image histogram and automatically obtain the number of significant local maxima of the histogram. This number indicates the number of different regions in the image. Second, we derive a simple and fast calculation to segment an image composed of distinct multiple regions. Then, we split an image into multiple regions according to the previous procedure. Finally, we show the efficiency of our method by comparing it with other previous methods. Sunhee Kim, Myungjoo Kang |
IEEE Trans. Image Process. | 1 |
| 2011 | A Corpus-Based Study of English Pronunciation Variations
Sunhee Kim, Kyuwhan Lee, Minhwa Chung |
INTERSPEECH | 1 |
| 2008 | Effects of allophones on the performance of Korean speech recognition
Hyejin Hong, Sunhee Kim, Minhwa Chung |
INTERSPEECH | 2 |
| 2007 | Evaluating two versions of the momel pitch modelling algorithm on a corpus of read speech in KoreanabstractThe Momel algorithm provides an automatic factoring of raw fundamental frequency into two components: a microprosodic component, corresponding to local variations of pitch caused by the phonetic nature of the speech segments and a macroprosodic component corresponding to the overall pitch pattern of the utterance which is then represented as a sequence of pitch targets. An earlier evaluation estimated the overall efficiency of the algorithm (F-measure) at around 95% on a corpus of read speech for 5 European languages and at around 93% for a corpus of spontaneous speech. In this paper we present the results of the evaluation of the output of two versions of the Momel algorithm as compared with manually corrected pitch targets for a corpus of just over 2 hours of read speech in Korean (40 continuous 5-sentence passages, each read by 5 male and 5 female speakers). The results show that the new version of the Momel algorithm performs systematically better than the earlier version. Daniel Hirst, Hyongsil Cho, Sunhee Kim, Hyunji Yu |
INTERSPEECH | 3 |
| 2005 | Durational characteristics of Korean Lombard speech
Sunhee Kim |
INTERSPEECH | 1 |
| 2004 | Phonology of exceptions for for Korean grapheme-to-phoneme conversionabstractBeing an essential part of a Korean speech recognition system and a Text-To-Speech (TTS) system, a Korean Grapheme-toPhoneme conversion system is generally composed of a set of regular rules and an exceptions dictionary [1, 2, 3]. The exceptions have been recorded in the dictionary in a simple and random manner, whereas the researches on the regular rules have been actively progressed. This paper presents a systematic description of the exceptions for a Grapheme-toPhoneme conversion system based on the analysis of entries of a lexical dictionary [4] from the phonological point of view, showing that the exceptions are related with certain limited phonological phenomena. Sunhee Kim |
INTERSPEECH | 1 |
| 2004 | A Korean grapheme-to-phoneme conversion system using selection procedure for exceptions
Sunhee Kim, Ju-Eun Ahn, Soon-Hyob Kim, Yang-Hee Lee |
INTERSPEECH | 1 |