VLDB 2026 Research / reviewers in the wild / expert
Jangwon Kim
dblp:49/8761
· DBLP profile ↗
30ranked-venue papers
15as first author
12since 2021 · last 2026
0000-0003-0228-3502ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 13 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bayesian policy distillation: Towards lightweight and fast neural policy networks
Jangwon Kim, Yoonsu Jang, Yoonhee Gil, Soohee Han |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Radon averaging: A practical approach for designing rotation-invariant models
Jangwon Kim, Sanghyun Ryoo, Junkee Hong, Soohee Han |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Shaping Q -values right: A distributional normalized actor-critic approach
Jangwon Kim, Soohee Han |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Efficient knowledge distillation with emphasized Radon features for wafer bin map classification
Sanghyun Ryoo, Jangwon Kim, Jaehyung Cho, Jongyul Lee, Junkee Hong, Soohee Han |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Provable generalization of clipped double Q -learning for variance reduction and sample efficiency
Jangwon Kim, Jiseok Jeong, Soohee Han |
Neurocomputing | 1 |
| 2026 | Reinforcement learning via conservative agent for environments with random delays
Jangwon Kim, Jiseok Jeong, Soohee Han |
Neural Networks | 2 |
| 2025 | A local patch regression-based generative model for urban flood prediction in data-poor areas
Jangwon Kim, Kyungjun Kim, Soohee Han |
Expert Syst. Appl. | 3 |
| 2025 | Overcoming intermittent instability in reinforcement learning via gradient norm preservation
Jangwon Kim, Soohee Han |
Inf. Sci. | 3 |
| 2025 | A dataset and benchmark for hospital course summarization with adapted large language modelsabstractOBJECTIVE: Brief hospital course (BHC) summaries are clinical documents that summarize a patient's hospital stay. While large language models (LLMs) depict remarkable capabilities in automating real-world tasks, their capabilities for healthcare applications such as synthesizing BHCs from clinical notes have not been shown. We introduce a novel preprocessed dataset, the MIMIC-IV-BHC, encapsulating clinical note and BHC pairs to adapt LLMs for BHC synthesis. Furthermore, we introduce a benchmark of the summarization performance of 2 general-purpose LLMs and 3 healthcare-adapted LLMs. MATERIALS AND METHODS: Using clinical notes as input, we apply prompting-based (using in-context learning) and fine-tuning-based adaptation strategies to 3 open-source LLMs (Clinical-T5-Large, Llama2-13B, and FLAN-UL2) and 2 proprietary LLMs (Generative Pre-trained Transformer [GPT]-3.5 and GPT-4). We evaluate these LLMs across multiple context-length inputs using natural language similarity metrics. We further conduct a clinical study with 5 clinicians, comparing clinician-written and LLM-generated BHCs across 30 samples, focusing on their potential to enhance clinical decision-making through improved summary quality. We compare reader preferences for the original and LLM-generated summary using Wilcoxon signed-rank tests. We further request optional qualitative feedback from clinicians to gain deeper insights into their preferences, and we present the frequency of common themes arising from these comments. RESULTS: The Llama2-13B fine-tuned LLM outperforms other domain-adapted models given quantitative evaluation metrics of Bilingual Evaluation Understudy (BLEU) and Bidirectional Encoder Representations from Transformers (BERT)-Score. GPT-4 with in-context learning shows more robustness to increasing context lengths of clinical note inputs than fine-tuned Llama2-13B. Despite comparable quantitative metrics, the reader study depicts a significant preference for summaries generated by GPT-4 with in-context learning compared to both Llama2-13B fine-tuned summaries and the original summaries (P<.001), highlighting the need for qualitative clinical evaluation. DISCUSSION AND CONCLUSION: We release a foundational clinically relevant dataset, the MIMIC-IV-BHC, and present an open-source benchmark of LLM performance in BHC synthesis from clinical notes. We observe high-quality summarization performance for both in-context proprietary and fine-tuned open-source LLMs using both quantitative metrics and a qualitative clinical reader study. Our research effectively integrates elements from the data assimilation pipeline: our methods use (1) clinical data sources to integrate, (2) data translation, and (3) knowledge creation, while our evaluation strategy paves the way for (4) deployment. Asad Aali, Dave Van Veen, Yamin Ishraq Arefeen, Jason Hom, Christian Bluethgen, Eduardo Pontes Reis, Sergios Gatidis, Namuun Clifford, Joseph Daws, Arash S. Tehrani, Jangwon Kim, Akshay Chaudhari |
J. Am. Medical Informatics Assoc. | 11 |
| 2023 | Belief Projection-Based Reinforcement Learning for Environments with Delayed FeedbackabstractWe present a novel actor-critic algorithm for an environment with delayed feedback, which addresses the state-space explosion problem of conventional approaches. Conventional approaches use an augmented state constructed from the last observed state and actions executed since visiting the last observed state. Using the augmented state space, the correct Markov decision process for delayed environments can be constructed; however, this causes the state space to explode as the number of delayed timesteps increases, leading to slow convergence. Our proposed algorithm, called Belief-Projection-Based Q-learning (BPQL), addresses the state-space explosion problem by evaluating the values of the critic for which the input state size is equal to the original state-space size rather than that of the augmented one. We compare BPQL to traditional approaches in continuous control tasks and demonstrate that it significantly outperforms other algorithms in terms of asymptotic performance and sample efficiency. We also show that BPQL solves long-delayed environments, which conventional approaches are unable to do. Jangwon Kim, Jiwook Kang, Jongchan Baek, Soohee Han |
NeurIPS | 1 |
| 2022 | Leveraging Task Transferability to Meta-learning for Clinical Section Classification with Limited DataabstractIdentifying sections is one of the critical components of understanding medical information from unstructured clinical notes and developing assistive technologies for clinical notewriting tasks.Most state-of-the-art text classification systems require thousands of indomain text data to achieve high performance.However, collecting in-domain and recent clinical note data with section labels is challenging given the high level of privacy and sensitivity.This paper proposes an algorithmic way to improve the task transferability of meta-learningbased text classification in order to address the issue of low-resource target data.Specifically, we explore how to make the best use of the source dataset and propose a unique task transferability measure named Normalized Negative Conditional Entropy (NNCE).Leveraging the NNCE, we develop strategies for selecting clinical categories and sections from source task data to boost cross-domain meta-learning accuracy.Experimental results show that our task selection strategies improve section classification accuracy significantly compared to meta-learning algorithms. Zhuohao Chen, Jangwon Kim, Ram Bhakta, Mustafa Y. Sir |
ACL (1) | 2 |
| 2022 | Cmri2spec: Cine MRI Sequence to Spectrogram Synthesis via A Pairwise Heterogeneous TranslatorabstractMultimodal representation learning using visual movements from cine magnetic resonance imaging (MRI) and their acoustics has shown great potential to learn shared representation and to predict one modality from another. Here, we propose a new synthesis framework to translate from cine MRI sequences to spectrograms with a limited dataset size. Our framework hinges on a novel fully convolutional heterogeneous translator, with a 3D CNN encoder for efficient sequence encoding and a 2D transpose convolution decoder. In addition, a pairwise correlation of the samples with the same speech word is utilized with a latent space representation disentanglement scheme. Furthermore, an adversarial training approach with generative adversarial networks is incorporated to provide enhanced realism on our generated spectrograms. Our experimental results, carried out with a total of 63 cine MRI sequences alongside speech acoustics, show that our framework improves synthesis accuracy, compared with competing methods. Our framework thereby has shown the potential to aid in better understanding the relationship between the two modalities. Xiaofeng Liu 0001, Fangxu Xing, Maureen Stone 0001, Jerry L. Prince, Jangwon Kim, Georges El Fakhri, Jonghye Woo |
ICASSP | 5 |
| 2020 | Vocal tract shaping of emotional speech
Jangwon Kim, Asterios Toutios, Sungbok Lee, Shri Narayanan |
Comput. Speech Lang. | 1 |
| 2017 | Database of Volumetric and Real-Time Vocal Tract MRI for Speech Science
Tanner Sorensen, Z.-I. Skordilis, Asterios Toutios, Yoon-Chul Kim, Yinghua Zhu, Jangwon Kim, Adam C. Lammert, Vikram Ramanarayanan, Louis Goldstein, Dani Byrd, Krishna S. Nayak, Shri Narayanan |
INTERSPEECH | 6 |
| 2016 | Pathological speech processing: State-of-the-art, current challenges, and future directionsabstractThe study of speech pathology involves evaluation and treatment of speech production related disorders affecting phonation, fluency, intonation and aeromechanical components of respiration. Recently, speech pathology has garnered special interest amongst machine learning and signal processing (ML-SP) scientists. This growth in interest is led by advances in novel data collection technology, data science, speech processing and computational modeling. These in turn have enabled scientists in better understanding both the causes and effects of pathological speech conditions. In this paper, we review the application of machine learning and signal processing techniques to speech pathology and specifically focus on three different aspects. First, we list challenges such as controlling subjectivity in pathological speech assessments and patient variability in the application of ML-SP tools to the domain. Second, we discuss feature design methods and machine learning algorithms using a combination of domain knowledge and data driven methods. Finally, we present some case studies related to analysis of pathological speech and discuss their design. Rahul Gupta 0001, Theodora Chaspari, Jangwon Kim, Naveen Kumar 0004, Daniel Bone, Shri Narayanan |
ICASSP | 3 |
| 2016 | Illustrating the Production of the International Phonetic Alphabet Sounds Using Fast Real-Time Magnetic Resonance Imaging
Asterios Toutios, Sajan Goud Lingala, Colin Vaz, Jangwon Kim, John H. Esling, Patricia A. Keating, Matthew Gordon, Dani Byrd, Louis Goldstein, Krishna S. Nayak, Shri Narayanan |
INTERSPEECH | 4 |
| 2016 | Speaker verification based on the fusion of speech acoustics and inverted articulatory signals
Ming Li 0026, Jangwon Kim, Adam C. Lammert, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2015 | Automated evaluation of non-native English pronunciation quality: combining knowledge- and data-driven features at multiple time scalesabstractAutomatically evaluating pronunciation quality of non-native speech has seen tremendous success in both research and com-mercial settings, with applications in L2 learning. In this paper, submitted for the INTERSPEECH 2015 Degree of Nativeness Sub-Challenge, this problem is posed under a challenging cross-corpora setting using speech data drawn from multiple speakers from a variety of language backgrounds (L1) reading different English sentences. Since the perception of non-nativeness is re-alized at the segmental and suprasegmental linguistic levels, we explore a number of acoustic cues at multiple time scales. We experiment with both data-driven and knowledge-inspired fea-tures that capture degree of nativeness from pauses in speech, speaking rate, rhythm/stress, and goodness of phone pronunci-ation. One promising finding is that highly accurate automated assessment can be attained using a small diverse set of intuitive and interpretable features. Performance is further boosted by smoothing scores across utterances from the same speaker; our best system significantly outperforms the challenge baseline. Matthew Black, Daniel Bone, Z.-I. Skordilis, Rahul Gupta 0001, Pavlos Papadopoulos, Sandeep Nallan Chakravarthula, Bo Xiao 0003, Maarten Van Segbroeck, Jangwon Kim, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 10 |
| 2015 | Automatic estimation of parkinson's disease severity from diverse speech tasksabstractThe need for reliable, scalable and efficient diagnosis of Parkin-son’s Disease (PD) is a major clinical need. Automating the diagnosis can lead to more accurate and objective predictions as well as provide insights regarding the nature of Parkinson’s condition. This paper proposes a fully automated system to rate the severity (UPDRS-III scale) of PD from patients ’ speech. Specifically, the system captures atypicalities in an individ-ual’s voice when performing multiple diverse speaking tasks and makes a unified prediction of the PD severity. The perfor-mance is tested in a cross-data setting, with different subjects and dissimilar recording conditions. Results indicate that (i) effective features vary depending on the nature of the specific speech task, (ii) additional novel feature sets to detect distor-tions in Parkinson’s speech significantly improve the prediction accuracy from the Interspeech15 Challenge baseline system and (iii) our fusion system based on an unsupervised clustering tech-nique also improves the accuracy. Our system incorporates i-vector and functionals for segmental features, non-linear time series features, speech rhythm and automatic speech recogni-tion decoding based features. By its application on the Inter-speech15 eating condition challenge, the system also shows its potential for detecting other sources of speech variability. Jangwon Kim, Md. Nasir, Rahul Gupta 0001, Maarten Van Segbroeck, Daniel Bone, Matthew Black, Z.-I. Skordilis, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 1 |
| 2015 | Automatic intelligibility classification of sentence-level pathological speech
Jangwon Kim, Naveen Kumar 0004, Andreas Tsiartas, Ming Li 0026, Shri Narayanan |
Comput. Speech Lang. | 1 |
| 2014 | A study of invariant properties and variation patterns in the converter/distributor model for emotional speechabstractInvariant properties of vocal organ controls at an abstract level are crucial for better understanding and modeling of the speech production mechanism. Despite the large variability of articulatory movements at the execution level, the Converter/Distributor (C/D) model provides a systematic and comprehensive framework for the prosodic organization of speech production, based on the invariant properties of articulatory movements with the concept of “iceberg” region. The goal of this paper is two-fold: (i) to examine the invariant properties in the C/D model in emotional speech, and (ii) to understand emotion-dependent variation patterns of important parameters in the C/D model framework. Experimental results support the validity of strong linear relationship between the speed and excursion of critical articulatorsat the iceberg points for emotional speech. Also, emotion-dependent variation patterns of the C/D model parameters, (e.g., relatively smaller “shadow” angle and greater syllable magnitude for happiness) are reported. Finally, the emotion-dependent relationships between the abstract-level C/D model parameters and the surface-level parameters of the invariant articulatory behaviors are reported. Index Terms: emotional variation, temporal organization of speech, C/D model, invariant property Jangwon Kim, Donna Erickson, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 1 |
| 2014 | Estimation of the movement trajectories of non-crucial articulators based on the detection of crucial moments and physiological constraintsabstractThis study develops a mathematical model that estimates the movements of (linguistically) non-crucial articulators in speech production, which provides a systematic way to study the relationship between the behaviors of crucial and non-crucial articulators; crucial articulators are those essential for realizing a speech task. The underlying assumption of our model is that non-crucial articulatory movements are governed by the physiological constraints in relation to the corresponding crucial articulators as well as by the contextual constraint from the nearest crucial time of the non-crucial articulator. These constraints have been generally assumed in the speech production literature, but they have not been incorporated directly into articulatory models. The crucial articulatory moments in an utterance are automatically determined by a novel forced-alignment algorithm for articulatory trajectories, which uses the inherent physical properties of crucial articulatory movements. Experimental results suggest that the proposed algorithm is capable of estimating non-crucial articulatory positions well in both neutral and emotional speech, significantly better than the simple interpolation of crucial points. Index Terms: non-crucial articulators, articulatory modeling, emotional speech Jangwon Kim, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 1 |
| 2014 | Classification of cognitive load from speech using an i-vector frameworkabstractThe goal in this work is to automatically classify speakers ’ level of cognitive load (low, medium, high) from a standard battery of reading tasks requiring varying levels of working memory. This is a challenging machine learning problem because of the inherent difficulty in defining/measuring cognitive load and due to intra-/inter-speaker differences in how their effects are man-ifested in behavioral cues. We experimented with a number of static and dynamic features extracted directly from the audio signal (prosodic, spectral, voice quality) and from automatic speech recognition hypotheses (lexical information, speaking rate). Our approach to classification addressed the wide vari-ability and heterogeneity through speaker normalization and by adopting an i-vector framework that affords a systematic way to factorize the multiple sources of variability. Index Terms: computational paralinguistics, behavioral signal processing (BSP), prosody, ASR, i-vector, cognitive load Maarten Van Segbroeck, Ruchir Travadi, Colin Vaz, Jangwon Kim, Matthew Black, Alexandros Potamianos, Shri Narayanan |
INTERSPEECH | 4 |
| 2013 | Spatial and temporal alignment of multimodal human speech production data: Real time imaging, flesh point tracking and audioabstractIn speech production research, the integration of articulatory data derived from multiple measurement modalities can provide rich description of vocal tract dynamics by overcoming the limited spatio-temporal representations offered by individual modalities. This paper presents a spatial and temporal alignment method between two promising modalities using a corpus of TIMIT sentences obtained from the same speaker: flesh point tracking from Electromagnetic Articulography (EMA) that offers high temporal resolution but sparse spatial information and real time Magnetic Resonance Imaging (MRI) that offers good spatial details but at lower temporal rates. Spatial alignment is done by using palate tracking of EMA, but distortion in MRI audio and articulatory data variability make temporal alignment challenging. This paper proposes a novel alignment technique using joint acoustic-articulatory features which combines dynamic time warping and automatic feature extraction from MRI images. Experimental results show that the temporal alignment obtained using this technique is better (12% relative) than that using acoustic feature only. Jangwon Kim, Adam C. Lammert, Prasanta Kumar Ghosh, Shri Narayanan |
ICASSP | 1 |
| 2013 | Speaker verification based on fusion of acoustic and articulatory informationabstractWe propose a practical, feature-level fusion approach for com-bining acoustic and articulatory information in speaker ver-ification task. We find that concatenating articulation fea-tures obtained from the measured speech production data with conventional Mel-frequency cepstral coefficients (MFCCs) im-proves the overall speaker verification performance. However, since access to the measured articulatory data is impractical for real world speaker verification applications, we also ex-periment with estimated articulatory features obtained using acoustic-to-articulatory inversion technique. Specifically, we show that augmenting MFCCs with articulatory features ob-tained from subject-independent acoustic-to-articulatory inver-sion technique also significantly enhances the speaker verifi-cation performance. This performance boost could be due to the information about inter-speaker variation present in the es-timated articulatory features, especially at the mean and vari-ance level. Experimental results on the Wisconsin X-Ray Mi-crobeam database show that the proposed acoustic-estimated-articulatory fusion approach significantly outperforms the tra-ditional acoustic-only baseline, providing up to 10 % relative re-duction in Equal Error Rate (EER). We further show that we can achieve an additional 5 % relative reduction in EER after score-level fusion. Index Terms: speech production, speaker verification, articula-tion features, acoustic-to-articulatory inversion, biometrics Ming Li 0026, Jangwon Kim, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan |
INTERSPEECH | 2 |
| 2012 | Intelligibility classification of pathological speech using fusion of multiple high level descriptors
Jangwon Kim, Naveen Kumar 0004, Andreas Tsiartas, Ming Li 0026, Shri Narayanan |
INTERSPEECH | 1 |
| 2011 | An Exploratory Study of the Relations Between Perceived Emotion Strength and Articulatory KinematicsabstractAcoustic and articulatory behaviors underlying emotion strength perception are studied by analyzing acted emotional speech. Listeners evaluated emotion identity, strength and con-fidence. Parameters related to pitch, loudness and articulatory kinematics are associated with a 2-level (strong/weak) represen-tation of the emotion strength. Two-class discriminant analyses show averaged leave-one-out accuracies of 65.8 % and 63.8 % in the acoustic and articulatory domains, respectively. Two-factor ANOVA (emotion type/strength) indicates that the listeners as-sess the emotion strength based on the nature of perceived emo-tions in the arousal dimension. Only hot anger and happiness show significant differences in pitch use in the strength contrast. Such contrasts are also observed in tongue lowering and/or ad-vancing. The strength contrast by listeners may mainly rely upon pitch and loudness. However, interactions between the acoustic and articulatory parameters in strength perception are complex. Jangwon Kim, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 1 |
| 2010 | An exploratory study of manifolds of emotional speechabstractThis study explores manifold representations of emotionally modulated speech. The manifolds are derived in the articulatory space and two acoustic spaces (MFB and MFCC) using isometric feature mapping (Isomap) with data from an emotional speech corpus. Their effectiveness in representing emotional speech is tested based on the emotion classification accuracy. Results show that the effective manifold dimensions of the articulatory and MFB spaces are both about 5 while being greater in MFCC space. Also, the accuracies in the articulatory and MFB manifolds are close to those in the original spaces, but this is not the case for the MFCC. It is speculated that the manifold in the MFCC space is less structured, or more distorted, than others. Jangwon Kim, Sungbok Lee, Shri Narayanan |
ICASSP | 1 |
| 2010 | A study of interplay between articulatory movement and prosodic characteristics in emotional speech productionabstractThis paper investigates the interplay between articulatory move-ment and voice source activity as a function of emotions in speech production. Our hypothesis is that humans use differ-ent modulation methods in which articulatory movements and prosodic modulations are differently weighted across different emotions. This hypothesis was examined by joint analysis of the two domains, using two statistical representations: (1) the sample distribution comparison using two-sigma ellipses of the articulatory speed statistics and prosodic feature (pitch or inten-sity) statistics, (2) the comparison of correlation coefficients. In the articulatory-prosodic spaces, we find (1) distinctive weight-ing patterns for angry and happy emotional speech and (2) dis-tinctive correlation patterns depending on articulators and target emotions. These findings support the hypothesis that humans use different modulation methods of emphasizing articulatory motions and/or prosodic activities depending on emotion. Index Terms: interplay, articulatory control, prosodic control, emotional speech production, joint analysis Jangwon Kim, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 1 |
| 2009 | A detailed study of word-position effects on emotion expression in speechabstractWe investigate emotional effects on articulatory-acoustic speech characteristics with respect to word location within a sentence. We examined the hypothesis that emotional effect will vary based on word position by first examining articulatory features manually extracted from Electromagnetic articulography data. Initial articulatory data analyses indicated that the emotional effects on sentence medial words are significantly stronger than on initial words. To verify that observation further, we expanded our hypothesis testing to include both acoustic and articulatory data, and a consideration of an expanded set of words from different locations. Results suggest that emotional effects are generally more significant on sentence medial words than sentence initial and final words. This finding suggests that word location needs to be considered as a factor in emotional speech processing. Jangwon Kim, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 1 |