June-Woo Kim

dblp:227/2585 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0003-0111-300XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Noise-Agnostic Multitask Whisper Training for Reducing False Alarm Errors in Call-for-Help Detection
abstract
Keyword spotting is often implemented by keyword classifier to the encoder in acoustic models, enabling the classification of predefined or open vocabulary keywords. Although keyword spotting is a crucial task in various applications and can be extended to call-for-help detection in emergencies, however, the previous method often suffers from scalability limitations due to retraining required to introduce new keywords or adapt to changing contexts. We explore a simple yet effective approach that leverages off-the-shelf pretrained ASR models to address these challenges, especially in call-for-help detection scenarios. Furthermore, we observed a substantial increase in false alarms when deploying call-for-help detection system in real-world scenarios due to noise introduced by microphones or different environments. To address this, we propose a novel noise-agnostic multitask learning approach that integrates a noise classification head into the ASR encoder. Our method enhances the model’s robustness to noisy environments, leading to a significant reduction in false alarms and improved overall call-for-help performance. Despite the added complexity of multitask learning, our approach is computationally efficient and provides a promising solution for call-for-help detection in real-world scenarios.
Myeonghoon Ryu, June-Woo Kim, Minseok Oh, Suji Lee, Han Park
ICASSP2
2025 Language-Agnostic Suicidal Risk Detection Using Large Language Models
June-Woo Kim, Wonkyo Oh, Haram Yoon, Sung-Hoon Yoon 0002, Sang-Yeol Lee, Chan-Mo Yang
INTERSPEECH1
2025 Improving Respiratory Sound Classification with Architecture-Agnostic Knowledge Distillation from Ensembles
Miika Toikkanen, June-Woo Kim
INTERSPEECH2
2025 Adaptive Metadata-Guided Supervised Contrastive Learning for Domain Adaptation on Respiratory Sound Classification
abstract
Despite considerable advancements in deep learning, optimizing respiratory sound classification (RSC) models remains challenging. This is partly due to the bias from inconsistent respiratory sound recording processes and imbalanced representation of demographics, which leads to poor performance when a model trained with the dataset is applied to real-world use cases. RSC datasets usually include various metadata attributes describing certain aspects of the data, such as environmental and demographic factors. To address the issues caused by bias, we take advantage of the metadata provided by RSC datasets and explore approaches for metadata-guided domain adaptation. We thoroughly evaluate the effect of various metadata attributes and their combinations on a simple metadata-guided approach, but also introduce a more advanced method that adaptively rescales the suitable metadata combinations to improve domain adaptation during training. The findings indicate a robust reduction in domain dependency and improvement in detection accuracy on both ICBHI and our own dataset. Specifically, the implementation of our proposed methods led to an improved score of 84.97%, which signifies a substantial enhancement of 7.37% compared to the baseline model.
June-Woo Kim, Miika Toikkanen, Amin Jalali 0003, Hye-Ji Han, Wonwoo Shin, Ho-Young Jung
IEEE J. Biomed. Health Informatics1
2024 Stethoscope-Guided Supervised Contrastive Learning for Cross-Domain Adaptation on Respiratory Sound Classification
abstract
Despite the remarkable advances in deep learning technology, achieving satisfactory performance in lung sound classification remains a challenge due to the scarcity of available data. Moreover, the respiratory sound samples are collected from a variety of electronic stethoscopes, which could potentially introduce biases into the trained models. When a significant distribution shift occurs within the test dataset or in a practical scenario, it can substantially decrease the performance. To tackle this issue, we introduce cross-domain adaptation techniques, which transfer the knowledge from a source domain to a distinct target domain. In particular, by considering different stethoscope types as individual domains, we propose a novel stethoscope-guided supervised contrastive learning approach. This method can mitigate any domain-related disparities and thus enables the model to distinguish respiratory sounds of the recording variation of the stethoscope. The experimental results on the ICBHI dataset demonstrate that the proposed methods are effective in reducing the domain dependency and achieving the ICBHI Score of 61.71%, which is a significant improvement of 2.16% over the baseline.
June-Woo Kim, Sangmin Bae, Won-Yang Cho, Byungjo Lee, Ho-Young Jung
ICASSP1
2024 BTS: Bridging Text and Sound Modalities for Metadata-Aided Respiratory Sound Classification
June-Woo Kim, Miika Toikkanen, Yera Choi, Seoung-Eun Moon, Ho-Young Jung
INTERSPEECH1
2023 Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
abstract
Respiratory sound contains crucial information for the early diagnosis of fatal lung diseases.Since the COVID-19 pandemic, there has been a growing interest in contact-free medical care based on electronic stethoscopes.To this end, cutting-edge deep learning models have been developed to diagnose lung diseases; however, it is still challenging due to the scarcity of medical data.In this study, we demonstrate that the pretrained model on large-scale visual and audio datasets can be generalized to the respiratory sound classification task.In addition, we introduce a straightforward Patch-Mix augmentation, which randomly mixes patches between different samples, with Audio Spectrogram Transformer (AST).We further propose a novel and effective Patch-Mix Contrastive Learning to distinguish the mixed representations in the latent space.Our method achieves state-of-the-art performance on the ICBHI dataset, outperforming the prior leading score by an improvement of 4.08%.
Sangmin Bae, June-Woo Kim, Won-Yang Cho, Hyerim Baek, Soyoun Son, Byungjo Lee, Changwan Ha, Kyongpil Tae, Sungnyun Kim, Se-Young Yun
INTERSPEECH2
2020 Vocoder-free End-to-End Voice Conversion with Transformer Network
abstract
Mel-frequency filter bank (MFB) based approaches have the advantage of higher learning speeds compared to using the raw spectrum due to a smaller number of features. However, speech generators with the MFB approach require an additional computationally expensive vocoder for the training process. The pre- and post-processing needed by the MFB and the vocoder is not essential to convert human voices, because it is possible to use only the raw spectrum to generate different style of voices with clear pronunciation. In this paper, we introduce a vocoder-free end-to-end voice conversion method using a transformer network to alleviate the computational burden from additional pre- and post-processing. Our transformer-based architecture, which does not have any CNN or RNN layers, has shown the benefit of learning fast while solving the limitation of sequential computation of the conventional RNN. For this reason, our model is a fast and effective approach to convert realistic voices using raw spectra in a parallel manner to generate different style of voices with clear pronunciation. Furthermore, we can get an adapted MFB for speech recognition by multiplying the converted magnitude with the phase information, and therefore our conversion model is also suitable for speaker adaptation. We perform our voice conversion experiments on TIDIGITS-dataset using the naturalness, similarity, and clarity with Mean Opinion Score as metrics1.
June-Woo Kim, Ho-Young Jung, Minho Lee 0001
IJCNN1