EDBT 2026 Demo / reviewers in the wild / expert
You Jin Kim
dblp:177/0948
· DBLP profile ↗
15ranked-venue papers
4as first author
13since 2021 · last 2024
0000-0002-4952-5532ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker VerificationabstractIn the field of speaker verification, session or channel variability poses a significant challenge. While many contemporary methods aim to disentangle session information from speaker embeddings, we introduce a novel approach using an additional embedding to represent the session information. This is achieved by training an auxiliary network appended to the speaker embedding extractor which remains fixed in this training process. This results in two similarity scores: one for the speakers information and one for the session information. The latter score acts as a compensator for the former that might be skewed due to session variations. Our extensive experiments demonstrate that session information can be effectively compensated without retraining of the embedding extractor. Hee-Soo Heo, Kihyun Nam, Bong-Jin Lee, Youngki Kwon, You Jin Kim, Joon Son Chung |
ICASSP | 6 |
| 2024 | TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive LearningabstractThe goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures while learning effective representations has been less explored. In this work, we propose TalkNCE, a novel talk-aware contrastive loss. The loss is only applied to part of the full segments where a person on the screen is actually speaking. This encourages the model to learn effective representations through the natural correspondence of speech and facial movements. Our loss can be jointly optimized with the existing objectives for training ASD models without the need for additional supervision or training data. The experiments demonstrate that our loss can be easily integrated into the existing ASD frameworks, improving their performance. Our method achieves state-of-the-art performances on AVA-ActiveSpeaker and ASW datasets. Chaeyoung Jung, Suyeon Lee, Kihyun Nam, Kyeongha Rho, You Jin Kim, Youngjoon Jang 0001, Joon Son Chung |
ICASSP | 5 |
| 2023 | High-Resolution Embedding Extractor for Speaker DiarisationabstractSpeaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the sliding window approach, a segment easily includes two or more speakers owing to speaker change points. This study proposes a novel embedding extractor architecture, referred to as a high-resolution embedding extractor (HEE), which extracts multiple high-resolution embeddings from each speech segment. Hee consists of a feature-map extractor and an enhancer, where the enhancer with the self-attention mechanism is the key to success. The enhancer of HEE replaces the aggregation process; instead of a global pooling layer, the enhancer combines relative information to each frame via attention leveraging the global context. Extracted dense frame-level embeddings can each represent a speaker. Thus, multiple speakers can be represented by different frame-level features in each segment. We also propose an artificially generating mixture data training framework to train the proposed HEE. Through experiments on five evaluation sets, including four public datasets, the proposed HEE demonstrates at least 10% improvement on each evaluation set, except for one dataset, which we analyse that rapid speaker changes less exist. Hee-Soo Heo, Youngki Kwon, Bong-Jin Lee, You Jin Kim, Jee-Weon Jung |
ICASSP | 4 |
| 2023 | Advancing the Dimensionality Reduction of Speaker Embeddings for Speaker Diarisation: Disentangling Noise and Informing Speech ActivityabstractThe objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as noise, adversely affecting performance. Our previous work has proposed an auto-encoder-based dimensionality reduction module to help remove the redundant information. However, they do not explicitly separate such information and have also been found to be sensitive to hyper-parameter values. To this end, we propose two contributions to overcome these issues: (i) a novel dimensionality reduction framework that can disentangle spurious information from the speaker embeddings; (ii) the use of speech activity vector to prevent the speaker code from representing the background noise. Through a range of experiments conducted on four datasets, our approach consistently demonstrates the state-of-the-art performance among models without system fusion. You Jin Kim, Hee-Soo Heo, Jee-Weon Jung, Youngki Kwon, Bong-Jin Lee, Joon Son Chung |
ICASSP | 1 |
| 2023 | Absolute Decision Corrupts Absolutely: Conservative Online Speaker DiarisationabstractOur focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few misjudgements in the early phase of an input session can lead to catastrophic results. We hypothesise that cautiously increasing the number of estimated speakers is of paramount importance among many other factors. Thus, our proposed framework includes decreasing the number of speakers by one when the system judges that an increase in the past was faulty. We also adopt dual buffers, checkpoints and centroids, where checkpoints are combined with silhouette coefficients to estimate the number of speakers and centroids represent speakers. Again, we believe that more than one centroid can be generated from one speaker. Thus we design a clustering-based label matching technique to assign labels in realtime. The resulting system is lightweight yet surprisingly effective. The system demonstrates state-of-the-art performance on DIHARD II and III datasets, where it is also competitive in AMI and VoxConverse test sets. Youngki Kwon, Hee-Soo Heo, Bong-Jin Lee, You Jin Kim, Jee-Weon Jung |
ICASSP | 4 |
| 2023 | Curriculum Learning for Self-supervised Speaker Verification
Hee-Soo Heo, Jee-Weon Jung, Youngki Kwon, Bong-Jin Lee, You Jin Kim, Joon Son Chung |
INTERSPEECH | 6 |
| 2023 | Encoder-decoder Multimodal Speaker Change Detection
Jee-Weon Jung, Soonshin Seo, Hee-Soo Heo, Geonmin Kim, You Jin Kim, Youngki Kwon, Bong-Jin Lee |
INTERSPEECH | 5 |
| 2022 | Multi-Scale Speaker Embedding-Based Graph Attention Networks For Speaker DiarisationabstractThe objective of this work is effective speaker diarisation using multi-scale speaker embeddings. Typically, there is a trade-off between the ability to recognise short speaker segments and the discriminative power of the embedding, according to the segment length used for embedding extraction. To this end, recent works have proposed the use of multi-scale embeddings where segments with varying lengths are used. However, the scores are combined using a weighted summation scheme where the weights are fixed after the training phase, whereas the importance of segment lengths can differ within a single session.To address this issue, we present three key contributions in this paper: (1) we propose graph attention networks for multi-scale speaker diarisation; (2) we design scale indicators to utilise scale information of each embedding; (3) we adapt the attention-based aggregation to utilise a pre-computed affinity matrix from multi-scale embeddings.We demonstrate the effectiveness of our method in various datasets where the speaker confusion which constitutes the primary metric drops over 10% in average relative compared to the baseline. Youngki Kwon, Hee-Soo Heo, Jee-Weon Jung, You Jin Kim, Bong-Jin Lee, Joon Son Chung |
ICASSP | 4 |
| 2022 | Pushing the limits of raw waveform speaker recognitionabstractIn recent years, speaker recognition systems based on raw waveform inputs have received increasing attention.However, the performance of such systems are typically inferior to the state-of-the-art handcrafted feature-based counterparts, which demonstrate equal error rates under 1% on the popular VoxCeleb1 test set.This paper proposes a novel speaker recognition model based on raw waveform inputs.The model incorporates recent advances in machine learning and speaker verification, including the Res2Net backbone module and multi-layer feature aggregation.Our best model achieves an equal error rate of 0.89%, which is competitive with the state-of-the-art models based on handcrafted features, and outperforms the best model based on raw waveform inputs by a large margin.We also explore the application of the proposed model in the context of self-supervised learning framework.Our self-supervised model outperforms single phase-based existing works in this line of research.Finally, we show that self-supervised pre-training is effective for the semi-supervised scenario where we only have a small set of labelled training data, along with a larger set of unlabelled examples. Jee-Weon Jung, You Jin Kim, Hee-Soo Heo, Bong-Jin Lee, Youngki Kwon, Joon Son Chung |
INTERSPEECH | 2 |
| 2022 | Digestive Organ Recognition in Video Capsule Endoscopy Based on Temporal Segmentation Network
Yejee Shin, Taejoon Eo, Hyeongseop Rha, Dong Jun Oh, Geonhui Son, Jiwoong An, You Jin Kim, Dosik Hwang, Yun Jeong Lim |
MICCAI (8) | 7 |
| 2021 | Look Who's Talking: Active Speaker Detection in the WildabstractIn this work, we present a novel audio-visual dataset for active speaker detection in the wild.A speaker is considered active when his or her face is visible and the voice is audible simultaneously.Although active speaker detection is a crucial pre-processing step for many audio-visual tasks, there is no existing dataset of natural human speech to evaluate the performance of active speaker detection.We therefore curate the Active Speakers in the Wild (ASW) dataset which contains videos and co-occurring speech segments with dense speech activity labels.Videos and timestamps of audible segments are parsed and adopted from VoxConverse, an existing speaker diarisation dataset that consists of videos in the wild.Face tracks are extracted from the videos and active segments are annotated based on the timestamps of VoxConverse in a semi-automatic way.Two reference systems, a self-supervised system and a fully supervised one, are evaluated on the dataset to provide the baseline performances of ASW.Cross-domain evaluation is conducted in order to show the negative effect of dubbed videos in the training data. You Jin Kim, Hee-Soo Heo, Soyeon Choe, Soo-Whan Chung, Yoohwan Kwon, Bong-Jin Lee, Youngki Kwon, Joon Son Chung |
Interspeech | 1 |
| 2021 | Adapting Speaker Embeddings for Speaker DiarisationabstractThe goal of this paper is to adapt speaker embeddings for solving the problem of speaker diarisation. The quality of speaker embeddings is paramount to the performance of speaker diarisation systems. Despite this, prior works in the field have directly used embeddings designed only to be effective on the speaker verification task. In this paper, we propose three techniques that can be used to better adapt the speaker embeddings for diarisation: dimensionality reduction, attention-based embedding aggregation, and non-speech clustering. A wide range of experiments is performed on various challenging datasets. The results demonstrate that all three techniques contribute positively to the performance of the diarisation system achieving an average relative improvement of 25.07% in terms of diarisation error rate over the baseline. Youngki Kwon, Jee-Weon Jung, Hee-Soo Heo, You Jin Kim, Bong-Jin Lee, Joon Son Chung |
Interspeech | 4 |
| 2021 | End-To-End Lip Synchronisation Based on Pattern ClassificationabstractThe goal of this work is to synchronise audio and video of a talking face using deep neural network models. Existing works have trained networks on proxy tasks such as cross-modal similarity learning, and then computed similarities between audio and video frames using a sliding window approach. While these methods demonstrate satisfactory performance, the networks are not trained directly on the task. To this end, we propose an end-to-end trained network that can directly predict the offset between an audio stream and the corresponding video stream. The similarity matrix between the two modalities is first computed from the features, then the inference of the offset can be considered to be a pattern recognition problem where the matrix is considered equivalent to an image. The feature extractor and the classifier are trained jointly. We demonstrate that the proposed approach outperforms the previous work by a large margin on LRS2 and LRS3 datasets. You Jin Kim, Hee-Soo Heo, Soo-Whan Chung, Bong-Jin Lee |
SLT | 1 |
| 2018 | Interpretable Prediction of Vascular Diseases from Electronic Health Records via Deep Attention NetworksabstractPrecise prediction of severe diseases resulting in mortality is one of the main issues in medical fields. Even if pathological and radiological measurements provide competitive precision, they usually require large costs of time and expense to obtain and analyze the data for prediction. Recently, end-to-end approaches based on deep neural networks have been proposed, however, they still suffer from the low classification performance and difficulties of interpretation. In this study, we propose a novel disease prediction method, EHAN (EHR History-based prediction using Attention Network), based on the recurrent neural network (RNN) and attention mechanism. The proposed method incorporates (1) a bidirectional gated recurrent units (GRU) for automated sequential modeling, (2) attention mechanism for improving long-term dependence modeling, (3) RNN-based gradient-weighted class activation mapping (Grad-CAM) to visualize the class specific attention-weights. We conducted the experiments to predict the occurrence of risky disease containing cardiovascular and cerebrovascular diseases from more than 40,000 hypertension patients' electronic health records (EHR). The results showed that the proposed method outperformed the state-of-the-art model with respect to the various performance metrics. Furthermore, we confirmed that the proposed visualizing methods can be used to assist data-driven discovery. Seunghyun Park 0001, You Jin Kim, Jeong-Whun Kim, Jin Joo Park, Borim Ryu, Jung-Woo Ha 0001 |
BIBE | 2 |
| 2017 | Model Regularization of Deep Neural Networks for Robust Clinical Opinions Generation from General Blood Test ResultsabstractThe deep neural network (DNN) that models characteristics of general blood test (GBT) results was used in clinical opinions generation. The DNN that generates clinical opinions has the complex structure, which causes overfitting problem. The relatively small size of medical dataset also contributes to the occurrence of overfitting. In order to deal with overfitting, we apply two techniques that solve overfitting of DNN, which are dropout, and batch normalization. Dropout is inserted into the network in various ways in order to find out the optimal structure of the network. Batch normalization is also added in various ways for the same purpose. The experiment conducted on GBT dataset shows that DNNs with dropout and batch normalization outperform the simple DNN in generating clinical opinions for our GBT dataset. Besides, dropout shows slightly better performance compared to batch normalization. You Jin Kim, Han-Gyu Kim, Ho-Jin Choi |
MDM | 1 |