EDBT 2026 Demo / reviewers in the wild / expert
Ju-Seok Seong
dblp:330/9030
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2025
0009-0002-7009-9715ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Trainable Adaptive Score Normalization for Automatic Speaker VerificationabstractAdaptive S-norm (AS-norm) calibrates automatic speaker verification (ASV) scores by normalizing them utilize the scores of impostors which are similar to the input speaker. However, AS-norm does not involve any learning process, limiting its ability to provide appropriate regularization strength for various evaluation utterances. To address this limitation, we propose a trainable AS-norm (TAS-norm) that leverages learnable impostor embeddings (LIEs), which are used to compose the cohort. These LIEs are initialized to represent each speaker in a training dataset consisting of impostor speakers. Subsequently, LIEs are fine-tuned by simulating an ASV evaluation. We utilize a margin penalty during top-scoring IEs selection in fine-tuning to prevent non-impostor speakers from being selected. In our experiments with ECAPA-TDNN, the proposed TAS-norm observed 4.11% and 10.62% relative improvement in equal error rate and minimum detection cost function, respectively, on VoxCeleb1-O trial compared with standard AS-norm without using proposed LIEs. We further validated the effectiveness of the TAS-norm on additional ASV datasets comprising Persian and Chinese, demonstrating its robustness across different languages. Jeong-Hwan Choi, Ju-Seok Seong, Ye-Rin Jeoung, Joon-Hyuk Chang |
ICASSP | 2 |
| 2025 | Few-shot Keyword-incremental Learning Using Compositional InformationabstractRecognizing not only pre-defined keywords but also continuously expanding new keywords often with limited data has emerged as a main problem in recent keyword spotting research. To address this challenge few-shot class-incremental learning approaches have gained attention initially training models on sufficient data in a base session and then continuously adapting to recognize new classes with limited data. Recent focus has been on prototype-based calibration which fuses new prototypes with weighted base prototypes. However this method risks misclassification due to increased similarity between new and base classes. To mitigate this issue we propose a compositional feature-based calibration method. Instead of directly using base prototypes our approach extracts and utilizes rich compositional information from the initial session to enhance new class representations. Experimental results on two keyword spotting datasets demonstrate the superiority of our proposed method showing improved performance in recognizing initial and new keywords. Ilseok Kim, Ju-Seok Seong, Joon-Hyuk Chang |
ICASSP | 2 |
| 2025 | Enhancing Target-speaker Automatic Speech Recognition Using Multiple Speaker Embedding Extractors with Virtual Speaker Embedding
Ju-Seok Seong, Jeong-Hwan Choi, Ye-Rin Jeoung, Ilseok Kim, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2024 | Few-Shot Keyword-Incremental Learning with Total Calibration
Ilseok Kim, Ju-Seok Seong, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2023 | Extending Self-Distilled Self-Supervised Learning For Semi-Supervised Speaker VerificationabstractIn this study, we extend self-distillation with no labels (DINO), a successful self-supervised learning framework, by combining it with supervised classification (SC) for semi-supervised speaker verification with limited labeled data. We introduce a transfer learning framework that pre-trains and fine-tunes the encoder using DINO and SC, respectively, and a multitask learning framework that shares the encoder while having separate projection layers for both methods. To achieve lower inter-speaker similarity, we propose a joint learning framework sharing both the encoder and projection layer for DINO and SC. We also propose an auxiliary contrastive loss between embeddings derived from labeled and unlabeled utterances and introduce a two-stage learning strategy to apply margin penalty effectively. Experimental results on the VoxCeleb corpus indicate that the joint learning framework outperforms the other frameworks and is closest to achieving the performance of fully supervised learning. Jeong-Hwan Choi, Jehyun Kyung, Ju-Seok Seong, Ye-Rin Jeoung, Joon-Hyuk Chang |
ASRU | 3 |
| 2023 | Noise-Aware Target Extension with Self-Distillation for Robust Speech RecognitionabstractData augmentation using additive noise is a framework for robustly training automatic speech recognition models. To utilize noise information efficiently, previous studies used an additional branch to classify noise conditions. This added branch has a limited effect on the ASR because it performs independently of the ASR branch that classifies senones. In this paper, we propose a noise-aware target extension (NATE) that extends the senone target to contain noise awareness by jointly classifying the senone and noise in a single branch. In the inference stage, the output of the model is processed separately by the noise condition and then aggregated to match the senone posterior distribution. In addition, we combine NATE with self-distillation (NATEsd) to reduce the model parameters and avoid discrepancies between the outputs of training and inference. The effectiveness of the NATE method is validated on the two benchmark development and evaluation sets and simulated noisy test sets, resulting in significant improvements over the previous methods. Ju-Seok Seong, Jeong-Hwan Choi, Jehyun Kyung, Ye-Rin Jeoung, Joon-Hyuk Chang |
ICASSP | 1 |
| 2023 | Self-Distillation into Self-Attention Heads for Improving Transformer-based End-to-End Neural Speaker Diarization
Ye-Rin Jeoung, Jeong-Hwan Choi, Ju-Seok Seong, Jehyun Kyung, Joon-Hyuk Chang |
INTERSPEECH | 3 |
| 2023 | Improving Joint Speech and Emotion Recognition Using Global Style Tokens
Jehyun Kyung, Ju-Seok Seong, Jeong-Hwan Choi, Ye-Rin Jeoung, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2022 | Regularizing Transformer-based Acoustic Models by Penalizing Attention Weights
Mun-Hak Lee, Joon-Hyuk Chang, Sang-Eon Lee, Ju-Seok Seong, Chanhee Park, Haeyoung Kwon |
INTERSPEECH | 4 |