VLDB 2026 Research / reviewers in the wild / expert
Jeong-Hwan Choi
dblp:277/2489
· DBLP profile ↗
15ranked-venue papers
8as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Trainable Adaptive Score Normalization for Automatic Speaker VerificationabstractAdaptive S-norm (AS-norm) calibrates automatic speaker verification (ASV) scores by normalizing them utilize the scores of impostors which are similar to the input speaker. However, AS-norm does not involve any learning process, limiting its ability to provide appropriate regularization strength for various evaluation utterances. To address this limitation, we propose a trainable AS-norm (TAS-norm) that leverages learnable impostor embeddings (LIEs), which are used to compose the cohort. These LIEs are initialized to represent each speaker in a training dataset consisting of impostor speakers. Subsequently, LIEs are fine-tuned by simulating an ASV evaluation. We utilize a margin penalty during top-scoring IEs selection in fine-tuning to prevent non-impostor speakers from being selected. In our experiments with ECAPA-TDNN, the proposed TAS-norm observed 4.11% and 10.62% relative improvement in equal error rate and minimum detection cost function, respectively, on VoxCeleb1-O trial compared with standard AS-norm without using proposed LIEs. We further validated the effectiveness of the TAS-norm on additional ASV datasets comprising Persian and Chinese, demonstrating its robustness across different languages. Jeong-Hwan Choi, Ju-Seok Seong, Ye-Rin Jeoung, Joon-Hyuk Chang |
ICASSP | 1 |
| 2025 | Enhancing Target-speaker Automatic Speech Recognition Using Multiple Speaker Embedding Extractors with Virtual Speaker Embedding
Ju-Seok Seong, Jeong-Hwan Choi, Ye-Rin Jeoung, Ilseok Kim, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2024 | Efficient Speaker Embedding Extraction Using a Twofold Sliding Window Algorithm for Speaker Diarization
Jeong-Hwan Choi, Ye-Rin Jeoung, Ilseok Kim, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2024 | Efficient Lightweight Speaker Verification With Broadcasting CNN-Transformer and Knowledge Distillation Training of Self-Attention MapsabstractDeveloping a lightweight speaker embedding extractor (SEE) is crucial for the practical implementation of automatic speaker verification (ASV) systems. To this end, we recently introducedbroadcasting convolutional neural networks (CNNs)-meet-vision-Transformers(BC-CMT), a lightweight SEE that utilizes broadcasted residual learning (BRL) within the hybrid CNN-Transformer architecture to maintain a small number of model parameters. We proposed three BC-CMT-based SEE with three different sizes: BC-CMT-Tiny, -Small, and -Base. In this study, we extend our previously proposed BC-CMT by introducing an improved model architecture and a training strategy based on knowledge distillation (KD) using self-attention (SA) maps. First, to reduce the computational costs and latency of the BC-CMT, the two-dimensional (2D) SA operations in the BC-CMT, which calculate the SA maps in the frequency–time dimensions, are simplified to 1D SA operations that consider only temporal importance. Moreover, to enhance the SA capability of the BC-CMT, the group convolution layers in the SA block are adjusted to have smaller number of groups and are combined with the BRL operations. Second, to improve the training effectiveness of the modified BC-CMT-Tiny, the SA maps of a pretrained large BC-CMT-Base are used for the KD to guide those of a smaller BC-CMT-Tiny. Because the attention map sizes of the modified BC-CMT models do not depend on the number of frequency bins or convolution channels, the proposed strategy enables KD between feature maps with different sizes. The experimental results demonstrate that the proposed BC-CMT-Tiny model having 271.44K model parameters achieved 36.8% and 9.3% reduction in floating point operations on 1s signals and equal error rate (EER) on VoxCeleb 1 testset, respectively, compared to the conventional BC-CMT-Tiny. The CPU and GPU running time of the proposed BC-CMT-Tiny ranges of 1 to 10 s signals were 29.07 to 146.32 ms and 36.01 to 206.43 ms, respectively. The proposed KD further reduced the EER by 15.5% with improved attention capability. Jeong-Hwan Choi, Joon-Young Yang, Joon-Hyuk Chang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Extending Self-Distilled Self-Supervised Learning For Semi-Supervised Speaker VerificationabstractIn this study, we extend self-distillation with no labels (DINO), a successful self-supervised learning framework, by combining it with supervised classification (SC) for semi-supervised speaker verification with limited labeled data. We introduce a transfer learning framework that pre-trains and fine-tunes the encoder using DINO and SC, respectively, and a multitask learning framework that shares the encoder while having separate projection layers for both methods. To achieve lower inter-speaker similarity, we propose a joint learning framework sharing both the encoder and projection layer for DINO and SC. We also propose an auxiliary contrastive loss between embeddings derived from labeled and unlabeled utterances and introduce a two-stage learning strategy to apply margin penalty effectively. Experimental results on the VoxCeleb corpus indicate that the joint learning framework outperforms the other frameworks and is closest to achieving the performance of fully supervised learning. Jeong-Hwan Choi, Jehyun Kyung, Ju-Seok Seong, Ye-Rin Jeoung, Joon-Hyuk Chang |
ASRU | 1 |
| 2023 | Improving Transformer-Based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention HeadsabstractTransformer-based end-to-end neural speaker diarization (EEND) models utilize the multi-head self-attention (SA) mechanism to enable accurate speaker label prediction in overlapped speech regions. In this study, to enhance the training effectiveness of SA-EEND models, we propose the use of auxiliary losses for the SA heads of the transformer layers. Specifically, we assume that the attention weight matrices of an SA layer are redundant if their patterns are similar to those of the identity matrix. We then explicitly constrain such matrices to exhibit specific speaker activity patterns relevant to voice activity detection or overlapped speech detection tasks. Consequently, we expect the proposed auxiliary losses to guide the transformer layers to exhibit more diverse patterns in the attention weights, thereby reducing the assumed redundancies in the SA heads. The effectiveness of the proposed method is demonstrated using the simulated and CALLHOME datasets for two-speaker diarization tasks, reducing the diarization error rate of the conventional SA-EEND model by 32.58% and 17.11%, respectively. Ye-Rin Jeoung, Joon-Young Yang, Jeong-Hwan Choi, Joon-Hyuk Chang |
ICASSP | 3 |
| 2023 | Noise-Aware Target Extension with Self-Distillation for Robust Speech RecognitionabstractData augmentation using additive noise is a framework for robustly training automatic speech recognition models. To utilize noise information efficiently, previous studies used an additional branch to classify noise conditions. This added branch has a limited effect on the ASR because it performs independently of the ASR branch that classifies senones. In this paper, we propose a noise-aware target extension (NATE) that extends the senone target to contain noise awareness by jointly classifying the senone and noise in a single branch. In the inference stage, the output of the model is processed separately by the noise condition and then aggregated to match the senone posterior distribution. In addition, we combine NATE with self-distillation (NATEsd) to reduce the model parameters and avoid discrepancies between the outputs of training and inference. The effectiveness of the NATE method is validated on the two benchmark development and evaluation sets and simulated noisy test sets, resulting in significant improvements over the previous methods. Ju-Seok Seong, Jeong-Hwan Choi, Jehyun Kyung, Ye-Rin Jeoung, Joon-Hyuk Chang |
ICASSP | 2 |
| 2023 | Self-Distillation into Self-Attention Heads for Improving Transformer-based End-to-End Neural Speaker Diarization
Ye-Rin Jeoung, Jeong-Hwan Choi, Ju-Seok Seong, Jehyun Kyung, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2023 | Improving Joint Speech and Emotion Recognition Using Global Style Tokens
Jehyun Kyung, Ju-Seok Seong, Jeong-Hwan Choi, Ye-Rin Jeoung, Joon-Hyuk Chang |
INTERSPEECH | 3 |
| 2023 | HAD-ANC: A Hybrid System Comprising an Adaptive Filter and Deep Neural Networks for Active Noise Control
JungPhil Park, Jeong-Hwan Choi, Yungyeo Kim, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2022 | Improved CNN-Transformer using Broadcasted Residual Learning for Text-Independent Speaker Verification
Jeong-Hwan Choi, Joon-Young Yang, Ye-Rin Jeoung, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2022 | HYU Submission for the SASV Challenge 2022: Reforming Speaker Embeddings with Spoofing-Aware Conditioning
Jeong-Hwan Choi, Joon-Young Yang, Ye-Rin Jeoung, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2022 | Supervised Learning Approach for Explicit Spatial Filtering of SpeechabstractSpatial filtering of speech based on neural networks has been widely studied. However, existing approaches focus on improving signal extraction or separation performance, and how to define the signal in the direction-of-interest (DOI) for spatial filtering has not been investigated in detail. This study proposes a method to train neural networks for extracting directivity components of speech signals in the DOI. To this end, we formulate the problem by defining the DOI and its corresponding desired signal in a reverberant environment. Moreover, we demonstrate an on-the-fly training data generation procedure to feed the spatially diverse data to train the networks. The proposed method was evaluated with regard to spatial speech extraction and localization performance. In particular, it has been confirmed that the network trained with the proposed method using simulated datasets also functions for real recordings. Jeong-Hwan Choi, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 1 |
| 2021 | Short-Utterance Embedding Enhancement Method Based on Time Series Forecasting Technique for Text-Independent Speaker VerificationabstractShort-utterance embedding, which is a speaker embedding extracted from a short utterance, shows poor speaker verification performance due to insufficient speaker information. To address the problem, we propose a method to map the set of short-utterance embeddings to a set of long-utterance embeddings based on a neural network. Specifically, a speech utterance is cropped into multiple segments whose durations are gradually increasing, and the speaker embeddings are extracted from the sequence of cropped segments using a pre-trained speaker embedding extractor. Subsequently, the sequence of embeddings is divided into a group of short-utterances embeddings and that of long-utterance embeddings. In our method, a sequence-to-sequence model based forecasting technique is employed, where an encoder transforms the group of short-utterance embeddings to a fixed-dimensional vector, and then a decoder converts the vector into a group of long-utterance embeddings. Experimental results on the VoxCeleb and Speakers in the Wild datasets show that our method improves the text-independent speaker verification performance under short utterance condition. Jeong-Hwan Choi, Joon-Young Yang, Joon-Hyuk Chang |
ASRU | 1 |
| 2021 | MIMO Noise Suppression Preserving Spatial Cues for Sound Source Localization in Mobile RobotabstractIn this paper, a multi-input multi-output (MIMO) noise suppression (NS) algorithm for sound source localization (SSL) is proposed in mobile robot environment. Especially for the mobile robot, some ego noise (e.g., motor noise) is naturally generated and it becomes a dominant signal at microphone array, which seriously damages the performance of the SSL. Therefore, the proposed MIMO NS algorithm is designed not only to suppress the noise but also to preserve spatial information of sound sources. To this end, the recent MIMO time-domain audio separation network (TasNet) with preserving spatial cues for binaural source separation is extended to tetrahedral microphone array in the proposed MIMO NS algorithm. Furthermore, weighted phase error (WPE) between inter-channels is additionally employed in the loss function, which can improve the performance especially in preserving inter-channel time difference (ITD). The performance of the proposed approach is verified within the context of the robot vacuum cleaner, which shows from simulations and real experiments that the proposed approach can preserve the spatial information as well as remove the vacuum noise. Jeong-Hwan Choi, Jinyoung Son, Gyeong-Su Kim, Joon-Hyuk Chang |
ISCAS | 2 |