VLDB 2026 Research / reviewers in the wild / expert
Hyunsin Park
dblp:50/9205
· DBLP profile ↗
16ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-3556-5792ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech RepresentationsabstractRecent advances in zero-shot text-to-speech (TTS) have enabled accurate imitation of reference speech in terms of both speaking style and speaker timbre.However, achieving disentangled control over these aspects from separate references remains a challenging task.Several studies have proposed disentangled speech representations that decompose speech into interpretable attributes (e.g., timbre, prosody, and content), providing a promising foundation for TTS with attribute control from separate references.Yet, how to effectively integrate such representations into TTS systems to achieve independent and precise control remains underexplored.In this paper, we present FC-TTS, a zero-shot TTS framework that enables disentangled control of style and timbre by conditioning on two distinct reference utterances.Unlike existing systems that inherit limitations from those pre-trained disentangled representations, FC-TTS introduces key design strategies, including architectural choices, training framework, and auxiliary training objectives, which improve the reliability of attribute separation and dual-reference control.Experiments show that FC-TTS achieves high-fidelity synthesis and competitive zero-shot naturalness, while uniquely supporting consistent and independent manipulation of style and timbre.Audio samples are available at https Yoonhyung Lee, Hyunsin Park, Jinhwan Park |
ACL (1) | 2 |
| 2024 | FedHide: Federated Learning by Hiding in the Neighbors
Hyunsin Park, Sungrack Yun |
ECCV (70) | 1 |
| 2024 | Feature Diversification and Adaptation for Federated Domain Generalization
Seunghan Yang, Seokeon Choi, Hyunsin Park, Sungha Choi, Simyung Chang, Sungrack Yun |
ECCV (72) | 3 |
| 2024 | Balanced Learning for Multi-Domain Long-Tailed Speaker RecognitionabstractThis paper considers two types of imbalance problems commonly inherent in large-scale datasets: multiple domain and class imbalance. Class imbalance causes the algorithm to be biased toward the majority classes, and multiple-domain data results in significant performance disparities for different domains. To tackle these challenges, we propose a novel learning approach for multi-domain imbalanced datasets, featuring two techniques: (i) distribution-aware partial mask and (ii) domain-wise interprototype loss function. The distribution-aware partial mask selects negative class centers based on class-level distribution and domain labels, adjusting the ratio of positive and negative updates for prototype vectors and enhancing discriminative feature learning within each domain. Additionally, the domain-wise interprototype loss enforces orthogonality among prototype vectors within each domain, leading to increased discriminativeness. We demonstrate the superiority of our approach over baselines through experiments on publicly available speaker recognition datasets, including CN-Celeb and Mozilla Common Voice. Janghoon Cho, Hyunsin Park, Hyoungwoo Park, Seunghan Yang, Sungrack Yun |
ICASSP | 3 |
| 2023 | Progressive Random Convolutions for Single Domain GeneralizationabstractSingle domain generalization aims to train a generalizable model with only one source domain to perform well on arbitrary unseen target domains. Image augmentation based on Random Convolutions (RandConv), consisting of one convolution layer randomly initialized for each mini-batch, enables the model to learn generalizable visual representations by distorting local textures despite its simple and lightweight structure. However, RandConv has structural limitations in that the generated image easily loses semantics as the kernel size increases, and lacks the inherent diversity of a single convolution operation. To solve the problem, we propose a Progressive Random Convolution (Pro-RandConv) method that recursively stacks random convolution layers with a small kernel size instead of increasing the kernel size. This progressive approach can not only mitigate semantic distortions by reducing the influence of pixels away from the center in the theoretical receptive field, but also create more effective virtual domains by gradually increasing the style diversity. In addition, we develop a basic random convolution layer into a random convolution block including deformable offsets and affine transformation to support texture and contrast diversification, both of which are also randomly initialized. Without complex generators or adversarial learning, we demonstrate that our simple yet effective augmentation strategy outperforms state-of-the-art methods on single domain generalization benchmarks. Seokeon Choi, Debasmit Das, Sungha Choi, Seunghan Yang, Hyunsin Park, Sungrack Yun |
CVPR | 5 |
| 2022 | Multi-Head Modularization to Leverage Generalization Capability in Multi-Modal NetworksabstractIt has been crucial to leverage the rich information of multiple modalities in many tasks. Existing works have tried to design multi-modal networks with descent multi-modal fusion modules. Instead, we focus on improving generalization capability of multi-modal networks, especially the fusion module. Viewing the multi-modal data as different projections of information, we first observe that bad projection can cause poor generalization behaviors of multi-modal networks. Then, motivated by well-generalized network's low sensitivity to perturbation, we propose a novel multi-modal training method, multi-head modularization (MHM). We modularize a multi-modal network as a series of uni-modal embedding, multi-modal embedding, and task-specific head modules. Also, for training, we exploit multiple head modules learned with different datasets, swapping each other. From this, we can make the multi-modal embedding module robust to all the heads with different generalization behaviors. In testing phase, we select one of the head modules not to increase the computational cost. Owing to the perturbation of head modules, though including one selected head, the deployed network is more well-generalized compared to the simply end-to-end learned. We verify the effectiveness of MHM on various multi-modal tasks. We use the state-of-the-art methods as baselines, and show notable performance gain for all the baselines. Juntae Lee, Hyunsin Park, Sungrack Yun, Simyung Chang |
AAAI | 2 |
| 2022 | Domain Generalization with Relaxed Instance Frequency-wise Normalization for Multi-device Acoustic Scene ClassificationabstractWhile using two-dimensional convolutional neural networks (2D-CNNs) in image processing, it is possible to manipulate domain information using channel statistics, and instance normalization has been a promising way to get domain-invariant features. However, unlike image processing, we analyze that domain-relevant information in an audio feature is dominant in frequency statistics rather than channel statistics. Motivated by our analysis, we introduce Relaxed Instance Frequency-wise Normalization (RFN): a plug-and-play, explicit normalization module along the frequency axis which can eliminate instance-specific domain discrepancy in an audio feature while relaxing undesirable loss of useful discriminative information. Empirically, simply adding RFN to networks shows clear margins compared to previous domain generalization approaches on acoustic scene classification and yields improved robustness for multiple audio devices. Especially, the proposed RFN won the DCASE2021 challenge TASK1A, low-complexity acoustic scene classification with multiple devices, with a clear margin, and RFN is an extended work of our technical report. Byeonggeun Kim, Seunghan Yang, Jangho Kim, Hyunsin Park, Juntae Lee, Simyung Chang |
INTERSPEECH | 4 |
| 2021 | Subspectral Normalization for Neural Audio Data ProcessingabstractConvolutional Neural Networks are widely used in various machine learning domains. In image processing, the features can be obtained by applying 2D convolution to all spatial dimensions of the input. However, in the audio case, frequency domain input like Mel-Spectrogram has different and unique characteristics in the frequency dimension. Thus, there is a need for a method that allows the 2D convolution layer to handle the frequency dimension differently. In this work, we introduce SubSpectral Normalization (SSN), which splits the input frequency dimension into several groups (sub-bands) and performs a different normalization for each group. SSN also includes an affine transformation that can be applied to each group. Our method removes the inter-frequency deflection while the network learns a frequency-aware characteristic. In the experiments with audio data, we observed that SSN can efficiently improve the network’s performance. Simyung Chang, Hyoungwoo Park, Janghoon Cho, Hyunsin Park, Sungrack Yun, Kyuwoong Hwang |
ICASSP | 4 |
| 2021 | Federated Learning of User Verification Models Without Sharing EmbeddingsabstractWe consider the problem of training User Verification (UV) models in federated setup, where each user has access to the data of only one class and user embeddings cannot be shared with the server or other users. To address this problem, we propose Federated User Verification (FedUV), a framework in which users jointly learn a set of vectors and maximize the correlation of their instance embeddings with a secret linear combination of those vectors. We show that choosing the linear combinations from the codewords of an error-correcting code allows users to collaboratively train the model without revealing their embedding vectors. We present the experimental results for user verification with voice, face, and handwriting data and show that FedUV is on par with existing approaches, while not sharing the embeddings with other users or the server. Hossein Hosseini, Hyunsin Park, Sungrack Yun, Christos Louizos, Joseph B. Soriaga, Max Welling |
ICML | 2 |
| 2020 | Learning Augmentation Network via Influence FunctionsabstractData augmentation can impact the generalization performance of an image classification model in a significant way. However, it is currently conducted on the basis of trial and error, and its impact on the generalization performance cannot be predicted during training. This paper considers an influence function that predicts how generalization performance, in terms of validation loss, is affected by a particular augmented training sample. The influence function provides an approximation of the change in validation loss without actually comparing the performances that include and exclude the sample in the training process. Based on this function, a differentiable augmentation network is learned to augment an input training sample to reduce validation loss. The augmented sample is fed into the classification network, and its influence is approximated as a function of the parameters of the last fully-connected layer of the classification network. By backpropagating the influence to the augmentation network, the augmentation network parameters are learned. Experimental results on CIFAR-10, CIFAR-100, and ImageNet show that the proposed method provides better generalization performance than conventional data augmentation methods do. Hyunsin Park, Trung X. Pham, Chang Dong Yoo |
CVPR | 2 |
| 2020 | CNN-Based Learnable Gammatone Filterbank and Equal-Loudness Normalization for Environmental Sound ClassificationabstractFor environmental sound classification (ESC), this letter presents a learnable auditory filterbank based on a one-dimensional (1D) convolutional neural network with strong psychophysiological inductive bias in the form of a gammatone filterbank and an equal-loudness prompting normalization. In the past, a number of ESC methods based on learnable auditory features obtained by performing plain 1D convolutions on raw input waveforms for outperforming traditional handcrafted features such as a mel-frequency filterbank have been proposed. However, the large number of parameters involved in the convolutions suggests that these methods will not generalize better than a model defined by a smaller number of parameters, which is considered in this letter. Here, a learnable gammatone filterbank layer consisting of 1D kernels represented by a parametric form of the bandpass gammatone filters is proposed for acquiring a time-frequency representation of the raw waveform. A normalization with learnable parameters that control the trade-off between energy equalization and structure preservation in the spectro-temporal domain is proposed. To verify the effectiveness of the considered network and the normalization, ESC experiments on the ESC-50 and UrbanSound8K datasets were conducted. Compared to other state-of-the-art networks, the considered network performed better on the two datasets. In addition, an ensemble architecture achieved further performance improvement. Hyunsin Park, Chang Dong Yoo |
IEEE Signal Process. Lett. | 1 |
| 2017 | Melody extraction and detection through LSTM-RNN with harmonic sum lossabstractThis paper proposes a long short-term memory recurrent neural network (LSTM-RNN) for extracting melody and simultaneously detecting regions of melody from polyphonic audio using the proposed harmonic sum loss. The previous state-of-the-art algorithms have not been based on machine learning techniques and certainly not on deep architectures. The harmonics structure in melody is incorporated in the loss function to attain robustness against both octave mismatch and interference from background music. Experimental results show that the performance of the proposed method is better than or comparable to other state-of-the-art algorithms. Hyunsin Park, Chang Dong Yoo |
ICASSP | 1 |
| 2017 | Complex Video Scene Analysis Using Kernelized-Collaborative Behavior Pattern Learning Based on Hierarchical Representative Object BehaviorsabstractThis paper considers an unsupervised learning algorithm that can automatically discover key behavior patterns to characterize a complex video scene. For behavior features (bFs) extracted at multiple spatial-temporal scales, an optimization problem is formulated to cluster bFs in their scales while establishing a collaborative nonlinear relationship in the form of a kernel regression function among clustered bFs across different spatial scales. The relationship allows features extracted in one scale to be considered as contextual information in the analysis of another scale. This optimization problem is solved using linear programming to reduce computational complexity. The proposed algorithm is evaluated on four crowded traffic scenes and two sports video data sets. Experimental results show that the proposed algorithm achieves a better performance compared with the current state-of-the-art algorithms in terms of video segmentation accuracy. Sanghyuk Park, Hyunsin Park, Chang Dong Yoo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Face alignment using cascade Gaussian process regression treesabstractIn this paper, we propose a face alignment method that uses cascade Gaussian process regression trees (cGPRT) constructed by combining Gaussian process regression trees (GPRT) in a cascade stage-wise manner. Here, GPRT is a Gaussian process with a kernel defined by a set of trees. The kernel measures the similarity between two inputs as the number of trees where the two inputs fall in the same leaves. Without increasing prediction time, the prediction of cGPRT can be performed in the same framework as the cascade regression trees (CRT) but with better generalization. Features for GPRT are designed using shape-indexed difference of Gaussian (DoG) filter responses sampled from local retinal patterns to increase stability and to attain robustness against geometric variances. Compared with the previous CRT-based face alignment methods that have shown state-of-the-art performances, cGPRT using shape-indexed DoG features performed best on the HELEN and 300-W datasets which are the most challenging dataset today. Hyunsin Park, Chang Dong Yoo |
CVPR | 2 |
| 2012 | Sparsity Sharing Embedding for Face Verification
Hyunsin Park, Junyoung Chung, Youngook Song, Chang Dong Yoo |
ACCV (2) | 2 |
| 2012 | Phoneme Classification using Constrained Variational Gaussian Process Dynamical SystemabstractThis paper describes a new acoustic model based on variational Gaussian process dynamical system (VGPDS) for phoneme classification. The proposed model overcomes the limitations of the classical HMM in modeling the real speech data, by adopting a nonlinear and nonparametric model. In our model, the GP prior on the dynamics function enables representing the complex dynamic structure of speech, while the GP prior on the emission function successfully models the global dependency over the observations. Additionally, we introduce variance constraint to the original VGPDS for mitigating sparse approximation error of the kernel matrix. The effectiveness of the proposed model is demonstrated with extensive experimental results including parameter estimation, classification performance on the synthetic and benchmark datasets. Hyunsin Park, Sungrack Yun, Sanghyuk Park, Jongmin Kim 0006, Chang Dong Yoo |
NIPS | 1 |