Sungrack Yun

dblp:67/8053 · DBLP profile ↗
← Back
34ranked-venue papers
4as first author
22since 2021 · last 2025
0000-0003-2462-3854ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 24 · 3 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning
abstract
Few-shot class incremental learning (FSCIL) enables the continual learning of new concepts with only a few training examples. In FSCIL, the model undergoes substantial updates, making it prone to forgetting previous concepts and overfitting to the limited new examples. Most recent trend is typically to disentangle the learning of the representation from the classification head of the model. A well-generalized feature extractor on the base classes (many examples and many classes) is learned, and then fixed during incremental learning. Arguing that the fixed feature extractor restricts the model’s adaptability to new classes, we introduce a novel FSCIL method to effectively address catastrophic forgetting and overfitting issues. Our method enables to seamlessly update the entire model with a few examples. We mainly propose a tripartite weight-space ensemble (Tri-WE). Tri-WE interpolates the base, immediately previous, and current models in weight-space, especially for the classification heads of the models. Then, it collaboratively maintains knowledge from the base and previous models. In addition, we recognize the challenges of distilling generalized representations from the previous model from scarce data. Hence, we suggest a regularization loss term using amplified data knowledge distillation. Simply intermixing the few-shot data, we can produce richer data enabling the distillation of critical knowledge from the previous model. Consequently, we attain state-of-the-art results on the miniImageNet, CUB200, and CIFAR100 datasets.
Juntae Lee, Munawar Hayat, Sungrack Yun
CVPR3
2025 ConsNoTrainLoRA: Data-driven Weight Initialization of Low-Rank Adapters Using Constraints
Debasmit Das, Hyoungwoo Park, Munawar Hayat, Seokeon Choi, Sungrack Yun, Fatih Porikli
ICCV5
2025 Steering Guidance for Personalized Text-to-Image Diffusion Models
abstract
Personalizing text-to-image diffusion models is crucial for adapting the pre-trained models to specific target concepts, enabling diverse image generation. However, fine-tuning with few images introduces an inherent trade-off between aligning with the target distribution (e.g., subject fidelity) and preserving the broad knowledge of the original model (e.g., text editability). Existing sampling guidance methods, such as classifier-free guidance (CFG) and autoguidance (AG), fail to effectively guide the output toward well-balanced space: CFG restricts the adaptation to the target distribution, while AG compromises text alignment. To address these limitations, we propose personalization guidance, a simple yet effective method leveraging an unlearned weak model conditioned on a null text prompt. Moreover, our method dynamically controls the extent of unlearning in a weak model through weight interpolation between pre-trained and fine-tuned models during inference. Unlike existing guidance methods, which depend solely on guidance scales, our method explicitly steers the outputs toward a balanced latent space without additional computational overhead. Experimental results demonstrate that our proposed guidance can improve text alignment and target distribution fidelity, integrating seamlessly with various fine-tuning strategies.
Seokeon Choi, Hyoungwoo Park, Sungrack Yun
ICCV4
2025 MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
abstract
Generation of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to this is the lack of a a dedicated benchmark. To address this, we introduce MultiHuman-Testbench, a novel benchmark for rigorously evaluating generative models for multi-human generation. The benchmark comprises 1800 samples, including carefully curated text prompts, describing a range of simple to complex human actions. These prompts are matched with a total of 5,550 unique human face images, sampled uniformly to ensure diversity across age, ethnic background, and gender. Alongside captions, we provide human-selected pose conditioning images which accurately match the prompt. We propose a multi-faceted evaluation suite employing four key metrics to quantify face count, ID similarity, prompt alignment, and action detection. We conduct a thorough evaluation of a diverse set of models, including zero-shot approaches and training-based methods, with and without regional priors. We also propose novel techniques to incorporate image and region isolation using human segmentation and Hungarian matching, significantly improving ID similarity. Our proposed benchmark and key findings provide valuable insights and a standardized tool for advancing research in multi-human image generation.
Shubhankar Borse, Seokeon Choi, Sunghyun Park 0005, Jeongho Kim 0007, Shreya Kadambi, Risheek Garrepalli, Sungrack Yun, Durga Malladi, Fatih Porikli
NeurIPS7
2024 FedHide: Federated Learning by Hiding in the Neighbors
Hyunsin Park, Sungrack Yun
ECCV (70)2
2024 Feature Diversification and Adaptation for Federated Domain Generalization
Seunghan Yang, Seokeon Choi, Hyunsin Park, Sungha Choi, Simyung Chang, Sungrack Yun
ECCV (72)6
2024 Balanced Learning for Multi-Domain Long-Tailed Speaker Recognition
abstract
This paper considers two types of imbalance problems commonly inherent in large-scale datasets: multiple domain and class imbalance. Class imbalance causes the algorithm to be biased toward the majority classes, and multiple-domain data results in significant performance disparities for different domains. To tackle these challenges, we propose a novel learning approach for multi-domain imbalanced datasets, featuring two techniques: (i) distribution-aware partial mask and (ii) domain-wise interprototype loss function. The distribution-aware partial mask selects negative class centers based on class-level distribution and domain labels, adjusting the ratio of positive and negative updates for prototype vectors and enhancing discriminative feature learning within each domain. Additionally, the domain-wise interprototype loss enforces orthogonality among prototype vectors within each domain, leading to increased discriminativeness. We demonstrate the superiority of our approach over baselines through experiments on publicly available speaker recognition datasets, including CN-Celeb and Mozilla Common Voice.
Janghoon Cho, Hyunsin Park, Hyoungwoo Park, Seunghan Yang, Sungrack Yun
ICASSP6
2024 Hollowed Net for On-Device Personalization of Text-to-Image Diffusion Models
abstract
Recent advancements in text-to-image diffusion models have enabled the personalization of these models to generate custom images from textual prompts. This paper presents an efficient LoRA-based personalization approach for on-device subject-driven generation, where pre-trained diffusion models are fine-tuned with user-specific data on resource-constrained devices. Our method, termed Hollowed Net, enhances memory efficiency during fine-tuning by modifying the architecture of a diffusion U-Net to temporarily remove a fraction of its deep layers, creating a hollowed structure. This approach directly addresses on-device memory constraints and substantially reduces GPU memory requirements for training, in contrast to previous methods that primarily focus on minimizing training steps and reducing the number of parameters to update. Additionally, the personalized Hollowed Net can be transferred back into the original U-Net, enabling inference without additional memory overhead. Quantitative and qualitative analyses demonstrate that our approach not only reduces training memory to levels as low as those required for inference but also maintains or improves personalization performance compared to existing methods.
Wonguk Cho, Seokeon Choi, Debasmit Das, Matthias Reisser, Taesup Kim, Sungrack Yun, Fatih Porikli
NeurIPS6
2023 Progressive Random Convolutions for Single Domain Generalization
abstract
Single domain generalization aims to train a generalizable model with only one source domain to perform well on arbitrary unseen target domains. Image augmentation based on Random Convolutions (RandConv), consisting of one convolution layer randomly initialized for each mini-batch, enables the model to learn generalizable visual representations by distorting local textures despite its simple and lightweight structure. However, RandConv has structural limitations in that the generated image easily loses semantics as the kernel size increases, and lacks the inherent diversity of a single convolution operation. To solve the problem, we propose a Progressive Random Convolution (Pro-RandConv) method that recursively stacks random convolution layers with a small kernel size instead of increasing the kernel size. This progressive approach can not only mitigate semantic distortions by reducing the influence of pixels away from the center in the theoretical receptive field, but also create more effective virtual domains by gradually increasing the style diversity. In addition, we develop a basic random convolution layer into a random convolution block including deformable offsets and affine transformation to support texture and contrast diversification, both of which are also randomly initialized. Without complex generators or adversarial learning, we demonstrate that our simple yet effective augmentation strategy outperforms state-of-the-art methods on single domain generalization benchmarks.
Seokeon Choi, Debasmit Das, Sungha Choi, Seunghan Yang, Hyunsin Park, Sungrack Yun
CVPR6
2023 Few-Shot Common Action Localization via Cross-Attentional Fusion of Context and Temporal Dynamics
abstract
The goal of this paper is to localize action instances in a long untrimmed query video using just meager trimmed support videos representing a common action whose class information is not given. In this task, it is crucial to mine reliable temporal cues representing a common action from handful support videos. In our work, we develop an attention mechanism using cross-correlation. Based on this cross-attention, we first transform the support videos into query video’s context to emphasize query-relevant important frames, and suppress less relevant ones. Next, we summarize sub-sequences of support video frames to represent temporal dynamics in coarse temporal granularity, which is then propagated to the fine-grained support video features through the cross-attention. In each case, the cross-attentions are applied to each support video in the individual-to-all strategy to balance heterogeneity and compatibility of the support videos. In contrast, the candidate instances in the query video are lastly attended by the resulting support video features, at once. In addition, we also develop a relational classifier head based on the query and support video representations. We show the effectiveness of our work with the state-of-the-art (SOTA) performance in benchmark datasets (ActivityNet1.3 and THUMOS14), and analyze each component extensively.
Juntae Lee, Mihir Jain, Sungrack Yun
ICCV3
2023 Label Shift Adapter for Test-Time Adaptation under Covariate and Label Shifts
abstract
Test-time adaptation (TTA) aims to adapt a pre-trained model to the target domain in a batch-by-batch manner during inference. While label distributions often exhibit imbalances in real-world scenarios, most previous TTA approaches typically assume that both source and target domain datasets have balanced label distribution. Due to the fact that certain classes appear more frequently in certain domains (e.g., buildings in cities, trees in forests), it is natural that the label distribution shifts as the domain changes. However, we discover that the majority of existing TTA methods fail to address the coexistence of covariate and label shifts. To tackle this challenge, we propose a novel label shift adapter that can be incorporated into existing TTA approaches to deal with label shifts during the TTA process effectively. Specifically, we estimate the label distribution of the target domain to feed it into the label shift adapter. Subsequently, the label shift adapter produces optimal parameters for target label distribution. By predicting only the parameters for a part of the pre-trained source model, our approach is computationally efficient and can be easily applied, regardless of the model architectures. Through extensive experiments, we demonstrate that integrating our strategy with TTA approaches leads to substantial performance improvements under the joint presence of label and covariate shifts.
Sunghyun Park 0005, Seunghan Yang, Jaegul Choo, Sungrack Yun
ICCV4
2023 Multi-Scale Temporal Feature Fusion for Few-Shot Action Recognition
abstract
The aim of this paper is to recognize actions of interest that are given by a few support videos in testing (query) videos. The focus of our approach is to develop a novel temporal enrichment module where the features describing local temporal contexts in videos are enhanced by collaboratively merging important information in frame-level (no temporal context) features. We call this module a multi-scale temporal feature fusion (MSTFF) module. Utilizing multiple MSTFF modules varying the scope of local temporal context extraction, we can obtain discriminative video representation which is crucial in the few-shot tasks where support videos are not sufficient to describe an action class. For stable learning of a model with MSTFF and the performance boost, we also learn a local temporal context-level auxiliary classifier in parallel with the main classifier. We analyze the proposed components to demonstrate their importance. We achieve state-of-the-art on three few-shot action recognition benchmarks: Something-Something V2 (SSv2), HMDB51, and Kinetics.
Juntae Lee, Sungrack Yun
ICIP2
2022 Multi-Head Modularization to Leverage Generalization Capability in Multi-Modal Networks
abstract
It has been crucial to leverage the rich information of multiple modalities in many tasks. Existing works have tried to design multi-modal networks with descent multi-modal fusion modules. Instead, we focus on improving generalization capability of multi-modal networks, especially the fusion module. Viewing the multi-modal data as different projections of information, we first observe that bad projection can cause poor generalization behaviors of multi-modal networks. Then, motivated by well-generalized network's low sensitivity to perturbation, we propose a novel multi-modal training method, multi-head modularization (MHM). We modularize a multi-modal network as a series of uni-modal embedding, multi-modal embedding, and task-specific head modules. Also, for training, we exploit multiple head modules learned with different datasets, swapping each other. From this, we can make the multi-modal embedding module robust to all the heads with different generalization behaviors. In testing phase, we select one of the head modules not to increase the computational cost. Owing to the perturbation of head modules, though including one selected head, the deployed network is more well-generalized compared to the simply end-to-end learned. We verify the effectiveness of MHM on various multi-modal tasks. We use the state-of-the-art methods as baselines, and show notable performance gain for all the baselines.
Juntae Lee, Hyunsin Park, Sungrack Yun, Simyung Chang
AAAI3
2022 Improving Test-Time Adaptation Via Shift-Agnostic Weight Regularization and Nearest Source Prototypes
Sungha Choi, Seunghan Yang, Seokeon Choi, Sungrack Yun
ECCV (33)4
2022 ConFeSS: A Framework for Single Source Cross-Domain Few-Shot Learning
Debasmit Das, Sungrack Yun, Fatih Porikli
ICLR2
2022 Domain Agnostic Few-shot Learning for Speaker Verification
abstract
Deep learning models for verification systems often fail to generalize to new users and new environments, even though they learn highly discriminative features.To address this problem, we propose a few-shot domain generalization framework that learns to tackle distribution shift for new users and new domains.Our framework consists of domain-specific and domainaggregation networks, which are the experts on specific and combined domains, respectively.By using these networks, we generate episodes that mimic the presence of both novel users and novel domains in the training phase to eventually produce better generalization.To save memory, we reduce the number of domain-specific networks by clustering similar domains together.Upon extensive evaluation on artificially generated noise domains, we can explicitly show generalization ability of our framework.In addition, we apply our proposed methods to the existing competitive architecture on the standard benchmark, which shows further performance improvements.
Seunghan Yang, Debasmit Das, Janghoon Cho, Hyoungwoo Park, Sungrack Yun
INTERSPEECH5
2022 Leaky Gated Cross-Attention for Weakly Supervised Multi-Modal Temporal Action Localization
abstract
As multiple modalities sometimes have a weak complementary relationship, multi-modal fusion is not always beneficial for weakly supervised action localization. Hence, to attain the adaptive multi-modal fusion, we propose a leaky gated cross-attention mechanism. In our work, we take the multi-stage cross-attention as the baseline fusion module to obtain multi-modal features. Then, for the stages of each modality, we design gates to decide the dependency on the other modality. For each input frame, if two modalities have a strong complementary relationship, the gate selects the cross-attended feature, otherwise the non-attended feature. Also, the proposed gate allows the non-selected feature to escape through it with a small intensity, we call it leaky gate. This leaky feature makes effective regularization of the selected major feature. Therefore, our leaky gating makes cross-attention more adaptable and robust even when the modalities have a weak complementary relationship. The proposed leaky gated cross-attention provides a modality fusion module that is generally compatible with various temporal action localization methods. To show its effectiveness, we do extensive experimental analysis and apply the proposed method to boost the performance of the state-of-the-art methods on two benchmark datasets (ActivityNet1.2 and THUMOS14).
Juntae Lee, Sungrack Yun, Mihir Jain
WACV2
2021 Subspectral Normalization for Neural Audio Data Processing
abstract
Convolutional Neural Networks are widely used in various machine learning domains. In image processing, the features can be obtained by applying 2D convolution to all spatial dimensions of the input. However, in the audio case, frequency domain input like Mel-Spectrogram has different and unique characteristics in the frequency dimension. Thus, there is a need for a method that allows the 2D convolution layer to handle the frequency dimension differently. In this work, we introduce SubSpectral Normalization (SSN), which splits the input frequency dimension into several groups (sub-bands) and performs a different normalization for each group. SSN also includes an affine transformation that can be applied to each group. Our method removes the inter-frequency deflection while the network learns a frequency-aware characteristic. In the experiments with audio data, we observed that SSN can efficiently improve the network’s performance.
Simyung Chang, Hyoungwoo Park, Janghoon Cho, Hyunsin Park, Sungrack Yun, Kyuwoong Hwang
ICASSP5
2021 Prototype-Based Personalized Pruning
abstract
Nowadays, as edge devices such as smartphones become prevalent, there are increasing demands for personalized services. However, traditional personalization methods are not suitable for edge devices because retraining or finetuning is needed with limited personal data. Also, a full model might be too heavy for edge devices with limited resources. Unfortunately, model compression methods which can handle the model complexity issue also require the retraining phase. These multiple training phases generally need huge computational cost during on-device learning which can be a burden to edge devices. In this work, we propose a dynamic personalization method called prototype-based personalized pruning (PPP). PPP considers both ends of personalization and model efficiency. After training a network, PPP can easily prune the network with a prototype representing the characteristics of personal data and it performs well without retraining or finetuning. We verify the usefulness of PPP on a couple of tasks in computer vision and Keyword spotting.
Jangho Kim, Simyung Chang, Sungrack Yun, Nojun Kwak
ICASSP3
2021 Efficient Action Recognition via Dynamic Knowledge Propagation
abstract
Efficient action recognition has become crucial to extend the success of action recognition to many real-world applications. Contrary to most existing methods, which mainly focus on selecting salient frames to reduce the computation cost, we focus more on making the most of the selected frames. To this end, we employ two networks of different capabilities that operate in tandem to efficiently recognize actions. Given a video, the lighter network processes more frames while the heavier one only processes a few. In order to enable the effective interaction between the two, we propose dynamic knowledge propagation based on a cross-attention mechanism. This is the main component of our framework that is essentially a student-teacher architecture, but as the teacher model continues to interact with the student model during inference, we call it a dynamic student-teacher framework. Through extensive experiments, we demonstrate the effectiveness of each component of our framework. Our method outperforms competing state-of-the-art methods on two video datasets: ActivityNet-v1.3 and Mini-Kinetics.
Hanul Kim 0001, Mihir Jain, Juntae Lee, Sungrack Yun, Fatih Porikli
ICCV4
2021 Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action Localization
Juntae Lee, Mihir Jain, Hyoungwoo Park, Sungrack Yun
ICLR4
2021 Federated Learning of User Verification Models Without Sharing Embeddings
abstract
We consider the problem of training User Verification (UV) models in federated setup, where each user has access to the data of only one class and user embeddings cannot be shared with the server or other users. To address this problem, we propose Federated User Verification (FedUV), a framework in which users jointly learn a set of vectors and maximize the correlation of their instance embeddings with a secret linear combination of those vectors. We show that choosing the linear combinations from the codewords of an error-correcting code allows users to collaboratively train the model without revealing their embedding vectors. We present the experimental results for user verification with voice, face, and handwriting data and show that FedUV is on par with existing approaches, while not sharing the embeddings with other users or the server.
Hossein Hosseini, Hyunsin Park, Sungrack Yun, Christos Louizos, Joseph B. Soriaga, Max Welling
ICML3
2019 An End-to-End Text-Independent Speaker Verification Framework with a Keyword Adversarial Network
abstract
This paper presents an end-to-end text-independent speaker verification framework by jointly considering the speaker embedding (SE) network and automatic speech recognition (ASR) network.The SE network learns to output an embedding vector which distinguishes the speaker characteristics of the input utterance, while the ASR network learns to recognize the phonetic context of the input.In training our speaker verification framework, we consider both the triplet loss minimization and adversarial gradient of the ASR network to obtain more discriminative and text-independent speaker embedding vectors.With the triplet loss, the distances between the embedding vectors of the same speaker are minimized while those of different speakers are maximized.Also, with the adversarial gradient of the ASR network, the text-dependency of the speaker embedding vector can be reduced.In the experiments, we evaluated our speaker verification framework using the LibriSpeech and CHiME 2013 dataset, and the evaluation results show that our speaker verification framework shows lower equal error rate and better textindependency compared to the other approaches.
Sungrack Yun, Janghoon Cho, Jungyun Eum, Wonil Chang, Kyuwoong Hwang
INTERSPEECH1
2017 Speaker Clustering by Iteratively Finding Discriminative Feature Space and Cluster Labels
Sungrack Yun, Hye Jin Jang, Taesu Kim
INTERSPEECH1
2012 Joint Kernel Learning for Supervised Image Segmentation
Jongmin Kim 0006, Youngjoo Seo, Sanghyuk Park, Sungrack Yun, Chang Dong Yoo
ACCV (1)4
2012 Phoneme Classification using Constrained Variational Gaussian Process Dynamical System
abstract
This paper describes a new acoustic model based on variational Gaussian process dynamical system (VGPDS) for phoneme classification. The proposed model overcomes the limitations of the classical HMM in modeling the real speech data, by adopting a nonlinear and nonparametric model. In our model, the GP prior on the dynamics function enables representing the complex dynamic structure of speech, while the GP prior on the emission function successfully models the global dependency over the observations. Additionally, we introduce variance constraint to the original VGPDS for mitigating sparse approximation error of the kernel matrix. The effectiveness of the proposed model is demonstrated with extensive experimental results including parameter estimation, classification performance on the synthetic and benchmark datasets.
Hyunsin Park, Sungrack Yun, Sanghyuk Park, Jongmin Kim 0006, Chang Dong Yoo
NIPS2
2012 Loss-Scaled Large-Margin Gaussian Mixture Models for Speech Emotion Classification
abstract
This paper considers a learning framework for speech emotion classification using a discriminant function based on Gaussian mixture models (GMMs). The GMM parameter set is estimated by margin scaling with a loss function to reduce the risk of predicting emotions with high loss. Here, the loss function is defined as a function of a distance metric using the Watson and Tellegen's emotion model. Margin scaling is known to have good generalization ability and can be considered appropriate for emotion modeling where the parameter set is likely to be over-fitted to the training data set whose characteristics may differ from those of the testing data set. Our learning framework is formulated as a constrained optimization problem which is solved using semi-definite programming. Three tasks were evaluated: acted emotion classification, natural emotion classification, and cross database emotion classification. In each task, four loss functions were evaluated. In all experiments, results consistently show that margin scaling improves the classification accuracy over other learning frameworks based on the maximum-likelihood, maximum mutual information and max-margin framework without margin scaling. Experiment results also show that margin scaling substantially reduces the overall loss compared to the max-margin framework without margin scaling.
Sungrack Yun, Chang Dong Yoo
IEEE Trans. Speech Audio Process.1
2011 Learning a discriminative visual codebook using homonym scheme
abstract
This paper studies a method for learning a discriminative visual codebook for various computer vision tasks such as image categorization and object recognition. The performance of various computer vision tasks depends on the construction of the code book which is a table of visual-words (i.e. codewords). This paper proposed a learning criterion for constructing a discriminative codebook, and it is solved by the homonym scheme which splits codeword regions by labels. A codebook is learned based on the proposed homonym scheme such that its histogram can be used to discriminate objects of different labels. The traditional codebook based on the k-means is compared against the learned codebook on two well-known datasets (Caltech 101, ETH-80) and a dataset we constructed using google images. We show that the learned codebook consistently outperforms the traditional codebook.
SeungRyul Baek, Chang Dong Yoo, Sungrack Yun
ICASSP3
2011 Large Margin Discriminative Semi-Markov Model for Phonetic Recognition
abstract
This paper considers a large margin discriminative semi-Markov model (LMSMM) for phonetic recognition. The hidden Markov model (HMM) framework that is often used for phonetic recognition assumes only local statistical dependencies between adjacent observations, and it is used to predict a label for each observation without explicit phone segmentation. On the other hand, the semi-Markov model (SMM) framework allows simultaneous segmentation and labeling of sequential data based on a segment-based Markovian structure that assumes statistical dependencies among all the observations within a phone segment. For phonetic recognition which is inherently a joint segmentation and labeling problem, the SMM framework has the potential to perform better than the HMM framework at the expense of slight increase in computational complexity. The SMM framework considered in this paper is based on a non-probabilistic discriminant function that is linear in the joint feature map which attempts to capture long-range statistical dependencies among observations. The parameters of the discriminant function are estimated by a large margin learning framework for structured prediction. The parameter estimation problem in hand leads to an optimization problem with many margin constraints, and this constrained optimization problem is solved using a stochastic gradient descent algorithm. The proposed LMSMM outperformed the large margin discriminative HMM in the TIMIT phonetic recognition task.
Sungwoong Kim, Sungrack Yun, Chang Dong Yoo
IEEE Trans. Speech Audio Process.2
2010 Largemargin training of semi-Markov model for phonetic recognition
abstract
This paper considers a large margin training of semi-Markov model (SMM) for phonetic recognition. The SMM framework is better suited for phonetic recognition than the hidden Markov model (HMM) framework in that the SMM framework is capable of simultaneously segmenting the uttered speech into phones and labeling the segment-based features. In this paper, the SMM framework is used to define a discriminant function that is linear in the joint feature map which attempts to capture the long-range statistical dependencies within a segment and between adjacent segments of variable length. The parameters of the discriminant function are estimated by a large margin learning criterion for structured prediction. The parameter estimation problem, which is an optimization problem with many margin constraints, is solved by using a stochastic subgradient descent algorithm. The proposed large margin SMM outperforms the large margin HMM on the TIMIT corpus.
Sungwoong Kim, Sungrack Yun, Chang Dong Yoo
ICASSP2
2010 Parametric emotional singing voice synthesis
abstract
This paper describes an algorithm to control the expressed emotion of a synthesized song. Based on the database of various melodies sung neutrally with restricted set of words, hidden semi-Markov models (HSMMs) of notes ranging from E3 to G5 are constructed for synthesizing singing voice. Three steps are taken in the synthesis: (1) Pitch and duration are determined according to the notes indicated by the musical score; (2) Features are sampled from appropriate HSMMs with the duration set to the maximum probability; (3) Singing voice is synthesized by the mel-log spectrum approximation (MLSA) filter using the sampled features as parameters of the filter. Emotion of a synthesized song is controlled by varying the duration and the vibrato parameters according to the Thayer's mood model. Perception test is performed to evaluate the synthesized song. The results show that the algorithm can control the expressed emotion of a singing voice given a neutral singing voice database.
Younsung Park, Sungrack Yun, Chang Dong Yoo
ICASSP2
2010 Wearable sensor activity analysis using semi-Markov models with a grammar
Owen Thomas, Peter Sunehag, Gideon Dror, Sungrack Yun, Sungwoong Kim, Matthew W. Robards, Alexander J. Smola, Daniel Green, Philo Saunders
Pervasive Mob. Comput.4
2009 Speech emotion recognition via a max-margin framework incorporating a loss function based on the Watson and Tellegen's emotion model
abstract
This paper considers a method for speech emotion recognition by a max-margin framework incorporating a loss function based on a well-known model called theWatson and Tellegen's emotion model. Each emotion is modeled by a single-state hidden Markov model (HMM) that is trained by maximizing the minimum separation margin between emotions, and the margin is scaled by a loss function. The framework is optimized by the semi-definite programming. Experiments were performed to evaluate the framework using the Berlin database of emotional speech. The framework performed better than other conventional training criteria for HMM such as maximum likelihood estimation and maximum mutual information estimation.
Sungrack Yun, Chang Dong Yoo
ICASSP1
2004 Hybrid utterance verification based on n-best models and model derived from kulback-leibler divergence
abstract
In this paper, utterance verification based on hybrid scores obtained from three pairs of models is investigated. The three models considered are the on-line garbage model, the antiword function model and a model derived using Kullback-Leibler divergence. The performance of utterance verification algorithm using hypothesis testing depends on the accuracy of the estimate of the alternative hypothesis. The three models offer different perspectives in the probability estimation of the alternate hypothesis. Performance comparison between hybrid scores using different model pairs is made. In addition, performance improvement over conventional algorithm is experimentally verified.
Minho Jin, Gyucheol Jang, Sungrack Yun, Chang Dong Yoo
INTERSPEECH3