VLDB 2026 Research / reviewers in the wild / expert
Chang Dong Yoo
dblp:31/7819 · also Chang D. Yoo
· DBLP profile ↗
157ranked-venue papers
3as first author
65since 2021 · last 2026
0000-0002-0756-7179ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 113 · 3 first-author · 37 since 2021Artificial intelligence and machine learning · 87 · 50 since 2021Security and privacy · 5Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards unsupervised speech recognition without pronunciation models
Junrui Ni, Liming Wang 0003, Yang Zhang 0001, Kaizhi Qian, Heting Gao, Mark Hasegawa-Johnson, James R. Glass, Chang Dong Yoo |
Speech Commun. | 8 |
| 2025 | Human-Centric AI: From Explainability and Trustworthiness to Actionable EthicsabstractTo address the potential risks of AI while supporting innovation and ensuring responsible adoption, there is an urgent need for clear governance frameworks grounded in human-centric values. It is imperative that AI systems operate in ways that are transparent, trustworthy, and ethically sound. Developing truly human-centric AI goes beyond technical innovation. It requires interdisciplinary collaboration and diverse perspectives. This workshop will explore key challenges and emerging solutions in the development of human-centric AI, with a focus on explainability, trustworthiness, fairness, and privacy. We welcome both theoretical contributions and practical case studies that demonstrate how human-centered principles are realized in real-world AI systems. The official workshop webpage is available at https://xai.kaist.ac.kr/Workshop/hcai2025/, which provides comprehensive information about the program. Jaesik Choi, Bohyung Han, Myoung-Wan Koo, Kyungman Bae, Chang Dong Yoo, Simon S. Woo, Wojciech Samek |
CIKM | 5 |
| 2025 | ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-OnabstractThis paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Masked Diffusion Transformer (MDT) for improved handling of both global garment context and fine-grained details. The IVTON task involves seamlessly superimposing a garment from one image onto a person in another, creating a realistic depiction of the person wearing the specified garment. Unlike conventional diffusion-based virtual try-on models that depend on large pre-trained U-Net architectures, ITA-MDT leverages a lightweight, scalable transformer-based denoising diffusion model with a mask latent modeling scheme, achieving competitive results while reducing computational overhead. A key component of ITA-MDT is the Image-Timestep Adaptive Feature Aggregator (ITAFA), a dynamic feature aggregator that combines all of the features from the image encoder into a unified feature of the same size, guided by diffusion timestep and garment image complexity. This enables adaptive weighting of features, allowing the model to emphasize either global information or fine-grained details based on the requirements of the denoising stage. Additionally, the Salient Region Extractor (SRE) module is presented to identify complex region of the garment to provide high-resolution local information to the denoising model as an additional condition alongside the global information of the full garment image. This targeted conditioning strategy enhances detail preservation of fine details in highly salient garment regions, optimizing computational resources by avoiding unnecessarily processing entire garment image. Comparative evaluations confirms that ITA-MDT improves efficiency while maintaining strong performance, reaching state-of-the-art results in several metrics. Our project page is available at https://jiwoohong93.github.io/ita-mdt/. Ji Woo Hong, Tri Ton, Trung X. Pham, Gwanhyeong Koo, Sunjae Yoon, Chang Dong Yoo |
CVPR | 6 |
| 2025 | Sample Efficient Reinforcement Learning via Large Vision Language Model DistillationabstractRecent research highlights the potential of multi-modal foundation models in tackling complex decision-making challenges. However, their large parameters make real-world deployment resource-intensive and often impractical for constrained systems. Reinforcement learning (RL) shows promise for task-specific agents but suffers from high sample complexity, limiting practical applications. To address these challenges, we introduce LVLM to Policy (LVLM2P), a novel framework that distills knowledge from large vision-language models (LVLM) into more efficient RL agents. Our approach leverages the LVLM as a teacher, providing instructional actions based on trajectories collected by the RL agent, which helps reduce less meaningful exploration in the early stages of learning, thereby significantly accelerating the agent’s learning progress. Additionally, by leveraging the LVLM to suggest actions directly from visual observations, we eliminate the need for manual textual descriptors of the environment, enhancing applicability across diverse tasks. Experiments show that LVLM2P significantly enhances the sample efficiency of baseline RL algorithms. The code is available at https://github.com/i22024/LVLM2P Tung Minh Luu, Younghwan Lee, Chang Dong Yoo |
ICASSP | 4 |
| 2025 | Reward Generation via Large Vision-Language Model in Offline Reinforcement LearningabstractIn offline reinforcement learning (RL), learning from fixed datasets presents a promising solution for domains where real-time interaction with the environment is expensive or risky. However, designing dense reward signals for offline dataset requires significant human effort and domain expertise. Reinforcement learning with human feedback (RLHF) has emerged as an alternative, but it remains costly due to the human-in-the-loop process, prompting interest in automated reward generation models. To address this, we propose Reward Generation via Large Vision-Language Models (RG-VLM), which leverages the reasoning capabilities of LVLMs to generate rewards from offline data without human involvement. RG-VLM improves generalization in long-horizon tasks and can be seamlessly integrated with the sparse reward signals to enhance task performance, demonstrating its potential as an auxiliary reward signal. Younghwan Lee, Tung Minh Luu, Chang Dong Yoo |
ICASSP | 4 |
| 2025 | TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-To-Audio SynthesisabstractThis paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transformers, which offer stable training and continuous transformations for enhanced synchronization and audio quality, TARO introduces two key innovations: (1) Timestep-Adaptive Representation Alignment (TRA), which dynamically aligns latent representations by adjusting alignment strength based on the noise schedule, ensuring smooth evolution and improved fidelity, and (2) Onset-Aware Conditioning (OAC), which integrates onset cues that serve as sharp event-driven markers of audio-relevant visual moments to enhance synchronization with dynamic visual events. Extensive experiments on the VGGSound and Landscape datasets demonstrate that TARO outperforms prior methods, achieving relatively 53% lower Frechet Distance (FD), 29% lower Frechet Audio Distance (FAD), and a 97.19% Alignment Accuracy, highlighting its superior audio quality and synchronization precision. Tri Ton, Ji Woo Hong, Chang Dong Yoo |
ICCV | 3 |
| 2025 | Occlusion-Robust Stylization for Drawing-Based 3D Animationabstract3D animation aims to generate a 3D animated video from an input image and a target 3D motion sequence. Recent advances in image-to-3D models enable the creation of animations directly from user-hand drawings. Distinguished from conventional 3D animation, drawing-based 3D animation is crucial to preserve artist's unique style properties, such as rough contours and distinct stroke patterns. However, recent methods still exhibit quality deterioration in style properties, especially under occlusions caused by overlapping body parts, leading to contour flickering and stroke blurring. This occurs due to a `stylization pose gap' between training and inference in stylization networks designed to preserve drawing styles in drawing-based 3D animation systems. The stylization pose gap denotes that input target poses used to train the stylization network are always in occlusion-free poses, while target poses encountered in an inference include diverse occlusions under dynamic motions. To this end, we propose Occlusion-robust Stylization Framework (OSF) for drawing-based 3D animation. We found that while employing object's edge can be effective input prior for guiding stylization, it becomes notably inaccurate when occlusions occur at inference. Thus, our proposed OSF provides occlusion-robust edge guidance for stylization network using optical flow, ensuring a consistent stylization even under occlusions. Furthermore, OSF operates in a single run instead of the previous two-stage method, achieving 2.4x faster inference and 2.1x less memory. Sunjae Yoon, Gwanhyeong Koo, Younghwan Lee, Ji Woo Hong, Chang Dong Yoo |
ICCV | 5 |
| 2025 | MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound GenerationabstractWe introduce MDSGen, a novel framework for vision-guided open-domain sound generation optimized for model parameter size, memory consumption, and inference speed. This framework incorporates two key innovations: (1) a redundant video feature removal module that filters out unnecessary visual information, and (2) a temporal-aware masking strategy that leverages temporal context for enhanced audio generation accuracy. In contrast to existing resource-heavy Unet-based models, MDSGen employs denoising masked diffusion transformers, facilitating efficient generation without reliance on pre-trained diffusion models. Evaluated on the benchmark VGGSound dataset, our smallest model (5M parameters) achieves 97.9% alignment accuracy, using 172x fewer parameters, 371% less memory, and offering 36x faster inference than the current 860M-parameter state-of-the-art model (93.9% accuracy). The larger model (131M parameters) reaches nearly 99% accuracy while requiring 6.5x fewer parameters. These results highlight the scalability and effectiveness of our approach. The code is available at https://bit.ly/mdsgen. Trung X. Pham, Tri Ton, Chang Dong Yoo |
ICLR | 3 |
| 2025 | Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language ModelsabstractIn the broader context of deep learning, Multimodal Large Language Models have achieved significant breakthroughs by leveraging powerful Large Language Models as a backbone to align different modalities into the language space. A prime exemplification is the development of Video Large Language Models (Video-LLMs). While numerous advancements have been proposed to enhance the video understanding capabilities of these models, they are predominantly trained on questions generated directly from video content. However, in real-world scenarios, users often pose questions that extend beyond the informational scope of the video, highlighting the need for Video-LLMs to assess the relevance of the question. We demonstrate that even the best-performing Video-LLMs fail to reject unfit questions-not necessarily due to a lack of video understanding, but because they have not been trained to identify and refuse such questions. To address this limitation, we propose alignment for answerability, a framework that equips Video-LLMs with the ability to evaluate the relevance of a question based on the input video and appropriately decline to answer when the question exceeds the scope of the video, as well as an evaluation framework with a comprehensive set of metrics designed to measure model behavior before and after alignment. Furthermore, we present a pipeline for creating a dataset specifically tailored for alignment for answerability, leveraging existing video-description paired datasets. Eunseop Yoon, Hee Suk Yoon, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICLR | 4 |
| 2025 | FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow FieldsabstractDrag-based editing allows precise object manipulation through point-based control, offering user convenience. However, current methods often suffer from a geometric inconsistency problem by focusing exclusively on matching user-defined points, neglecting the broader geometry and leading to artifacts or unstable edits. We propose FlowDrag, which leverages geometric information for more accurate and coherent transformations. Our approach constructs a 3D mesh from the image, using an energy function to guide mesh deformation based on user-defined drag points. The resulting mesh displacements are projected into 2D and incorporated into a UNet denoising process, enabling precise handle-to-target point alignment while preserving structural integrity. Additionally, existing drag-editing benchmarks provide no ground truth, making it difficult to assess how accurately the edits match the intended transformations. To address this, we present VFD (VidFrameDrag) benchmark dataset, which provides ground-truth frames using consecutive shots in a video dataset. FlowDrag outperforms existing drag-based editing methods on both VFD Bench and DragBench. Gwanhyeong Koo, Sunjae Yoon, Younghwan Lee, Ji Woo Hong, Chang Dong Yoo |
ICML | 5 |
| 2025 | Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language ModelsabstractDesigning effective reward functions remains a fundamental challenge in reinforcement learning (RL), as it often requires extensive human effort and domain expertise. While RL from human feedback has been successful in aligning agents with human intent, acquiring high-quality feedback is costly and labor-intensive, limiting its scalability. Recent advancements in foundation models present a promising alternative–leveraging AI-generated feedback to reduce reliance on human supervision in reward learning. Building on this paradigm, we introduce ERL-VLM, an enhanced rating-based RL method that effectively learns reward functions from AI feedback. Unlike prior methods that rely on pairwise comparisons, ERL-VLM queries large vision-language models (VLMs) for absolute ratings of individual trajectories, enabling more expressive feedback and improved sample efficiency. Additionally, we propose key enhancements to rating-based RL, addressing instability issues caused by data imbalance and noisy labels. Through extensive experiments across both low-level and high-level control tasks, we demonstrate that ERL-VLM significantly outperforms existing VLM-based reward generation methods. Our results demonstrate the potential of AI feedback for scaling RL with minimal human intervention, paving the way for more autonomous and efficient reward learning. Tung Minh Luu, Younghwan Lee, Chang Dong Yoo |
ICML | 6 |
| 2025 | ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference OptimizationabstractWe introduce ConfPO, a method for preference learning in Large Language Models (LLMs) that identifies and optimizes preference-critical tokens based solely on the training policy’s confidence, without requiring any auxiliary models or compute. Unlike prior Direct Alignment Algorithms (DAAs) such as Direct Preference Optimization (DPO), which uniformly adjust all token probabilities regardless of their relevance to preference, ConfPO focuses optimization on the most impactful tokens. This targeted approach improves alignment quality while mitigating overoptimization (i.e., reward hacking) by using the KL divergence budget more efficiently. In contrast to recent token-level methods that rely on credit-assignment models or AI annotators, raising concerns about scalability and reliability, ConfPO is simple, lightweight, and model-free. Experimental results on challenging alignment benchmarks, including AlpacaEval 2 and Arena-Hard, demonstrate that ConfPO consistently outperforms uniform DAAs across various LLMs, delivering better alignment with zero additional computational overhead. Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson, Sungwoong Kim, Chang Dong Yoo |
ICML | 5 |
| 2025 | The Interspeech 2025 Speech Accessibility Project Challenge
Xiuwen Zheng 0003, Bornali Phukan, Jonghwan Na, Edward Cutrell, Kyu J. Han, Mark Hasegawa-Johnson, Pan-Pan Jiang, Aadhrik Kuila, Colin Lea, Bob MacDonald, Gautam Varma Mantena, Venkatesh Ravichandran, Leda Sari, Katrin Tomanek, Chang Dong Yoo, Chris Zwilling |
INTERSPEECH | 15 |
| 2025 | SiamCTC: Learning Speech Representations through Monotonic Temporal AlignmentabstractSelf-supervised speech representation learning has made significant progress through Siamese networks, which leverage different views of the same input. However, existing methods often require frame-wise alignment between these views, overlooking the broader linguistic context invariance across different speaking styles. We introduce SiamCTC, a framework that integrates Siamese networks with Connectionist Temporal Classification (CTC) to learn speech representations without strict frame-level correspondence. By employing CTC loss to establish flexible, monotonic alignments between differing temporal realizations of the same content, SiamCTC accommodates speed perturbations and other temporal augmentations. This design relaxes frame-wise constraints while preserving temporal coherence and enhancing robustness to speaking-rate variations in downstream tasks. Our experiments demonstrate that SiamCTC leads to more adaptable speech representations, particularly at diverse speaking rates. SooHwan Eom, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 3 |
| 2025 | Policy Learning from Large Vision-Language Model Feedback Without Reward ModelingabstractOffline reinforcement learning (RL) provides a powerful framework for training robotic agents using pre-collected, suboptimal datasets, eliminating the need for costly, time-consuming, and potentially hazardous online interactions. This is particularly useful in safety-critical real-world applications, where online data collection is expensive and impractical. However, existing offline RL algorithms typically require reward labeled data, which introduces an additional bottleneck: reward function design is itself costly, labor-intensive, and requires significant domain expertise. In this paper, we introduce PLARE, a novel approach that leverages large vision-language models (VLMs) to provide guidance signals for agent training. Instead of relying on manually designed reward functions, PLARE queries a VLM for preference labels on pairs of visual trajectory segments based on a language task description. The policy is then trained directly from these preference labels using a supervised contrastive preference learning objective, bypassing the need to learn explicit reward models. Through extensive experiments on robotic manipulation tasks from the MetaWorld, PLARE achieves performance on par with or surpassing existing state-of-the-art VLM-based reward generation methods. Furthermore, we demonstrate the effectiveness of PLARE in real-world manipulation tasks with a physical robot, further validating its practical applicability. Tung Minh Luu, Younghwan Lee, Chang Dong Yoo |
IROS | 4 |
| 2025 | A Gradient Guidance Perspective on Stepwise Preference Optimization for Diffusion ModelsabstractDirect Preference Optimization (DPO) is a key framework for aligning text-to-image models with human preferences, extended by Stepwise Preference Optimization (SPO) to leverage intermediate steps for preference learning, generating more aesthetically pleasing images with significantly less computational cost. While effective, SPO's underlying mechanisms remain underexplored. In light of this, we critically re-examine SPO by formalizing its mechanism as gradient guidance. This new lens shows that SPO uses biased temporal weighting, giving too little weight to later generative steps, and unlike likelihood centric views it reveals substantial noise in the gradient estimates. Leveraging these insights, our GradSPO algorithm introduces a simplified loss and a targeted, variance-informed noise reduction strategy, enhancing training stability. Evaluations on SD 1.5 and SDXL show GradSPO substantially outperforms leading baselines in human preference, yielding images with markedly improved aesthetics and semantic faithfulness, leading to more robust alignment. Code and models are available at https://github.com/JoshuaTTJ/GradSPO. Joshua Tian Jin Tee, Hee Suk Yoon, Abu Hanif Muhammad Syarubany, Eunseop Yoon, Chang Dong Yoo |
NeurIPS | 5 |
| 2025 | Continual Learning: Forget-Free Winning Subnetworks for Video RepresentationsabstractInspired by the Lottery Ticket Hypothesis (LTH), which highlights the existence of efficient subnetworks within larger, dense networks, a high-performing Winning Subnetwork (WSN) in terms of task performance under appropriate sparsity conditions is considered for various continual learning tasks. It leverages pre-existing weights from dense networks to achieve efficient learning in Task Incremental Learning (TIL) and Task-agnostic Incremental Learning (TaIL) scenarios. In Few-Shot Class Incremental Learning (FSCIL), a variation of WSN referred to as the Soft subnetwork (SoftNet) is designed to prevent overfitting when the data samples are scarce. Furthermore, the sparse reuse of WSN weights is considered for Video Incremental Learning (VIL). The use of Fourier Subneural Operator (FSO) within WSN is considered. It enables compact encoding of videos and identifies reusable subnetworks across varying bandwidths. We have integrated FSO into different architectural frameworks for continual learning, including VIL, TIL, and FSCIL. Our comprehensive experiments demonstrate FSO's effectiveness, significantly improving task performance at various convolutional representational levels. Specifically, FSO enhances higher-layer performance in TIL and FSCIL and lower-layer performance in VIL. Haeyong Kang, Jaehong Yoon, Sung Ju Hwang, Chang Dong Yoo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data AugmentationabstractData augmentation is a crucial component in training neural networks to overcome the limitation imposed by data size, and several techniques have been studied for time series. Although these techniques are effective in certain tasks, they have yet to be generalized to time series benchmarks. We find that current data augmentation techniques ruin the core information contained within the frequency domain. To address this issue, we propose a simple strategy to preserve spectral information (SimPSI) in time series data augmentation. SimPSI preserves the spectral information by mixing the original and augmented input spectrum weighted by a preservation map, which indicates the importance score of each frequency. Specifically, our experimental contributions are to build three distinct preservation maps: magnitude spectrum, saliency map, and spectrum-preservative map. We apply SimPSI to various time series data augmentations and evaluate its effectiveness across a wide range of time series benchmarks. Our experimental results support that SimPSI considerably enhances the performance of time series data augmentations by preserving core spectral information. The source code used in the paper is available at https://github.com/Hyun-Ryu/simpsi. Hyun Ryu, Sunjae Yoon, Hee Suk Yoon, Eunseop Yoon, Chang Dong Yoo |
AAAI | 5 |
| 2024 | FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-rigid Editing
Gwanhyeong Koo, Sunjae Yoon, Ji Woo Hong, Chang Dong Yoo |
ECCV (63) | 4 |
| 2024 | Implicit Steganography Beyond the Constraints of Modality
Sojeong Song, Seoyun Yang, Chang Dong Yoo, Junmo Kim 0002 |
ECCV (86) | 3 |
| 2024 | DNI: Dilutional Noise Initialization for Diffusion Video Editing
Sunjae Yoon, Gwanhyeong Koo, Ji Woo Hong, Chang Dong Yoo |
ECCV (48) | 4 |
| 2024 | BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation
Hee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee, Kang Zhang 0008, Yu-Jung Heo, Du-Seong Chang, Chang Dong Yoo |
ECCV (31) | 7 |
| 2024 | AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech RecognitionabstractIn Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapping. While earlier efforts have tried to combine the CTC loss with an entropy maximization regularization term to mitigate this issue, they employed a constant weighting term on the regularization during the training, which we find may not be optimal. In this work, we introduce Adaptive Maximum Entropy Regularization (AdaMER), a technique that can modulate the impact of entropy regularization throughout the training process. This approach not only refines ASR model training but ensures that as training proceeds, predictions display the desired model confidence. SooHwan Eom, Eunseop Yoon, Hee Suk Yoon, Chanwoo Kim 0001, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 6 |
| 2024 | G2PU: Grapheme-To-Phoneme Transducer with Speech UnitsabstractMost phoneme transcripts are generated using forced alignment: typically a grapheme-to-phoneme transducer (G2P) is applied to text sequences to generate candidate phoneme transcripts, which are then time-aligned to the waveform using an acoustic model. This paper demonstrates, for the first time, simultaneous optimization of the G2P, the acoustic model, and the acoustic alignment to a corpus. To this end, we propose G2PU, a joint CTC-attention model consisting of an encoder-decoder G2P network and an encoder-CTC unit-to-phoneme (U2P) network, where the units are extracted from speech. We demonstrate that the G2P and U2P, operating in parallel, produce lower phone error rates than those of state-of-the-art open-source G2P and forced alignment systems. Furthermore, although the G2P and U2P are trained using parallel speech and text, their synergy can be generalized to text-only test corpora if we also train a grapheme-to-unit (G2U) network that generates speech units from text in the absence of parallel speech. Our G2PU model is trained using phoneme transcripts generated by a teacher G2P tool. Our experiments on Chinese and Japanese show that G2PU reduces phoneme error rate by 7% to 29% relative compared to its teacher. Finally, we include case studies to provide insights into the system’s workings. Heting Gao, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 3 |
| 2024 | Wavelet-Guided Acceleration of Text Inversion in Diffusion-Based Image EditingabstractIn the field of image editing, Null-text Inversion (NTI) enables fine-grained editing while preserving the structure of the original image by optimizing null embeddings during the DDIM sampling process. However, the NTI process is time-consuming, taking more than two minutes per image. To address this, we introduce an innovative method that maintains the principles of the NTI while accelerating the image editing process. We propose the WaveOpt-Estimator, which determines the text optimization endpoint based on frequency characteristics. Utilizing wavelet transform analysis to identify the image’s frequency characteristics, we can limit text optimization to specific timesteps during the DDIM sampling process. By adopting the Negative-Prompt Inversion (NPI) concept, a target prompt representing the original image serves as the initial text value for optimization. This approach maintains performance comparable to NTI while reducing the average editing time by over 80% compared to the NTI method. Our method presents a promising approach for efficient, high-quality image editing based on diffusion models. Gwanhyeong Koo, Sunjae Yoon, Chang Dong Yoo |
ICASSP | 3 |
| 2024 | Unsupervised Speech Recognition with N-skipgram and Positional Unigram MatchingabstractTraining unsupervised speech recognition systems presents challenges due to GAN-associated instability, misalignment between speech and text, and significant memory demands. To tackle these challenges, we introduce a novel ASR system, ESPUM. This system harnesses the power of lower-order N-skipgrams (up to N = 3) combined with positional unigram statistics gathered from a small batch of samples. Evaluated on the TIMIT benchmark, our model showcases competitive performance in ASR and phoneme segmentation tasks. Access our publicly available code at https://github.com/lwang114/GraphUnsupASR. Liming Wang 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 3 |
| 2024 | Querying Easily Flip-flopped Samples for Deep Active LearningabstractActive learning, a paradigm within machine learning, aims to select and query unlabeled data to enhance model performance strategically. A crucial selection strategy leverages the model's predictive uncertainty, reflecting the informativeness of a data point. While the sample's distance to the decision boundary intuitively measures predictive uncertainty, its computation becomes intractable for complex decision boundaries formed in multiclass classification tasks. This paper introduces the *least disagree metric* (LDM), the smallest probability of predicted label disagreement. We propose an asymptotically consistent estimator for LDM under mild assumptions. The estimator boasts computational efficiency and straightforward implementation for deep learning models using parameter perturbation. The LDM-based active learning algorithm queries unlabeled data with the smallest LDM, achieving state-of-the-art *overall* performance across various datasets and deep architectures, as demonstrated by the experimental results. Seong Jin Cho, Gwangsu Kim, Jinwoo Shin, Chang Dong Yoo |
ICLR | 5 |
| 2024 | Progressive Fourier Neural Representation for Sequential Video CompilationabstractNeural Implicit Representation (NIR) has recently gained significant attention due to its remarkable ability to encode complex and high-dimensional data into representation space and easily reconstruct it through a trainable mapping function. However, NIR methods assume a one-to-one mapping between the target data and representation models regardless of data relevancy or similarity. This results in poor generalization over multiple complex data and limits their efficiency and scalability. Motivated by continual learning, this work investigates how to accumulate and transfer neural implicit representations for multiple complex video data over sequential encoding sessions. To overcome the limitation of NIR, we propose a novel method, Progressive Fourier Neural Representation (PFNR), that aims to find an adaptive and compact sub-module in Fourier space to encode videos in each training session. This sparsified neural encoding allows the neural network to hold free weights, enabling an improved adaptation for future videos. In addition, when learning a representation for a new video, PFNR transfers the representation of previous videos with frozen weights. This design allows the model to continuously accumulate high-quality neural representations for multiple videos while ensuring lossless decoding that perfectly preserves the learned representations for previous videos. We validate our PFNR method on the UVG8/17 and DAVIS50 video sequence benchmarks and achieve impressive performance gains over strong continual learning baselines. Haeyong Kang, Jaehong Yoon, Dahyun Kim 0002, Sung Ju Hwang, Chang Dong Yoo |
ICLR | 5 |
| 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature DispersionabstractIn deep learning, test-time adaptation has gained attention as a method for model fine-tuning without the need for labeled data. A prime exemplification is the recently proposed test-time prompt tuning for large-scale vision-language models such as CLIP. Unfortunately, these prompts have been mainly developed to improve accuracy, overlooking the importance of calibration, which is a crucial aspect for quantifying prediction uncertainty. However, traditional calibration methods rely on substantial amounts of labeled data, making them impractical for test-time scenarios. To this end, this paper explores calibration during test-time prompt tuning by leveraging the inherent properties of CLIP. Through a series of observations, we find that the prompt choice significantly affects the calibration in CLIP, where the prompts leading to higher text feature dispersion result in better-calibrated predictions. Introducing the Average Text Feature Dispersion (ATFD), we establish its relationship with calibration error and present a novel method, Calibrated Test-time Prompt Tuning (C-TPT), for optimizing prompts during test-time with enhanced calibration. Through extensive experiments on different CLIP architectures and datasets, we show that C-TPT can effectively improve the calibration of test-time prompt tuning without needing labeled data. The code is publicly accessible at https://github.com/hee-suk-yoon/C-TPT. Hee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee, Mark Hasegawa-Johnson, Yingzhen Li, Chang Dong Yoo |
ICLR | 6 |
| 2024 | Cross-view Masked Diffusion Transformers for Person Image SynthesisabstractWe present X-MDPT ($\underline{Cross}$-view $\underline{M}$asked $\underline{D}$iffusion $\underline{P}$rediction $\underline{T}$ransformers), a novel diffusion model designed for pose-guided human image generation. X-MDPT distinguishes itself by employing masked diffusion transformers that operate on latent patches, a departure from the commonly-used Unet structures in existing works. The model comprises three key modules: 1) a denoising diffusion Transformer, 2) an aggregation network that consolidates conditions into a single vector for the diffusion process, and 3) a mask cross-prediction module that enhances representation learning with semantic information from the reference image. X-MDPT demonstrates scalability, improving FID, SSIM, and LPIPS with larger models. Despite its simple design, our model outperforms state-of-the-art approaches on the DeepFashion dataset while exhibiting efficiency in terms of training parameters, training time, and inference speed. Our compact 33MB model achieves an FID of 7.42, surpassing a prior Unet latent diffusion approach (FID 8.07) using only $11\times$ fewer parameters. Our best model surpasses the pixel-based diffusion with $\frac{2}{3}$ of the parameters and achieves $5.43 \times$ faster inference. The code is available at https://github.com/trungpx/xmdpt. Trung X. Pham, Kang Zhang 0008, Chang Dong Yoo |
ICML | 3 |
| 2024 | FRAG: Frequency Adapting Group for Diffusion Video EditingabstractIn video editing, the hallmark of a quality edit lies in its consistent and unobtrusive adjustment. Modification, when integrated, must be smooth and subtle, preserving the natural flow and aligning seamlessly with the original vision. Therefore, our primary focus is on overcoming the current challenges in high quality edit to ensure that each edit enhances the final product without disrupting its intended essence. However, quality deterioration such as blurring and flickering is routinely observed in recent diffusion video editing systems. We confirm that this deterioration often stems from high-frequency leak: the diffusion model fails to accurately synthesize high-frequency components during denoising process. To this end, we devise Frequency Adapting Group (FRAG) which enhances the video quality in terms of consistency and fidelity by introducing a novel receptive field branch to preserve high-frequency components during the denoising process. FRAG is performed in a model-agnostic manner without additional training and validates the effectiveness on video editing benchmarks (i.e., TGVE, DAVIS). Sunjae Yoon, Gwanhyeong Koo, Chang Dong Yoo |
ICML | 4 |
| 2024 | LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
Eunseop Yoon, Hee Suk Yoon, John B. Harvill, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 5 |
| 2024 | Predictive Coding for Decision TransformerabstractRecent work in offline reinforcement learning (RL) has demonstrated the effectiveness of formulating decision-making as return-conditioned supervised learning. Notably, the decision transformer (DT) architecture has shown promise across various domains. However, despite its initial success, DTs have underperformed on several challenging datasets in goal-conditioned RL. This limitation stems from the inefficiency of return conditioning for guiding policy learning, particularly in unstructured and suboptimal datasets, resulting in DTs failing to effectively learn temporal compositionality. Moreover, this problem might be further exacerbated in long-horizon sparse-reward tasks. To address this challenge, we propose the Predictive Coding for Decision Transformer (PCDT) framework, which leverages generalized future conditioning to enhance DT methods. PCDT utilizes an architecture that extends the DT framework, conditioned on predictive codings, enabling decision-making based on both past and future factors, thereby improving generalization. Through extensive experiments on eight datasets from the AntMaze and FrankaKitchen environments, our proposed method achieves performance on par with or surpassing existing popular value-based and transformer-based methods in offline goal-conditioned RL. Furthermore, we also evaluate our method on a goal-reaching task with a physical robot. Tung Minh Luu, Chang Dong Yoo |
IROS | 3 |
| 2024 | Mitigating Adversarial Perturbations for Deep Reinforcement Learning via Vector QuantizationabstractRecent studies reveal that well-performing reinforcement learning (RL) agents in training often lack resilience against adversarial perturbations during deployment. This highlights the importance of building a robust agent before deploying it in the real world. Most prior works focus on developing robust training-based procedures to tackle this problem, including enhancing the robustness of the deep neural network component itself or adversarially training the agent on strong attacks. In this work, we instead study an input transformation-based defense for RL. Specifically, we propose using a variant of vector quantization (VQ) as a transformation for input observations, which is then used to reduce the space of adversarial attacks during testing, resulting in the transformed observations being less affected by attacks. Our method is computationally efficient and seamlessly integrates with adversarial training, further enhancing the robustness of RL agents against adversarial attacks. Through extensive experiments in multiple environments, we demonstrate that using VQ as the input transformation effectively defends against adversarial attacks on the agent’s observations. Tung Minh Luu, Joshua Tian Jin Tee, Sungwoon Kim, Chang Dong Yoo |
IROS | 5 |
| 2024 | TPC: Test-time Procrustes Calibration for Diffusion-based Human Image AnimationabstractHuman image animation aims to generate a human motion video from the inputs of a reference human image and a target motion video. Current diffusion-based image animation systems exhibit high precision in transferring human identity into targeted motion, yet they still exhibit irregular quality in their outputs. Their optimal precision is achieved only when the physical compositions (i.e., scale and rotation) of the human shapes in the reference image and target pose frame are aligned. In the absence of such alignment, there is a noticeable decline in fidelity and consistency. Especially, in real-world environments, this compositional misalignment commonly occurs, posing significant challenges to the practical usage of current systems. To this end, we propose Test-time Procrustes Calibration (TPC), which enhances the robustness of diffusion-based image animation systems by maintaining optimal performance even when faced with compositional misalignment, effectively addressing real-world scenarios. The TPC provides a calibrated reference image for the diffusion model, enhancing its capability to understand the correspondence between human shapes in the reference and target images. Our method is simple and can be applied to any diffusion-based image animation system in a model-agnostic manner, improving the effectiveness at test time without additional training. Sunjae Yoon, Gwanhyeong Koo, Younghwan Lee, Chang Dong Yoo |
NeurIPS | 4 |
| 2024 | Scalable SoftGroup for 3D Instance Segmentation on Point CloudsabstractThis paper considers a network referred to as SoftGroup for accurate and scalable 3D instance segmentation. Existing state-of-the-art methods produce hard semantic predictions followed by grouping instance segmentation results. Unfortunately, errors stemming from hard decisions propagate into the grouping, resulting in poor overlap between predicted instances and ground truth and substantial false positives. To address the abovementioned problems, SoftGroup allows each point to be associated with multiple classes to mitigate the uncertainty stemming from semantic prediction. It also suppresses false positive instances by learning to categorize them as background. Regarding scalability, the existing fast methods require computational time on the order of tens of seconds on large-scale scenes, which is unsatisfactory and far from applicable for real-time. Our finding is that the$k$-Nearest Neighbor ($k$-NN) module, which serves as the prerequisite of grouping, introduces a computational bottleneck. SoftGroup is extended to resolve this computational bottleneck, referred to as SoftGroup++. The proposed SoftGroup++ reduces time complexity with octree$k$-NN and reduces search space with class-aware pyramid scaling and late devoxelization. Experimental results on various indoor and outdoor datasets demonstrate the efficacy and generality of the proposed SoftGroup and SoftGroup++. Their performances surpass the best-performing baseline by a large margin (6%$\sim$16%) in terms of AP$_{50}$. On datasets with large-scale scenes, SoftGroup++ achieves a 6× speed boost on average compared to SoftGroup. Furthermore, SoftGroup can be extended to perform object detection and panoptic segmentation with nontrivial improvements over existing methods. Thang Vu, Kookhoi Kim, Tung Minh Luu, Junyeong Kim, Chang Dong Yoo |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | A Theory of Unsupervised Speech RecognitionabstractUnsupervised speech recognition (ASR-U) is the problem of learning automatic speech recognition (ASR) systems from unpaired speech-only and text-only corpora.While various algorithms exist to solve this problem, a theoretical framework is missing to study their properties and address such issues as sensitivity to hyperparameters and training instability.In this paper, we proposed a general theoretical framework to study the properties of ASR-U systems based on random matrix theory and the theory of neural tangent kernels.Such a framework allows us to prove various learnability conditions and sample complexity bounds of ASR-U.Extensive ASR-U experiments on synthetic languages with three classes of transition graphs provide strong empirical evidence for our theory (code available at cactuswith- thoughts/UnsupASRTheory.git). Liming Wang 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ACL (1) | 3 |
| 2023 | Counterfactual Two-Stage Debiasing For Video Corpus Moment RetrievalabstractVideo Corpus Moment Retrieval aims to select a temporal video moment pertinent to a given language query from a large video corpus. Existing systems are prone to rely on a retrieval bias as a shortcut, which hinders the systems from accurately learning vision-language association. The retrieval bias is spurious correlations between query and scene. For a given query, systems tend to retrieve incorrectly correlated scenes due to biased annotations that have predominant binding in a dataset. To this end, we present a Counterfactual Two-stage Debiasing Learning (CTDL), which incorporates a counterfactual bias network that intentionally learns the retrieval bias by providing a shortcut to learn the spurious correlation between keyword and scene, and performs two-stage debiasing learning that mitigates the bias via contrasting factual retrievals with counterfactually biased retrievals. Extensive experiments show the effectiveness of CTDL paradigm. Sunjae Yoon, Ji Woo Hong, SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Daehyeok Kim, Junyeong Kim, Chanwoo Kim 0001, Chang Dong Yoo |
ICASSP | 9 |
| 2023 | SCANet: Scene Complexity Aware Network for Weakly-Supervised Video Moment RetrievalabstractVideo moment retrieval aims to localize moments in video corresponding to a given language query. To avoid the expensive cost of annotating the temporal moments, weakly-supervised VMR (wsVMR) systems have been studied. For such systems, generating a number of proposals as moment candidates and then selecting the most appropriate proposal has been a popular approach. These proposals are assumed to contain many distinguishable scenes in a video as candidates. However, existing proposals of wsVMR systems do not respect the varying numbers of scenes in each video, where the proposals are heuristically determined irrespective of the video. We argue that the retrieval system should be able to counter the complexities caused by varying numbers of scenes in each video. To this end, we present a novel concept of a retrieval system referred to as Scene Complexity Aware Network (SCANet), which measures the ‘scene complexity' of multiple scenes in each video and generates adaptive proposals responding to variable complexities of scenes in each video. Experimental results on three retrieval benchmarks (i.e. Charades-STA, ActivityNet, TVR) achieve state-of-the-art performances and demonstrate the effectiveness of incorporating the scene complexity. Sunjae Yoon, Gwanhyeong Koo, Dahyun Kim 0002, Chang Dong Yoo |
ICCV | 4 |
| 2023 | On the Soft-Subnetwork for Few-Shot Class Incremental Learning
Haeyong Kang, Jaehong Yoon, Sultan Rizky Hikmawan Madjid, Sung Ju Hwang, Chang Dong Yoo |
ICLR | 5 |
| 2023 | ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
Hee Suk Yoon, Joshua Tian Jin Tee, Eunseop Yoon, Sunjae Yoon, Gwangsu Kim, Yingzhen Li, Chang Dong Yoo |
ICLR | 7 |
| 2023 | Mitigating the Exposure Bias in Sentence-Level Grapheme-to-Phoneme (G2P) Transduction
Eunseop Yoon, Hee Suk Yoon, Dhananjaya Gowda, SooHwan Eom, Daehyeok Kim, John B. Harvill, Heting Gao, Mark Hasegawa-Johnson, Chanwoo Kim 0001, Chang Dong Yoo |
INTERSPEECH | 10 |
| 2022 | Fast and Efficient MMD-Based Fair PCA via Optimization over Stiefel ManifoldabstractThis paper defines fair principal component analysis (PCA) as minimizing the maximum mean discrepancy (MMD) between the dimensionality-reduced conditional distributions of different protected classes. The incorporation of MMD naturally leads to an exact and tractable mathematical formulation of fairness with good statistical properties. We formulate the problem of fair PCA subject to MMD constraints as a non-convex optimization over the Stiefel manifold and solve it using the Riemannian Exact Penalty Method with Smoothing (REPMS). Importantly, we provide a local optimality guarantee and explicitly show the theoretical effect of each hyperparameter in practical settings, extending previous results. Experimental comparisons based on synthetic and UCI datasets show that our approach outperforms prior work in explained variance, fairness, and runtime. Gwangsu Kim, Mahbod Olfat, Mark Hasegawa-Johnson, Chang Dong Yoo |
AAAI | 5 |
| 2022 | Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech RecognitionabstractPhonemes are defined by their relationship to words: changing a phoneme changes the word.Learning a phoneme inventory with little supervision has been a longstanding challenge with important applications to underresourced speech technology.In this paper, we bridge the gap between the linguistic and statistical definition of phonemes and propose a novel neural discrete representation learning model for self-supervised learning of phoneme inventory with raw speech and word labels.Given the availability of phoneme segmentation and some mild conditions, we prove that the phoneme inventory learned by our approach converges to the true one with an exponentially low error rate.Moreover, in experiments on TIMIT and Mboshi benchmarks, our approach consistently learns a better phonemelevel representation and achieves a lower error rate in a zero-resource phoneme recognition task than previous state-of-the-art selfsupervised representation learning algorithms. Liming Wang 0003, Siyuan Feng 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ACL (1) | 4 |
| 2022 | SoftGroup for 3D Instance Segmentation on Point CloudsabstractExisting state-of-the-art 3D instance segmentation methods perform semantic segmentation followed by grouping. The hard predictions are made when performing semantic segmentation such that each point is associated with a single class. However, the errors stemming from hard decision propagate into grouping that results in (1) low overlaps between the predicted instance with the ground truth and (2) substantial false positives. To address the aforementioned problems, this paper proposes a 3D instance segmentation method referred to as SoftGroup by performing bottom-up soft grouping followed by top-down refinement. SoftGroup allows each point to be associated with multiple classes to mitigate the problems stemming from semantic prediction errors and suppresses false positive instances by learning to categorize them as background. Experimental results on different datasets and multiple evaluation metrics demonstrate the efficacy of SoftGroup. Its performance surpasses the strongest prior method by a significant margin of$+6.2\%$on the ScanNet v2 hidden test set and$+6.8\%$on S3DIS Area 5 in terms of$AP_{50}$. Soft-Group is also fast, running at 345ms per scan with a sin-gle Titan X on ScanNet v2 dataset. The source code and trained models for both datasets are available at https://github.com/thangvubk/SoftGroup.git. Thang Vu, Kookhoi Kim, Tung Minh Luu, Chang Dong Yoo |
CVPR | 5 |
| 2022 | Dual Temperature Helps Contrastive Learning Without Many Negative Samples: Towards Understanding and Simplifying MoCoabstractContrastive learning (CL) is widely known to require many negative samples, 65536 in MoCo for instance, for which the performance of a dictionary-free framework is often inferior because the negative sample size (NSS) is limited by its mini-batch size (MBS). To decouple the NSS from the MBS, a dynamic dictionary has been adopted in a large volume of CL frameworks, among which arguably the most popular one is MoCo family. In essence, MoCo adopts a momentum-based queue dictionary, for which we perform a fine-grained analysis of its size and consistency. We point out that InfoNCE loss used in MoCo implicitly attract anchors to their corresponding positive sample with various strength of penalties and identify such inter-anchor hardness-awareness property as a major reason for the necessity of a large dictionary. Our findings motivate us to simplify MoCo v2 via the removal of its dictionary as well as momentum. Based on an InfoNCE with the proposed dual temperature, our simplified frameworks, Sim-MoCo and SimCo, outperform MoCo v2 by a visible margin. Moreover, our work bridges the gap between CL and non-CL frameworks, contributing to a more unified under-standing of these two mainstream frameworks in SSL. Code is available at: https://bit.ly/3LkQbaT. Chaoning Zhang, Kang Zhang 0008, Trung X. Pham, Axi Niu, Zhinan Qiao, Chang Dong Yoo, In-So Kweon |
CVPR | 6 |
| 2022 | Selective Query-Guided Debiasing for Video Corpus Moment Retrieval
Sunjae Yoon, Ji Woo Hong, Eunseop Yoon, Dahyun Kim 0002, Junyeong Kim, Hee Suk Yoon, Chang Dong Yoo |
ECCV (36) | 7 |
| 2022 | Decoupled Adversarial Contrastive Learning for Self-supervised Adversarial Robustness
Chaoning Zhang, Kang Zhang 0008, Chenshuang Zhang, Axi Niu, Jiu Feng, Chang Dong Yoo, In-So Kweon |
ECCV (30) | 6 |
| 2022 | Information-Theoretic Text Hallucination Reduction for Video-grounded DialogueabstractVideo-grounded Dialogue (VGD) aims to decode an answer sentence to a question regarding a given video and dialogue context. Despite the recent success of multi-modal reasoning to generate answer sentences, existing dialogue systems still suffer from a text hallucination problem, which denotes indiscriminate text-copying from input texts without an understanding of the question. This is due to learning spurious correlations from the fact that answer sentences in the dataset usually include the words of input texts, thus the VGD system excessively relies on copying words from input texts by hoping those words to overlap with ground-truth texts. Hence, we design Text Hallucination Mitigating (THAM) framework, which incorporates Text Hallucination Regularization (THR) loss derived from the proposed information-theoretic text hallucination measurement approach. Applying THAM with current dialogue systems validates the effectiveness on VGD benchmarks (i.e., AVSD@DSTC7 and AVSD@DSTC8) and shows enhanced interpretability. Sunjae Yoon, Eunseop Yoon, Hee Suk Yoon, Junyeong Kim, Chang Dong Yoo |
EMNLP | 5 |
| 2022 | Semantic Association Network for Video Corpus Moment RetrievalabstractThis paper considers Semantic Association Network (SAN) for Video Corpus Moment Retrieval (VCMR) which localizes temporal moment that best corresponds to the given text query in a corpus of videos. Collaborations among common semantics from multi-modal inputs are essential for effectively understanding video together with subtitle and text query. For this collaboration, SAN associates common semantics within the same modality (by Intra Semantic Association) and across different modalities (by Inter Semantic Association) with dedicated module referred to as Modality Semantic Association (MSA). SAN surpasses existing state-of-the-art performance on the TVR and DiDeMo benchmark datasets. Extensive ablation studies and qualitative analyses show the effectiveness of the proposed model. Dahyun Kim 0002, Sunjae Yoon, Ji Woo Hong, Chang Dong Yoo |
ICASSP | 4 |
| 2022 | How Does SimSiam Avoid Collapse Without Negative Samples? A Unified Understanding with Self-supervised Contrastive Learning
Chaoning Zhang, Kang Zhang 0008, Chenshuang Zhang, Trung X. Pham, Chang Dong Yoo, In-So Kweon |
ICLR | 5 |
| 2022 | Forget-free Continual Learning with Winning SubnetworksabstractInspired by Lottery Ticket Hypothesis that competitive subnetworks exist within a dense network, we propose a continual learning method referred to as Winning SubNetworks (WSN), which sequentially learns and selects an optimal subnetwork for each task. Specifically, WSN jointly learns the model weights and task-adaptive binary masks pertaining to subnetworks associated with each task whilst attempting to select a small set of weights to be activated (winning ticket) by reusing weights of the prior subnetworks. The proposed method is inherently immune to catastrophic forgetting as each selected subnetwork model does not infringe upon other subnetworks. Binary masks spawned per winning ticket are encoded into one N-bit binary digit mask, then compressed using Huffman coding for a sub-linear increase in network capacity with respect to the number of tasks. Haeyong Kang, Rusty Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon, Mark Hasegawa-Johnson, Sung Ju Hwang, Chang Dong Yoo |
ICML | 7 |
| 2022 | Frame-Level Stutter Detection
John B. Harvill, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 3 |
| 2022 | Seamless equal accuracy ratio for inclusive CTC speech recognitionabstractConcerns have been raised regarding performance disparity in automatic speech recognition (ASR) systems as they provide unequal transcription accuracy for different user groups defined by different attributes that include gender, dialect, and race. In this paper, we propose “equal accuracy ratio”, a novel inclusiveness measure for ASR systems that can be seamlessly integrated into the standard connectionist temporal classification (CTC) training pipeline of an end-to-end neural speech recognizer to increase the recognizer’s inclusiveness. We also create a novel multi-dialect benchmark dataset to study the inclusiveness of ASR, by combining data from existing corpora in seven dialects of English (African American, General American, Latino English, British English, Indian English, Afrikaaner English, and Xhosa English). Experiments on this multi-dialect corpus show that using the equal accuracy ratio as a regularization term along with CTC loss, succeeds in lowering the accuracy gap between user groups and reduces the recognition error rate compared with a non-regularized baseline. Experiments on additional speech corpora that have different user groups also confirm our findings. Heting Gao, Sunghun Kang, Rusty Mina, Dias Issa, John B. Harvill, Leda Sari, Mark Hasegawa-Johnson, Chang Dong Yoo |
Speech Commun. | 9 |
| 2021 | Structured Co-reference Graph Attention for Video-grounded DialogueabstractA video-grounded dialogue system referred to as the Structured Co-reference Graph Attention (SCGA) is presented for decoding the answer sequence to a question regarding a given video while keeping track of the dialogue context. Although recent efforts have made great strides in improving the quality of the response, performance is still far from satisfactory. The two main challenging issues are as follows: (1) how to deduce co-reference among multiple modalities and (2) how to reason on the rich underlying semantic structure of video with complex spatial and temporal dynamics. To this end, SCGA is based on (1) Structured Co-reference Resolver that performs dereferencing via building a structured graph over multiple modalities, (2) Spatio-temporal Video Reasoner that captures local-to-global dynamics of video via gradually neighboring graph attention. SCGA makes use of pointer network to dynamically replicate parts of the question for decoding the answer sequence. The validity of the proposed SCGA is demonstrated on AVSD@DSTC7 and AVSD@DSTC8 datasets, a challenging video-grounded dialogue benchmarks, and TVQA dataset, a large-scale videoQA benchmark. Our empirical results show that SCGA outperforms other state-of-the-art dialogue systems on both benchmarks, while extensive ablation study and qualitative analysis reveal performance gain and improved interpretability. Junyeong Kim, Sunjae Yoon, Dahyun Kim 0002, Chang Dong Yoo |
AAAI | 4 |
| 2021 | Semantic Grouping Network for Video CaptioningabstractThis paper considers a video caption generating network referred to as Semantic Grouping Network (SGN) that attempts (1) to group video frames with discriminating word phrases of partially decoded caption and then (2) to decode those semantically aligned groups in predicting the next word. As consecutive frames are not likely to provide unique information, prior methods have focused on discarding or merging repetitive information based only on the input video. The SGN learns an algorithm to capture the most discriminating word phrases of the partially decoded caption and a mapping that associates each phrase to the relevant video frames - establishing this mapping allows semantically related frames to be clustered, which reduces redundancy. In contrast to the prior methods, the continuous feedback from decoded words enables the SGN to dynamically update the video representation that adapts to the partially decoded caption. Furthermore, a contrastive attention loss is proposed to facilitate accurate alignment between a word phrase and video frames without manual annotations. The SGN achieves state-of-the-art performances by outperforming runner-up methods by a margin of 2.1%p and 2.4%p in a CIDEr-D score on MSVD and MSR-VTT datasets, respectively. Extensive experiments demonstrate the effectiveness and interpretability of the SGN. Hobin Ryu, Sunghun Kang, Haeyong Kang, Chang Dong Yoo |
AAAI | 4 |
| 2021 | SCNet: Training Inference Sample Consistency for Instance SegmentationabstractCascaded architectures have brought significant performance improvement in object detection and instance segmentation. However, there are lingering issues regarding the disparity in the Intersection-over-Union (IoU) distribution of the samples between training and inference. This disparity can potentially exacerbate detection accuracy. This paper proposes an architecture referred to as Sample Consistency Network (SCNet) to ensure that the IoU distribution of the samples at training time is close to that at inference time. Furthermore, SCNet incorporates feature relay and utilizes global contextual information to further reinforce the reciprocal relationships among classifying, detecting, and segmenting sub-tasks. Extensive experiments on the standard COCO dataset reveal the effectiveness of the proposed method over multiple evaluation metrics, including box AP, mask AP, and inference speed. In particular, while running 38\% faster, the proposed SCNet improves the AP of the box and mask predictions by respectively 1.3 and 2.3 points compared to the strong Cascade Mask R-CNN baseline. Code is available at https://github.com/thangvubk/SCNet. Thang Vu, Haeyong Kang, Chang Dong Yoo |
AAAI | 3 |
| 2021 | Synthesis of New Words for Improved Dysarthric Speech Recognition on an Expanded VocabularyabstractDysarthria is a condition where people experience a reduction in speech intelligibility due to a neuromotor disorder. Previous works in dysarthric speech recognition have focused on accurate recognition of words encountered in training data. Due to the rarity of dysarthria in the general population, a relatively small amount of publicly-available training data exists for dysarthric speech. The number of unique words in these datasets is small, so ASR systems trained with existing dysarthric speech data are limited to recognition of those words. In this paper, we propose a data augmentation method using voice conversion that allows dysarthric ASR systems to accurately recognize words outside of the training set vocabulary. We demonstrate that a small amount of dysarthric speech data can be used to capture the relevant vocal characteristics of a speaker with dysarthria through a parallel voice conversion system. We show that it’s possible to synthesize utterances of new words that were never recorded by speakers with dysarthria, and that these synthesized utterances can be used to train a dysarthric ASR system. John B. Harvill, Dias Issa, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 4 |
| 2021 | Robust Maml: Prioritization Task Buffer with Adaptive Learning Process for Model-Agnostic Meta-LearningabstractModel agnostic meta-learning (MAML) is a popular state-of-the-art meta-learning algorithm that provides good weight initialization of a model given a variety of learning tasks. The model initialized by provided weight can be fine-tuned to an unseen task despite only using a small amount of samples and within a few adaptation steps. MAML is simple and versatile but requires costly learning rate tuning and careful design of the task distribution which affects its scalability and generalization. This paper proposes a more robust MAML based on an adaptive learning scheme and a prioritization task buffer (PTB) referred to as Robust MAML (RMAML) for improving scalability of training process and alleviating the problem of distribution mismatch. RMAML uses gradient-based hyper-parameter optimization to automatically find the optimal learning rate and uses the PTB to gradually adjust training task distribution toward testing task distribution over the course of training. Experimental results on meta reinforcement learning environments demonstrate a substantial performance gain as well as being less sensitive to hyper-parameter choice and robust to distribution mismatch. Tung Minh Luu, Trung X. Pham, Sanzhar Rakhimkul, Chang Dong Yoo |
ICASSP | 5 |
| 2021 | Learning Imbalanced Datasets With Maximum Margin LossabstractA learning algorithm referred to as Maximum Margin (MM) is proposed for considering the class-imbalance data learning issue: the deep model tends to predict the majority classes rather than the minority ones. For better generalization on the minority classes, the proposed Maximum Margin (MM) loss function is newly designed by minimizing a margin-based generalization bound through the shifting decision bound. As a prior study, the theoretically principled label-distribution-aware margin (LDAM) loss had been successfully applied with classical strategies such as re-weighting or re-sampling. However, the maximum margin loss function has not been investigated so far. In this study, we evaluate the two types of hard maximum margin-based decision boundary shift with training schedule on artificially imbalanced CIFAR-10 /100 and show the effectiveness. Haeyong Kang, Thang Vu, Chang Dong Yoo |
ICIP | 3 |
| 2021 | Sphererpn: Learning Spheres For High-Quality Region Proposals On 3d Point Clouds Object DetectionabstractA bounding box commonly serves as the proxy for 2D object detection. However, extending this practice to 3D detection raises sensitivity to localization error. This problem is acute on flat objects since small localization error may lead to low overlaps between the prediction and ground truth. To address this problem, this paper proposes Sphere Region Proposal Network (SphereRPN) which detects objects by learning spheres as opposed to bounding boxes. We demonstrate that spherical proposals are more robust to localization error compared to bounding boxes. The proposed SphereRPN is not only accurate but also fast. Experiment results on the standard ScanNet dataset show that the proposed SphereRPN outperforms the previous state-of-the-art methods by a large margin while being $2 \times$ to $7 \times$ faster. The code will be made publicly available. Thang Vu, Kookhoi Kim, Haeyong Kang, Xuan Thanh Nguyen, Tung Minh Luu, Chang Dong Yoo |
ICIP | 6 |
| 2021 | Weakly-Supervised Moment Retrieval Network for Video Corpus Moment RetrievalabstractThis paper proposes Weakly-supervised Moment Retrieval Network (WMRN) for Video Corpus Moment Retrieval (VCMR), which retrieves pertinent temporal moments related to natural language query in a large video corpus. Previous methods for VCMR require full supervision of temporal boundary information for training, which involves a labor-intensive process of annotating the boundaries in a large number of videos. To leverage this, the proposed WMRN performs VCMR in a weakly-supervised manner, where WMRN is learned without ground-truth labels but only with video and text queries. For weakly-supervised VCMR, WMRN addresses the following two limitations of prior methods: (1) Blurry attention over video features due to redundant video candidate proposals generation, (2) Insufficient learning due to weak supervision only with video-query pairs. To this end, WMRN is based on (1) Text Guided Proposal Generation (TGPG) that effectively generates text guided multi-scale video proposals in the prospective region related to query, and (2) Hard Negative Proposal Sampling (HNPS) that enhances video-language alignment via extracting negative video proposals in positive video sample for contrastive learning. Experimental results show that WMRN achieves state-of-the-art performance on TVR and DiDeMo benchmarks in the weakly-supervised setting. To validate the attainments of proposed components of WMRN, comprehensive ablation studies and qualitative analysis are conducted. Sunjae Yoon, Dahyun Kim 0002, Ji Woo Hong, Junyeong Kim, Kookhoi Kim, Chang Dong Yoo |
ICIP | 6 |
| 2021 | Sample-efficient Reinforcement Learning Representation Learning with Curiosity Contrastive Forward Dynamics ModelabstractDeveloping an agent in reinforcement learning (RL) that is capable of performing complex control tasks directly from high-dimensional observation such as raw pixels is a challenge as efforts still need to be made towards improving sample efficiency and generalization of RL algorithm. This paper considers a learning framework for a Curiosity Contrastive Forward Dynamics Model (CCFDM) to achieve a more sample-efficient RL based directly on raw pixels. CCFDM incorporates a forward dynamics model (FDM) and performs contrastive learning to train its deep convolutional neural network-based image encoder (IE) to extract conducive spatial and temporal information to achieve a more sample efficiency for RL. In addition, during training, CCFDM provides intrinsic rewards, produced based on FDM prediction error, and encourages the curiosity of the RL agent to improve exploration. The diverge and less-repetitive observations provided by both our exploration strategy and data augmentation available in contrastive learning improve not only the sample efficiency but also the generalization . Performance of existing model-free RL methods such as Soft Actor-Critic built on top of CCFDM outperforms prior state-of-the-art pixel-based RL methods on the DeepMind Control Suite benchmark. Tung Minh Luu, Thang Vu, Chang Dong Yoo |
IROS | 4 |
| 2021 | Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question AnsweringabstractKiran Ramnath, Leda Sari, Mark Hasegawa-Johnson, Chang Yoo. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Kiran Ramnath, Leda Sari, Mark Hasegawa-Johnson, Chang Dong Yoo |
NAACL-HLT | 4 |
| 2021 | Counterfactually Fair Automatic Speech RecognitionabstractWidely used automatic speech recognition (ASR) systems have been empirically demonstrated in various studies to be unfair, having higher error rates for some groups of users than others. One way to define fairness in ASR is to require that changing the demographic group affiliation of any individual (e.g., changing their gender, age, education or race) should not change the probability distribution across possible speech-to-text transcriptions. In the paradigm of counterfactual fairness, all variables independent of group affiliation (e.g., the text being read by the speaker) remain unchanged, while variables dependent on group affiliation (e.g., the speaker's voice) are counterfactually modified. Hence, we approach the fairness of ASR by training the ASR to minimize change in its outcome probabilities despite a counterfactual change in the individual's demographic attributes. Starting from the individualized counterfactual equal odds criterion, we provide relaxations to it and compare their performances for connectionist temporal classification (CTC) based end-to-end ASR systems. We perform our experiments on the Corpus of Regional African American Languages (CORAAL) and the LibriSpeech dataset to accommodate for differences due to gender, age, education, and race. We show that with counterfactual training, we can reduce average character error rates while achieving lower performance gap between demographic groups, and lower error standard deviation among individuals. Leda Sari, Mark Hasegawa-Johnson, Chang Dong Yoo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Modality Shifting Attention Network for Multi-Modal Video Question AnsweringabstractThis paper considers a network referred to as Modality Shifting Attention Network (MSAN) for Multimodal Video Question Answering (MVQA) task. MSAN decomposes the task into two sub-tasks: (1) localization of temporal moment relevant to the question, and (2) accurate prediction of the answer based on the localized moment. The modality required for temporal localization may be different from that for answer prediction, and this ability to shift modality is essential for performing the task. To this end, MSAN is based on (1) the moment proposal network (MPN) that attempts to locate the most appropriate temporal moment from each of the modalities, and also on (2) the heterogeneous reasoning network (HRN) that predicts the answer using an attention mechanism on both modalities. MSAN is able to place importance weight on the two modalities for each sub-task using a component referred to as Modality Importance Modulation (MIM). Experimental results show that MSAN outperforms previous state-of-the-art by achieving 71.13\% test accuracy on TVQA benchmark dataset. Extensive ablation studies and qualitative analysis are conducted to validate various components of the network. Junyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim 0003, Chang Dong Yoo |
CVPR | 5 |
| 2020 | Learning Augmentation Network via Influence FunctionsabstractData augmentation can impact the generalization performance of an image classification model in a significant way. However, it is currently conducted on the basis of trial and error, and its impact on the generalization performance cannot be predicted during training. This paper considers an influence function that predicts how generalization performance, in terms of validation loss, is affected by a particular augmented training sample. The influence function provides an approximation of the change in validation loss without actually comparing the performances that include and exclude the sample in the training process. Based on this function, a differentiable augmentation network is learned to augment an input training sample to reduce validation loss. The augmented sample is fed into the classification network, and its influence is approximated as a function of the parameters of the last fully-connected layer of the classification network. By backpropagating the influence to the augmentation network, the augmentation network parameters are learned. Experimental results on CIFAR-10, CIFAR-100, and ImageNet show that the proposed method provides better generalization performance than conventional data augmentation methods do. Hyunsin Park, Trung X. Pham, Chang Dong Yoo |
CVPR | 4 |
| 2020 | VLANet: Video-Language Alignment Network for Weakly-Supervised Video Moment Retrieval
Minuk Ma, Sunjae Yoon, Junyeong Kim, Youngjoon Lee, Sunghun Kang, Chang Dong Yoo |
ECCV (28) | 6 |
| 2020 | GAPNet: Generic-Attribute-Pose Network For Fine-Grained Visual Categorization Using Multi-Attribute Attention ModuleabstractThis paper proposes a multi-task learning framework for fine-grained visual categorization (FGVC) referred to as Generic-Attribute-Pose Network (GAPNet) that is capable of attending discriminating parts depending on the pose and part-attribute of an object using multi-attribute attention. FGVC is a challenging task that involves categorical data with small inter-class variation and large intra-class variation. Multi-Attribute Attention Module (MAAM) guides the GAPNet to focus on multiple parts of the image feature by emphasizing appropriate feature channels given both pose and part-attribute features. Experiments on Caltech-UCSD Birds and NABirds datasets demonstrate that GAPNet is competitive with other state-of-the-art methods, and ablation study on GAPNet conditioned on pose and part-attribute feature shows that GAPNet performs best when conditioned on both pose and part-attribute features. Minjeong Ju, Hobin Ryu, Sangkeun Moon, Chang Dong Yoo |
ICIP | 4 |
| 2020 | CNN-Based Learnable Gammatone Filterbank and Equal-Loudness Normalization for Environmental Sound ClassificationabstractFor environmental sound classification (ESC), this letter presents a learnable auditory filterbank based on a one-dimensional (1D) convolutional neural network with strong psychophysiological inductive bias in the form of a gammatone filterbank and an equal-loudness prompting normalization. In the past, a number of ESC methods based on learnable auditory features obtained by performing plain 1D convolutions on raw input waveforms for outperforming traditional handcrafted features such as a mel-frequency filterbank have been proposed. However, the large number of parameters involved in the convolutions suggests that these methods will not generalize better than a model defined by a smaller number of parameters, which is considered in this letter. Here, a learnable gammatone filterbank layer consisting of 1D kernels represented by a parametric form of the bandpass gammatone filters is proposed for acquiring a time-frequency representation of the raw waveform. A normalization with learnable parameters that control the trade-off between energy equalization and structure preservation in the spectro-temporal domain is proposed. To verify the effectiveness of the considered network and the normalization, ESC experiments on the ESC-50 and UrbanSound8K datasets were conducted. Compared to other state-of-the-art networks, the considered network performed better on the two datasets. In addition, an ensemble architecture achieved further performance improvement. Hyunsin Park, Chang Dong Yoo |
IEEE Signal Process. Lett. | 2 |
| 2019 | Edge-Labeling Graph Neural Network for Few-Shot LearningabstractIn this paper, we propose a novel edge-labeling graph neural network (EGNN), which adapts a deep neural network on the edge-labeling graph, for few-shot learning. The previous graph neural network (GNN) approaches in few-shot learning have been based on the node-labeling framework, which implicitly models the intra-cluster similarity and the inter-cluster dissimilarity. In contrast, the proposed EGNN learns to predict the edge-labels rather than the node-labels on the graph that enables the evolution of an explicit clustering by iteratively updating the edge-labels with direct exploitation of both intra-cluster similarity and the inter-cluster dissimilarity. It is also well suited for performing on various numbers of classes without retraining, and can be easily extended to perform a transductive inference. The parameters of the EGNN are learned by episodic training with an edge-labeling loss to obtain a well-generalizable model for unseen low-data problem. On both of the supervised and semi-supervised few-shot image classification tasks with two benchmark datasets, the proposed EGNN significantly improves the performances over the existing GNNs. Jongmin Kim 0006, Taesup Kim, Sungwoong Kim, Chang Dong Yoo |
CVPR | 4 |
| 2019 | Progressive Attention Memory Network for Movie Story Question AnsweringabstractThis paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to answer the question is difficult as the movies are typically longer than an hour, (2) it has both video and subtitle where different questions require different modality to infer the answer. To overcome these challenges, PAMN involves three main features: (1) progressive attention mechanism that utilizes cues from both question and answer to progressively prune out irrelevant temporal parts in memory, (2) dynamic modality fusion that adaptively determines the contribution of each modality for answering the current question, and (3) belief correction answering scheme that successively corrects the prediction score on each candidate answer. Experiments on publicly available benchmark datasets, MovieQA and TVQA, demonstrate that each feature contributes to our movie story QA architecture, PAMN, and improves performance to achieve the state-of-the-art result. Qualitative analysis by visualizing the inference mechanism of PAMN is also provided. Junyeong Kim, Minuk Ma, Kyungsu Kim 0003, Chang Dong Yoo |
CVPR | 5 |
| 2019 | Few-Shot Associative Domain Adaptation for Surface Normal EstimationabstractThis paper considers a surface-normal learning algorithm referred to as few-shot kernel associative domain adaptation (FS-KADA) that reduces the domain shift between abundant synthetic source normals and a few real target normals. The FS-KADA takes an unpaired source and target samples as input and captures invariant representations. However, models trained on synthetically rendered normals do not perform well when accurately predicting real environmental normals due to the domain shift. To address this issue, a contextual weighting is considered for learning FS-KADA on the neighborhood of target ground truth, with kernel association in latent spaces and smoothing at predictions. FS-KADA is evaluated on both a real outdoor target dataset (SNOW) and real indoor datasets (NYUv2) using a synthetic indoor dataset (MLT). The state-of-the-art performance was observed on the SNOW dataset. The performance of FS-KADA using a single ground truth of a randomly selected pixel in each image of the NYUv2 is compared with others using the full ground truth. Haeyong Kang, Gwangsu Kim, Chang Dong Yoo |
ICIP | 3 |
| 2019 | Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question AnsweringabstractThis paper proposes a method to gain extra supervision via multi-task learning for multi-modal video question answering. Multi-modal video question answering is an important task that aims at the joint understanding of vision and language. However, establishing large scale dataset for multi-modal video question answering is expensive and the existing benchmarks are relatively small to provide sufficient supervision. To overcome this challenge, this paper proposes a multi-task learning method which is composed of three main components: (1) multi-modal video question answering network that answers the question based on the both video and subtitle feature, (2) temporal retrieval network that predicts the time in the video clip where the question was generated from and (3) modality alignment network that solves metric learning problem to find correct association of video and subtitle modalities. By simultaneously solving related auxiliary tasks with hierarchically shared intermediate layers, the extra synergistic supervisions are provided. Motivated by curriculum learning, multi-task ratio scheduling is proposed to learn easier task earlier to set inductive bias at the beginning of the training. The experiments on publicly available dataset TVQA shows state-of-the-art results, and ablation studies are conducted to prove the statistical validity. Junyeong Kim, Minuk Ma, Kyungsu Kim 0003, Chang Dong Yoo |
IJCNN | 5 |
| 2019 | Cascade RPN: Delving into High-Quality Region Proposal Network with Adaptive ConvolutionabstractThis paper considers an architecture referred to as Cascade Region Proposal Network (Cascade RPN) for improving the region-proposal quality and detection performance by systematically addressing the limitation of the conventional RPN that heuristically defines the anchors and aligns the features to the anchors. First, instead of using multiple anchors with predefined scales and aspect ratios, Cascade RPN relies on a single anchor per location and performs multi-stage refinement. Each stage is progressively more stringent in defining positive samples by starting out with an anchor-free metric followed by anchor-based metrics in the ensuing stages. Second, to attain alignment between the features and the anchors throughout the stages, adaptive convolution is proposed that takes the anchors in addition to the image features as its input and learns the sampled features guided by the anchors. A simple implementation of a two-stage Cascade RPN achieves 13.4 point AR higher than that of the conventional RPN, surpassing any existing region proposal methods. When adopting to Fast R-CNN and Faster R-CNN, Cascade RPN can improve the detection mAP by 3.1 and 3.5 points, respectively. The code will be made publicly available at https://github.com/thangvubk/Cascade-RPN. Thang Vu, Hyunjun Jang, Trung X. Pham, Chang Dong Yoo |
NeurIPS | 4 |
| 2018 | ImaGAN: Unsupervised Training of Conditional Joint CycleGAN for Transferring Style with Core Structures in Content Preserved
Kangmin Bae, Minuk Ma, Hyunjun Jang, Minjeong Ju, Hyoungwoo Park, Chang Dong Yoo |
ACCV (2) | 6 |
| 2018 | Pivot Correlational Neural Network for Multimodal Video Categorization
Sunghun Kang, Junyeong Kim, Hyunsoo Choi, Chang Dong Yoo |
ECCV (14) | 5 |
| 2018 | Action Recognition: First-and Second-Order 3D Feature in Bi-Directional Attention NetworkabstractThis paper considers a 3D convolutional neural network (CNN) that learns spatial and temporal regions of higher importance through a bi-direction long short-term memory (bi-LSTM) attention for action recognition. First- and second-order differences of spatially most relevant C3D features (sp-m-C3D) are obtained, and the concatenation of the two differences with the sp-m-C3D is used to generate a temporal attention on the sp-m-C3D using a bi-LSTM. Temporally most relevant sp-m-C3D features are fed into another bi-LSTM for action recognition. Essentially, the network learns spatial and temporal regions of high importance for action recognition. We evaluate the network on two public action recognition datasets: UCF-101 (YouTube Action) and HMDB51. The proposed network performs better compared to other state-of-the-art networks. Oh Chul Kwon, Junyeong Kim, Chang Dong Yoo |
ICIP | 3 |
| 2018 | Unsupervised Domain Adaptation for Object Detection Using Distribution Matching in Various Feature Level
Hyoungwoo Park, Minjeong Ju, Sangkeun Moon, Chang Dong Yoo |
IWDW | 4 |
| 2017 | Melody extraction and detection through LSTM-RNN with harmonic sum lossabstractThis paper proposes a long short-term memory recurrent neural network (LSTM-RNN) for extracting melody and simultaneously detecting regions of melody from polyphonic audio using the proposed harmonic sum loss. The previous state-of-the-art algorithms have not been based on machine learning techniques and certainly not on deep architectures. The harmonics structure in melody is incorporated in the loss function to attain robustness against both octave mismatch and interference from background music. Experimental results show that the performance of the proposed method is better than or comparable to other state-of-the-art algorithms. Hyunsin Park, Chang Dong Yoo |
ICASSP | 2 |
| 2017 | Multi-view visual speech recognition based on multi task learningabstractVisual speech recognition (VSR), also known as lipreading is a task that recognizes word or phrase using video clip of lip movement. Traditional VSR methods are limited in that they are based mostly on VSR of frontal view facial movement. This limitation should be relaxed to include lip movement from all angles. In this paper, we propose a pose-invariant network which can recognize word spoken from any arbitrary view input. The architecture that combines convolutional neural network (CNN) with bidirectional long short-term memory (LSTM) is trained in a multi-task manner such that the pose and the word spoken are jointly classified. Here, pose classification is considered as the auxiliary task. To comparatively evaluate the performance of the proposed multi-task learning, OuluVS2 benchmark dataset is considered. The experimental results show that the deep model learned based on the proposed multi-task learning method prove its advantage compared to previous single-view VSR methods and also previous multi-view lipreading methods. This deep model achieved recognition performance of 95.0% accuracy on OuluVS2 dataset. HouJeung Han, Sunghun Kang, Chang Dong Yoo |
ICIP | 3 |
| 2017 | Deep partial person re-identification via attention modelabstractThis paper considers a novel algorithm referred to as deep partial person re-identification (DPPR) for partial person re-identification where only a part of a person is observed and full body images are available for identification. The DPPR is based on an end-to-end deep model which make use of convolutional neural network (CNN), RoI Pooling layer and attention model. The RoI Pooling layer enables the extraction of feature vector corresponding to predefined part of input image. The attention model selects a subset of CNN feature vectors. For qualitative evaluation of proposed model, data from CUHK03 are randomly cropped in constructing p-CUHK03. Experimental results show that DPPR outperforms our baseline model on p-CUHK03. Junyeong Kim, Chang Dong Yoo |
ICIP | 2 |
| 2017 | Content adaptive video summarization using spatio-temporal featuresabstractThis paper proposes a video summarization method based on novel spatio-temporal features that combine motion magnitude, object class prediction, and saturation. Motion magnitude measures how much motion there is in a video. Object class prediction provides information about an object in a video. Saturation measures the colorfulness of a video. Con-volutional neural networks (CNNs) are incorporated for object class prediction. The sum of the normalized features per shot are ranked in descending order, and the summary is determined by the highest ranking shots. This ranking can be conditioned on the object class, and the high-ranking shots for different object classes are also proposed as a summary of the input video. The performance of the summarization method is evaluated on the SumMe datasets, and the results reveal that the proposed method achieves better performance than the summary of worst human and most other state-of-the-art video summarization methods. Chang Dong Yoo |
ICIP | 2 |
| 2017 | Complex Video Scene Analysis Using Kernelized-Collaborative Behavior Pattern Learning Based on Hierarchical Representative Object BehaviorsabstractThis paper considers an unsupervised learning algorithm that can automatically discover key behavior patterns to characterize a complex video scene. For behavior features (bFs) extracted at multiple spatial-temporal scales, an optimization problem is formulated to cluster bFs in their scales while establishing a collaborative nonlinear relationship in the form of a kernel regression function among clustered bFs across different spatial scales. The relationship allows features extracted in one scale to be considered as contextual information in the analysis of another scale. This optimization problem is solved using linear programming to reduce computational complexity. The proposed algorithm is evaluated on four crowded traffic scenes and two sports video data sets. Experimental results show that the proposed algorithm achieves a better performance compared with the current state-of-the-art algorithms in terms of video segmentation accuracy. Sanghyuk Park, Hyunsin Park, Chang Dong Yoo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Multimodal representation: Kneser-ney smoothing/skip-gram based neural language modelabstractFor image retrieval and caption generation, this paper considers a multimodal representation that associates image with its text description (caption) by defining a neural language model as the conditional probability of the next word given both n past words in a caption and the image that the caption describes. To address the data sparsity problem, the use of the Kneser-Ney smoothing and skip-gram models is examined by integrating each into the multimodal neural language model. A language model (LM) known as Kneser-Ney smoothing is based on absolute-discounting interpolation while skip-gram LM is based on n-grams organized by allowing intermediate tokens to be “skipped”. The multimodal representation is evaluated on the IAPR TC-12 dataset. Using perplexity and BLEU-n measures, both Kneser-Ney smoothing and skip-gram models are demonstrated to be more effective as approaches to addressing the data sparsity problem than the generic n-gram model used in previous multimodal representations. The modality-biased log-bilinear (MLBL-B) model is set as the base model in the experiment. Mingoo Song, Chang Dong Yoo |
ICIP | 2 |
| 2015 | Face alignment using cascade Gaussian process regression treesabstractIn this paper, we propose a face alignment method that uses cascade Gaussian process regression trees (cGPRT) constructed by combining Gaussian process regression trees (GPRT) in a cascade stage-wise manner. Here, GPRT is a Gaussian process with a kernel defined by a set of trees. The kernel measures the similarity between two inputs as the number of trees where the two inputs fall in the same leaves. Without increasing prediction time, the prediction of cGPRT can be performed in the same framework as the cascade regression trees (CRT) but with better generalization. Features for GPRT are designed using shape-indexed difference of Gaussian (DoG) filter responses sampled from local retinal patterns to increase stability and to attain robustness against geometric variances. Compared with the previous CRT-based face alignment methods that have shown state-of-the-art performances, cGPRT using shape-indexed DoG features performed best on the HELEN and 300-W datasets which are the most challenging dataset today. Hyunsin Park, Chang Dong Yoo |
CVPR | 3 |
| 2015 | Face detection using Local Hybrid PatternsabstractThis paper examines a novel binary feature referred to as the Local Hybrid Patterns (LHP) that is generated by mixing highly discriminative bits of the binary local pattern features (BLPFs) such as the Local Binary Patterns (LBP), Local Gradient Patterns (LGP), and Mean LBP (MLBP). Starting with the most discriminative BLPF selected, the LHP generating algorithm iteratively updates the bits of the selected BLPF by replacing the least discriminative bit with the most discriminative bit of all the candidate BLPFs. At the expense of a small increase in computation, the LHP is guaranteed to give smaller or equal empirical error compared to any BLPFs considered in the pool. Experimental comparison of different sets of features consistently shows that the LHP leads to better performance than previously proposed methods under the AdaBoost face detection framework on MIT+CMU and FDDB benchmark datasets. Chulhee Yun, Chang Dong Yoo |
ICASSP | 3 |
| 2015 | Dense Image Registration and Deformable Surface Reconstruction in Presence of Occlusions and Minimal TextureabstractDeformable surface tracking from monocular images is well-known to be under-constrained. Occlusions often make the task even more challenging, and can result in failure if the surface is not sufficiently textured. In this work, we explicitly address the problem of 3D reconstruction of poorly textured, occluded surfaces, proposing a framework based on a template-matching approach that scales dense robust features by a relevancy score. Our approach is extensively compared to current methods employing both local feature matching and dense template alignment. We test on standard datasets as well as on a new dataset (that will be made publicly available) of a sparsely textured, occluded surface. Our framework achieves state-of-the-art results for both well and poorly textured, occluded surfaces. Dat Tien Ngo, Sanghyuk Park, Anne Jorstad, Alberto Crivellaro, Chang Dong Yoo, Pascal Fua |
ICCV | 5 |
| 2015 | Face attribute classification using attribute-aware correlation map and gated convolutional neural networksabstractThis paper proposes a face attribute classification method based on attribute-aware correlation map and gated convolutional neural networks (CNN). The attribute-aware correlation map provides correlation information between pixel-location and attribute label, and each correlation map of an attribute provides information regarding regions where the relevant features should be extracted. Using the correlation maps of all the attributes, a number of most relevant face part regions are discovered. Based on the face part regions, gated columns of CNNs are simultaneously pre-trained on for face representations then fine-tuned for attribute classification. Here, each CNN column takes input from one of the regions discovered. The column of the CNN is gated such that in the backpropagation of the learning process, classification error due to less relevant attributes do not over influence the learning process. In the experiment, we manually labeled each image in the Labeled Faces in the Wild (LFW) benchmark dataset with 40 face attributes and obtained significant performance improvement over other state-of-the art methods. Sunghun Kang, Chang Dong Yoo |
ICIP | 3 |
| 2015 | Segment-wise online learning based on greedy algorithm for real-time multi-target trackingabstractThis paper proposes a tracklet-based algorithm for online multiple-target tracking. The algorithm performs tracking in three steps: (1) tracklet initialization, (2) tracklet refinement, and (3) tracklet association. Given detection responses, tracklets are initialized by finding a near-optimum path in the min-cost flow network using a greedy-based algorithm. Based on an appearance-based model, the tracklets are refined so that the detection responses within the tracklet become more homogeneous. Finally, the tracklets are linked based on a novel affinity measure, then by optimizing a min-cost flow network with links, the final tracks are generated. For real-time multi-target tracking, every step is processed in a segment-wise manner. On popular public datasets and strictly in an online fashion, the proposed multi-target tracking algorithm performed comparable to that of many state-of-the-art algorithms. Changhoon Lee, Chang Dong Yoo |
ICIP | 2 |
| 2015 | Underdetermined Convolutive BSS: Bayes Risk Minimization Based on a Mixture of Super-Gaussian Posterior ApproximationabstractThis paper considers the underdetermined blind source separation (BSS) of convolutively mixed super-Gaussian signals that include speech, audio, and various other sparse signals. Here, the separation is performed in three steps. In the first and second steps, the mixing matrix and the sources at each time-frequency location are estimated by minimizing the Bayes risk (or the posterior risk) with squared loss. In the final third step, the permutation alignment is conducted by considering the correlation between adjacent spectral bins as in many conventional algorithms. To overcome any computationally intractable integrations involving a complex-valued super-Gaussian source prior, the posterior distribution of the sources is approximated as a mixture of super-Gaussians. The posterior means of the mixing matrix and the sources are obtained with Metropolis-Hastings within Gibbs sampling and the weighted sum of individual super-Gaussians, respectively. Overall, this approximation leads to a separation that is computationally lighter than and as accurate as the algorithm without the approximation. The simulation results of the synthetically generated data in a virtual room with reverberation show that the estimates of the mixing matrix in the first step and the sources in the second step are more accurate than the estimates from the state-of-the-art algorithms in terms of the mixing error ratio (MER) and the signal-to-distortion ratio (SDR). The experiment was also conducted with recorded data in a real room environment using a public benchmark dataset. Results show that the proposed algorithm gives a better performance compared to the state-of-the-art algorithms in terms of the SDR. Janghoon Cho, Chang Dong Yoo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Joint Estimation of Pose and Face Landmark
Junyoung Chung, Chang Dong Yoo |
ACCV (4) | 3 |
| 2014 | Joint learning of foreground region labeling and depth orderingabstractThis paper considers a joint learning algorithm of foreground region labeling and depth ordering for 3D scene understanding. Given an object-level segmentation, the proposed algorithm classifies each region as either foreground or background while simultaneously infers the relative depth orders between every adjacent region pairs. For this, we consider a graph where regions are considered as nodes while boundaries between adjacent regions as edges, and the problem is formulated as jointly assigning binary labels to every nodes and edges via maximizing a unified linear discriminant function, under the constraints that make the resulting depth order to be always physically plausible. Instead of inferring region and edge labels separately, we infer them jointly by grouping them as a single variable referred to as triplet. Then, the problem is reformulated as multi-class triplet prediction to penalize the inconsistent labeling of regions and edges in a soft manner. As the discriminant function is linear, the parameters can be learned with structured support vector machine(S-SVM), and efficient inference using linear programming relaxation is possible. Experimental results show that the proposed joint inference algorithm improves both foreground region labeling and depth ordering performances. Youngjoo Seo, Jongmin Kim 0006, Hoyong Jang, Chang Dong Yoo |
ICASSP | 5 |
| 2014 | Greedy algorithm for real-time multi-object trackingabstractReal-time multi-object tracking is one of the most challenging problem in computer vision. Based on integer linear programming (ILP) formulation for object tracking [1, 2], we propose the greedy algorithm for real-time multi-object tracking. The experimental result shows that the performance of the proposed algorithm is competitive to the previous state-of-the-art algorithms on off-line video data, and its running time is also sufficient for analyzing the tracks of object in real-time video data. Changhoon Lee, Chang Dong Yoo |
ICIP | 3 |
| 2014 | Salient object detection using bipartite dictionaryabstractThis paper considers a bipartite dictionary based salient object detection algorithm that assigns one of two labels (object/background) to each superpixel of an image. The algorithm will iteratively find for each of the labels two dictionaries referred to as the bipartite dictionary, and the dictionaries will in turn update the labels of the superpixels based on the assumption that features of a particular label is better represented by the dictionary of its own label than by the dictionary of the other label. This iteration stops when convergence is reached, in other words, when there is no update. An objective function is formulated such that the bipartite dictionary and superpixel labels maximize inter-class reconstruction error while simultaneously minimize intra-class reconstruction error. The proposed algorithm is evaluated on the MSRA-1000 dataset. Experimental results show that the proposed algorithm performs better than state-of-the-art algorithms for the dataset when the initial conditions are set appropriately. We have also found that the proposed algorithm tends to highlight salient objects more uniformly than other algorithms. Yuna Seo, Chang Dong Yoo |
ICIP | 3 |
| 2014 | A hierarchical-structured dictionary learning for image classificationabstractThis paper proposes a hierarchical-structured discriminative dictionary learning algorithm for image classification. Hierarchical structure of the overall dictionary is learned such that the upper-level dictionaries are specific in representing patterns common across a wide set of class images while lower-level dictionaries are specific in representing patterns localized to a narrow set of class images. Therefore the root dictionary can represent patterns common to all classes, while the leaf dictionaries can represent patterns specific only to a single distinct class. The learned dictionary is efficient in its use of the bases, and leads to a more discriminative representation than that led by previous dictionaries which is devoid of any structure and contains redundant bases. This hierarchical-structured dictionary is learned by solving a constraint optimization problem that minimized reconstruction error of a given image while using dictionaries in the hierarchical structure pertaining only to the class of the image. Sparse representation is pursued in addition, and it acts a regularizer to improve generalization. The representation is as distinct as the paths to each of the class in the hierarchical structure are divergent. To evaluate the effectness of the hierarchical-structured dictionary, classification is performed on three benchmark datasets: Extended Yale B database, Caltech 101 and Caltech 256 dataset, and based on a common features, the proposed algorithm performs better than other state-of-the-art dictionary learning algorithms. Jaesik Yoon, Jinho Choi 0002, Chang Dong Yoo |
ICIP | 3 |
| 2014 | Image Segmentation UsingHigher-Order Correlation ClusteringabstractIn this paper, a hypergraph-based image segmentation framework is formulated in a supervised manner for many high-level computer vision tasks. To consider short- and long-range dependency among various regions of an image and also to incorporate wider selection of features, a higher-order correlation clustering (HO-CC) is incorporated in the framework. Correlation clustering (CC), which is a graph-partitioning algorithm, was recently shown to be effective in a number of applications such as natural language processing, document clustering, and image segmentation. It derives its partitioning result from a pairwise graph by optimizing a global objective function such that it simultaneously maximizes both intra-cluster similarity and inter-cluster dissimilarity. In the HO-CC, the pairwise graph which is used in the CC is generalized to a hypergraph which can alleviate local boundary ambiguities that can occur in the CC. Fast inference is possible by linear programming relaxation, and effective parameter learning by structured support vector machine is also possible by incorporating a decomposable structured loss function. Experimental results on various data sets show that the proposed HO-CC outperforms other state-of-the-art image segmentation algorithms. The HO-CC framework is therefore an efficient and flexible image segmentation framework. Sungwoong Kim, Chang Dong Yoo, Sebastian Nowozin, Pushmeet Kohli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | A maximum likelihood approach for underdetermined TDOA estimationabstractThis paper considers the estimation of time difference of arrival (TDOA) of multiple sparse sources when the number of sources is larger than that of the microphones. White Gaussian noise is assumed present at the microphone in addition to the instantaneously mixed sources. The TDOA estimate is obtained based on a maximum likelihood (ML) criteria, and the likelihood is obtained by marginalizing the joint probability over the sources. Explicit marginalization is mathematically intractable, thus the joint probability is approximated as a summation of several Dirac delta functions by assuming the time-frequency component of the source distribution to be a complex-valued super Gaussian, and the global maximum point of the marginalized joint probability is found by Markov chain Monte Carlo sampling. Experimental results show that the proposed algorithm outperforms TDOA estimation using a well-known Gaussian based approximation method in terms of root-mean-square error (RMSE). Janghoon Cho, Chang Dong Yoo |
ICASSP | 2 |
| 2013 | Task-Specific Image PartitioningabstractImage partitioning is an important preprocessing step for many of the state-of-the-art algorithms used for performing high-level computer vision tasks. Typically, partitioning is conducted without regard to the task in hand. We propose a task-specific image partitioning framework to produce a region-based image representation that will lead to a higher task performance than that reached using any task-oblivious partitioning framework and existing supervised partitioning framework, albeit few in number. The proposed method partitions the image by means of correlation clustering, maximizing a linear discriminant function defined over a superpixel graph. The parameters of the discriminant function that define task-specific similarity/dissimilarity among superpixels are estimated based on structured support vector machine (S-SVM) using task-specific training data. The S-SVM learning leads to a better generalization ability while the construction of the superpixel graph used to define the discriminant function allows a rich set of features to be incorporated to improve discriminability and robustness. We evaluate the learned task-aware partitioning algorithms on three benchmark datasets. Results show that task-aware partitioning leads to better labeling performance than the partitioning computed by the state-of-the-art general-purpose and supervised partitioning algorithms. We believe that the task-specific image partitioning paradigm is widely applicable to improving performance in high-level image understanding tasks. Sungwoong Kim, Sebastian Nowozin, Pushmeet Kohli, Chang Dong Yoo |
IEEE Trans. Image Process. | 4 |
| 2012 | Joint Kernel Learning for Supervised Image Segmentation
Jongmin Kim 0006, Youngjoo Seo, Sanghyuk Park, Sungrack Yun, Chang Dong Yoo |
ACCV (1) | 5 |
| 2012 | Sparsity Sharing Embedding for Face Verification
Hyunsin Park, Junyoung Chung, Youngook Song, Chang Dong Yoo |
ACCV (2) | 5 |
| 2012 | Phoneme Classification using Constrained Variational Gaussian Process Dynamical SystemabstractThis paper describes a new acoustic model based on variational Gaussian process dynamical system (VGPDS) for phoneme classification. The proposed model overcomes the limitations of the classical HMM in modeling the real speech data, by adopting a nonlinear and nonparametric model. In our model, the GP prior on the dynamics function enables representing the complex dynamic structure of speech, while the GP prior on the emission function successfully models the global dependency over the observations. Additionally, we introduce variance constraint to the original VGPDS for mitigating sparse approximation error of the kernel matrix. The effectiveness of the proposed model is demonstrated with extensive experimental results including parameter estimation, classification performance on the synthetic and benchmark datasets. Hyunsin Park, Sungrack Yun, Sanghyuk Park, Jongmin Kim 0006, Chang Dong Yoo |
NIPS | 5 |
| 2012 | Loss-Scaled Large-Margin Gaussian Mixture Models for Speech Emotion ClassificationabstractThis paper considers a learning framework for speech emotion classification using a discriminant function based on Gaussian mixture models (GMMs). The GMM parameter set is estimated by margin scaling with a loss function to reduce the risk of predicting emotions with high loss. Here, the loss function is defined as a function of a distance metric using the Watson and Tellegen's emotion model. Margin scaling is known to have good generalization ability and can be considered appropriate for emotion modeling where the parameter set is likely to be over-fitted to the training data set whose characteristics may differ from those of the testing data set. Our learning framework is formulated as a constrained optimization problem which is solved using semi-definite programming. Three tasks were evaluated: acted emotion classification, natural emotion classification, and cross database emotion classification. In each task, four loss functions were evaluated. In all experiments, results consistently show that margin scaling improves the classification accuracy over other learning frameworks based on the maximum-likelihood, maximum mutual information and max-margin framework without margin scaling. Experiment results also show that margin scaling substantially reduces the overall loss compared to the max-margin framework without margin scaling. Sungrack Yun, Chang Dong Yoo |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Variable grouping for energy minimizationabstractThis paper addresses the problem of efficiently solving large-scale energy minimization problems encountered in computer vision. We propose an energy-aware method for merging random variables to reduce the size of the energy to be minimized. The method examines the energy function to find groups of variables which are likely to take the same label in the minimum energy state and thus can be represented by a single random variable. We propose and evaluate a number of extremely efficient variable grouping strategies. Experimental results show that our methods result in a dramatic reduction in the computational cost and memory requirements (in some cases by a factor of one hundred) with almost no drop in the accuracy of the final result. Comparative evaluation with efficient super-pixel generation methods, which are commonly used in variable grouping, reveals that our methods are far superior both in terms of accuracy and running time. Taesup Kim, Sebastian Nowozin, Pushmeet Kohli, Chang Dong Yoo |
CVPR | 4 |
| 2011 | Learning a discriminative visual codebook using homonym schemeabstractThis paper studies a method for learning a discriminative visual codebook for various computer vision tasks such as image categorization and object recognition. The performance of various computer vision tasks depends on the construction of the code book which is a table of visual-words (i.e. codewords). This paper proposed a learning criterion for constructing a discriminative codebook, and it is solved by the homonym scheme which splits codeword regions by labels. A codebook is learned based on the proposed homonym scheme such that its histogram can be used to discriminate objects of different labels. The traditional codebook based on the k-means is compared against the learned codebook on two well-known datasets (Caltech 101, ETH-80) and a dataset we constructed using google images. We show that the learned codebook consistently outperforms the traditional codebook. SeungRyul Baek, Chang Dong Yoo, Sungrack Yun |
ICASSP | 2 |
| 2011 | A High Resolution Multiple Source Localization Based on Generalized Cumulant Structure (GCS) MatrixabstractThis paper considers a high-resolution multiple non-stationary and non-Gaussian source localization algorithm based on the proposed generalized cumulant structure (GCS) matrix that is constructed as a weighted sum of the second and fourth order cumulants of the sensor signals. The weight determines the rank and range space of the GCS matrix, and the range space of the GCS matrix should be same to the range space of the virtual array manifold matrix to estimate the true direction of arrival (DOA)s of the sources. To estimate the weight and the DOAs of sources, a rank constrained optimization problem is formulated. The optimal solution is computationally heavy, and for this reason a suboptimal solution is considered. With the weight set to an arbitrary value, singular value decomposition on the GCS matrix is performed to determine the singular matrix associated with the null space of the virtual array response matrix, and either this singular matrix or the singular matrix obtained using only the second order (SO) statistic is used to obtain the proposed spatial spectrum. Experimental results show that the proposed algorithm performs better than the recently proposed SO cumulant based algorithm for synthetic and real speech data. Jinho Choi 0002, Chang Dong Yoo |
INTERSPEECH | 2 |
| 2011 | Higher-Order Correlation Clustering for Image SegmentationabstractFor many of the state-of-the-art computer vision algorithms, image segmentation is an important preprocessing step. As such, several image segmentation algorithms have been proposed, however, with certain reservation due to high computational load and many hand-tuning parameters. Correlation clustering, a graph-partitioning algorithm often used in natural language processing and document clustering, has the potential to perform better than previously proposed image segmentation algorithms. We improve the basic correlation clustering formulation by taking into account higher-order cluster relationships. This improves clustering in the presence of local boundary ambiguities. We first apply the pairwise correlation clustering to image segmentation over a pairwise superpixel graph and then develop higher-order correlation clustering over a hypergraph that considers higher-order relations among superpixels. Fast inference is possible by linear programming relaxation, and also effective parameter learning framework by structured support vector machine is possible. Experimental results on various datasets show that the proposed higher-order correlation clustering outperforms other state-of-the-art image segmentation algorithms. Sungwoong Kim, Sebastian Nowozin, Pushmeet Kohli, Chang Dong Yoo |
NIPS | 4 |
| 2011 | Large Margin Discriminative Semi-Markov Model for Phonetic RecognitionabstractThis paper considers a large margin discriminative semi-Markov model (LMSMM) for phonetic recognition. The hidden Markov model (HMM) framework that is often used for phonetic recognition assumes only local statistical dependencies between adjacent observations, and it is used to predict a label for each observation without explicit phone segmentation. On the other hand, the semi-Markov model (SMM) framework allows simultaneous segmentation and labeling of sequential data based on a segment-based Markovian structure that assumes statistical dependencies among all the observations within a phone segment. For phonetic recognition which is inherently a joint segmentation and labeling problem, the SMM framework has the potential to perform better than the HMM framework at the expense of slight increase in computational complexity. The SMM framework considered in this paper is based on a non-probabilistic discriminant function that is linear in the joint feature map which attempts to capture long-range statistical dependencies among observations. The parameters of the discriminant function are estimated by a large margin learning framework for structured prediction. The parameter estimation problem in hand leads to an optimization problem with many margin constraints, and this constrained optimization problem is solved using a stochastic gradient descent algorithm. The proposed LMSMM outperformed the large margin discriminative HMM in the TIMIT phonetic recognition task. Sungwoong Kim, Sungrack Yun, Chang Dong Yoo |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Largemargin training of semi-Markov model for phonetic recognitionabstractThis paper considers a large margin training of semi-Markov model (SMM) for phonetic recognition. The SMM framework is better suited for phonetic recognition than the hidden Markov model (HMM) framework in that the SMM framework is capable of simultaneously segmenting the uttered speech into phones and labeling the segment-based features. In this paper, the SMM framework is used to define a discriminant function that is linear in the joint feature map which attempts to capture the long-range statistical dependencies within a segment and between adjacent segments of variable length. The parameters of the discriminant function are estimated by a large margin learning criterion for structured prediction. The parameter estimation problem, which is an optimization problem with many margin constraints, is solved by using a stochastic subgradient descent algorithm. The proposed large margin SMM outperforms the large margin HMM on the TIMIT corpus. Sungwoong Kim, Sungrack Yun, Chang Dong Yoo |
ICASSP | 3 |
| 2010 | Parametric emotional singing voice synthesisabstractThis paper describes an algorithm to control the expressed emotion of a synthesized song. Based on the database of various melodies sung neutrally with restricted set of words, hidden semi-Markov models (HSMMs) of notes ranging from E3 to G5 are constructed for synthesizing singing voice. Three steps are taken in the synthesis: (1) Pitch and duration are determined according to the notes indicated by the musical score; (2) Features are sampled from appropriate HSMMs with the duration set to the maximum probability; (3) Singing voice is synthesized by the mel-log spectrum approximation (MLSA) filter using the sampled features as parameters of the filter. Emotion of a synthesized song is controlled by varying the duration and the vibrato parameters according to the Thayer's mood model. Perception test is performed to evaluate the synthesized song. The results show that the algorithm can control the expressed emotion of a singing voice given a neutral singing voice database. Younsung Park, Sungrack Yun, Chang Dong Yoo |
ICASSP | 3 |
| 2010 | A maximum a posteriori sound source localization in reverberant and noisy conditionsabstractA maximum a posteriori sound source localization in reverberant and noisy conditions Jinho Choi 0002, Chang Dong Yoo |
INTERSPEECH | 2 |
| 2010 | Melody pitch estimation based on range estimation and candidate extraction using harmonic structure modelabstractThis paper proposes an algorithm to estimate the melody pitch line (the most dominant pitch sequence) of a given polyphonic audio based on melody range estimation and pitch candidate extraction using a harmonic structure model similar to that proposed by Goto. This paper defines melody pitch candidate as a list of pitch candidates that produces the best-fit harmonic models to the polyphonic audio. In many melody extraction algorithms proposed in the past, multiple-pitch extractor (MPE) is often performed for extracting melody pitch candidates; however, the MPE serves the purpose of estimating all pitches within a frame of a polyphonic audio and does not necessarily provide melody pitch candidates. The estimated weights of the harmonic structure model which must be obtained for extracting the pitch candidates are liable to octave error and strong low frequency interference, and therefore, certain refinement after the estimation must be performed. As a refinement, the algorithm measures the degree of harmonic fitness of each candidate. Furthermore, a melody pitch range is estimated to reduce false-positive pitch candidates. The melody pitch range is estimated based on the distribution of the best pitch candidates with long duration. Experimental results show that the proposed extraction algorithm performed better than many of the algorithms proposed in the past. Seokhwan Jo, Sihyun Joo, Chang Dong Yoo |
INTERSPEECH | 3 |
| 2010 | Psychoacoustically Constrained and Distortion Minimized Speech EnhancementabstractThis paper considers a psychoacoustically constrained and distortion minimized speech enhancement algorithm. Noise reduction, in general, leads to speech distortion, and a balanced tradeoff between noise reduction and speech distortion must be attained. A constrained optimization problem is set to reduce noise so that speech distortion is minimized while the sum of speech distortion and residual noise is kept below the masking threshold of the clean speech. Obtaining a solution to the optimization problem may be infeasible under certain conditions, and a slack variable is introduced to allow certain deviation from the constraint conditions. To estimate the power spectral density and also the masking threshold of clean speech, a speech model that assumes coexisting deterministic and stochastic components in speech is used. Experimental results show that the considered algorithm outperforms some of the more popular algorithms in terms of improvement in segmental signal-to-noise ratio (SegSNR), spectral distance (SD), modified Bark spectral distortion (MBSD), and mean opinion score (MOS). Seokhwan Jo, Chang Dong Yoo |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Distance Metric Learning for Content IdentificationabstractThis paper considers a distance metric learning (DML) algorithm for a fingerprinting system, which identifies a query content by finding the fingerprint in the database (DB) that measures the shortest distance to the query fingerprint. For a given training set consisting of original and distorted fingerprints, a distance metric equivalent to the lpnorm of the difference of two linearly projected fingerprints is learned by minimizing the false-positive rate (probability of perceptually dissimilar content to be identified as being similar) for a given false-negative rate (probability of perceptually similar content to be identified as being dissimilar). The learned metric can perform better than the often used lpdistance and improve the robustness against a set of unexpected distortions. In the experiments, the distance metric learned by the proposed algorithm performed better than those metrics learned by well-known DML algorithms for classification. Dalwon Jang, Chang Dong Yoo, Ton Kalker |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | Fingerprint matching based on distance metric learningabstractThis paper considers a method for learning a distance metric in a fingerprinting system which identifies a query content by measuring the distance between its fingerprint and a fingerprint stored in a database. A metric having a general form of the Mahalanobis distance is learned with the goal that the distance between fingerprints extracted from perceptually similar contents should be smaller than the distance between fingerprints extracted from perceptually dissimilar contents. The metric is learned by minimizing a cost function designed to achieve the goal. The cost function is convex, and the global minimum can be obtained using convex optimization. In our experiment, the distance metric learning is applied in an audio fingerprinting system, and it is experimentally shown that the learned distance metric improves the identification performance. Dalwon Jang, Chang Dong Yoo |
ICASSP | 2 |
| 2009 | Humming-based human verification and identificationabstractThis paper considers humming-based systems for human verification and identification. Humming of a target person is modeled as a Gaussian mixture model, and the matching score between a target model and humming is computed as the likelihood of humming given a target model. Verification is performed by comparing the matching score to the likelihood given a universal background model, and identification is performed by selecting the best-matched model. The verification and identification performances are evaluated using various acoustical features. The experimental results show that linear prediction cepstral coefficients and perceptually linear prediction coefficients are conducive to verification and identification, respectively. Minho Jin, Chang Dong Yoo |
ICASSP | 3 |
| 2009 | Psychoacoustically constrained and distortion minimized speech enhancement algorithmabstractA psychoacoustically constrained and distortion minimized speech enhancement algorithm is considered. In general, noise reduction leads to speech distortion, and thus, the goal of an enhancement algorithm should reduce noise and speech distortion so that both are inaudible. In this paper, a constrained optimization problem is formulated so that speech distortion is minimized while distortion that includes residual noise and speech distortion is kept below the masking threshold of the clean speech. Experimental results show that the algorithm considered in this paper outperforms some of the more popular algorithms in terms of improvement in segmental signal-to-noise ratio (SegSNR) and spectral distance (SD). Seokhwan Jo, Chang Dong Yoo |
ICASSP | 2 |
| 2009 | Speech emotion recognition via a max-margin framework incorporating a loss function based on the Watson and Tellegen's emotion modelabstractThis paper considers a method for speech emotion recognition by a max-margin framework incorporating a loss function based on a well-known model called theWatson and Tellegen's emotion model. Each emotion is modeled by a single-state hidden Markov model (HMM) that is trained by maximizing the minimum separation margin between emotions, and the margin is scaled by a loss function. The framework is optimized by the semi-definite programming. Experiments were performed to evaluate the framework using the Berlin database of emotional speech. The framework performed better than other conventional training criteria for HMM such as maximum likelihood estimation and maximum mutual information estimation. Sungrack Yun, Chang Dong Yoo |
ICASSP | 2 |
| 2009 | Robust Video Fingerprinting Based on Symmetric Pairwise BoostingabstractThis paper proposes a video fingerprinting method based on a novel binary fingerprint obtained using a feature selection algorithm called the symmetric pairwise boosting (SPB). The binary fingerprints are obtained by filtering and quantizing perceptually significant features extracted from an input video clip. The SPB algorithm, which is a generalization of the conventional asymmetric pairwise boosting (APB), selects appropriate filters and quantizers from a class of candidate filters and quantizers in such a way that perceptually similar and dissimilar pairs of video clips are correctly classified as matching and non-matching pairs, respectively. The binary form of the novel fingerprint makes it conducive to an efficient database search, and the experimental results show that the proposed method outperforms the APB-based video fingerprinting methods in terms of both robustness and discriminability. Sunil Lee, Chang Dong Yoo, Ton Kalker |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Pairwise boosted audio fingerprintabstractA novel binary audio fingerprint obtained by filtering and then quantizing the spectral centroids is proposed. A feature selection algorithm, coined pairwise boosting (PB), is used to determine the filters and quantizers by casting the fingerprinting problem of identifying a query audio clip into a binary classification problem. The PB algorithm selects the filters and quantizers which lead to accurate classification of matching and nonmatching audio pairs: a matching pair is an audio pair that should be classified as being identical, and a nonmatching pair is a pair that should be classified as being different. By iteratively reducing the classification error of both matching and nonmatching pairs, the PB algorithm improves both the robustness and discriminating ability. In our experiments, the proposed fingerprint outperformed previously reported binary fingerprints in terms of robustness and discriminating ability. In the experiment, we compared the performances of a number of distance measures. Dalwon Jang, Chang Dong Yoo, Sunil Lee, Sungwoong Kim, Ton Kalker |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | Quantum hashing for multimediaabstractIn this paper, a novel multimedia identification system based on quantum hashing is considered. Many traditional systems are based on binary hash which is obtained by encoding intermediate hash extracted from multimedia content. In the system considered, the intermediate hash values extracted from a query are encoded into quantum hash values by incorporating uncertainty in the binary hash values. For this, the intermediate hash difference between the query and its true-underlying content is considered as a random process. Then, the uncertainty is represented by the probability density estimate of the intermediate hash difference. The quantum hashing system is evaluated using both audio and video databases, and with marginal increment in computational cost, the quantum hashing system is shown to be more robust against various distortions than the binary hashing system using the same intermediate hash values. Minho Jin, Chang Dong Yoo |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2008 | Robust video fingerprinting based on affine covariant regionsabstractThis paper proposes a robust video fingerprinting method based on affine covariant regions. In video fingerprinting, a video clip is identified using short feature vectors referred to as fingerprints. In the proposed method, local fingerprints based on the centroid of gradient orientations are extracted from affine covariant regions detected in each key-frame. For the region detection, the maximally stable extremal region (MSER) detector which is considered to have high repeatability and low complexity is used. For reliable matching of the local fingerprints, only spatio-temporally consistent matches are taken into account. The experimental results show that the proposed method is robust against both geometric and non-geometric transformations. Sunil Lee, Chang Dong Yoo |
ICASSP | 2 |
| 2008 | Robust video fingerprinting based on 2D-OPCA of affine covariant regionsabstractThis paper proposes a robust video fingerprinting method based on 2-Dimensional Oriented Principal Component Analysis (2D-OPCA) of affine covariant regions. The goal of video fingerprinting is to identify a video clip using perceptual features called fingerprints. In the proposed method, to achieve the robustness against geometric transformations, fingerprints are extracted from local regions co-variant with a class of affine transformations. The detected affine covariant regions are normalized geometrically and photometrically, and local fingerprints are extracted by applying a novel discriminant analysis algorithm, 2D-OPCA to the normalized regions. For the reliable matching of local fingerprints, only spatio-temporally consistent matches are taken into account. The experimental results show that the proposed method is robust against both geometric and non-geometric transformations. Sunil Lee, Chang Dong Yoo |
ICIP | 2 |
| 2008 | Music genre classification using novel features and a weighted voting methodabstractThis paper proposes a novel music genre classification system based on two novel features and a weighted voting. The proposed features, modulation spectral flatness measure (MSFM) and modulation spectral crest measure (MSCM), represent the time-varying behavior of a music and indicate the beat strength. The weighted voting method determines the music genre by summarizing the classification results of consecutive time segments. Experimental results show that the proposed features give more accurate classification results when combined with traditional features than the octave-based modulation spectral contrast (OMSC) does in spite of short feature vector and that the weighted voting is more effective than statistical method and majority voting. Dalwon Jang, Minho Jin, Chang Dong Yoo |
ICME | 3 |
| 2008 | Development of a simple free viewpoint video systemabstractA simple free viewpoint video system which is able not only to display user-specified views at arbitrary angle but also to efficiently stream the necessary video over a network is described. Virtual view synthesis based on projection theory was implemented to obtain intermediate views between multiple cameras. The depth information, which is required for the virtual view synthesis, was obtained using the segmentation-based stereo matching algorithm. For realtime rendering, our system was optimized using Single Instructions, Multiple Data (SIMD) technology. For efficient streaming, a novel method of combining the video with the depth information is proposed. Seokhwan Jo, Yoonseob Kim, Chang Dong Yoo |
ICME | 4 |
| 2008 | Robust Video Fingerprinting for Content-Based Video IdentificationabstractVideo fingerprints are feature vectors that uniquely characterize one video clip from another. The goal of video fingerprinting is to identify a given video query in a database (DB) by measuring the distance between the query fingerprint and the fingerprints in the DB. The performance of a video fingerprinting system, which is usually measured in terms of pairwise independence and robustness, is directly related to the fingerprint that the system uses. In this paper, a novel video fingerprinting method based on the centroid of gradient orientations is proposed. The centroid of gradient orientations is chosen due to its pairwise independence and robustness against common video processing steps that include lossy compression, resizing, frame rate change, etc. A threshold used to reliably determine a fingerprint match is theoretically derived by modeling the proposed fingerprint as a stationary ergodic process, and the validity of the model is experimentally verified. The performance of the proposed fingerprint is experimentally evaluated and compared with that of other widely-used features. The experimental results show that the proposed fingerprint outperforms the considered features in the context of video fingerprinting. Sunil Lee, Chang Dong Yoo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Speech Enhancement Based on the Decomposition of Speech Into Deterministic and Stochastic Components and Psychoacoustic ModelabstractA novel speech enhancement algorithm based on both a decomposition of speech into coexisting deterministic and stochastic components and a psychoacoustic model is proposed. Noisy speech is first decomposed into deterministic and stochastic components, and then each component is enhanced preserving its individual characteristics. A psychoacoustic model is taken into account when enhancing the stochastic component which usually has much lower energy than the deterministic component. Simulation results show that the proposed algorithm performs better than some of the more popular algorithms in terms of segmental signal-to-noise ratio (SNR) and speech recognition rate. Seokhwan Jo, Chang Dong Yoo |
ICASSP (4) | 2 |
| 2007 | A Novel Adaptive Crosstalk Cancellation using Psychoacoustic Model for 3D AudioabstractIn rendering a virtual sound over two loudspeakers, adaptive inverse filtering is required for crosstalk cancellation. Although various adaptive algorithms have been proposed for crosstalk cancellation, few have been effective especially in time-varying environments where fast convergence rate is required. Until now, the least-mean square (LMS) algorithm known for its simplicity and robustness has been the predominant algorithm used, but its convergence rate is considered slow for colored inputs. In this paper, human perceptual characteristics which have never been incorporated in an LMS algorithm is introduced. In our experiment, the proposed algorithm achieved higher perceptual accuracy and faster convergence rate than the conventional LMS algorithm. Jun Jun Seong Kim, Sang-Gyun Kim, Chang Dong Yoo |
ICASSP (1) | 3 |
| 2007 | Boosted Binary Audio Fingerprint Based on Spectral Subband MomentsabstractAn audio fingerprinting system identifies an audio based on a unique feature vector called the audio fingerprint. The performance of an audio fingerprinting system is directly related to the fingerprint that the system uses. To reduce both the DB size and the DB search time, binary fingerprints are often used. However converting a real-valued fingerprint into a binary fingerprint results in loss of information and leads to severe degradation in performance. In this paper, an algorithm known as boosting is used as a binary conversion method which minimizes the degradation. The experimental results showed that the proposed binary audio fingerprint obtained by boosting the spectral subband moments outperformed some of the state-of-the-art binary audio fingerprints in the context of both robustness and pair-wise independence (reliability). Sungwoong Kim, Chang Dong Yoo |
ICASSP (1) | 2 |
| 2007 | Temporal Dynamics for Spectral Sub-Band Centroid Audio FingerprintsabstractMotivated by the effectual use of temporal information in speech recognition, we investigate the effectiveness of the temporal dynamics of the spectral sub-band centroid (SSC) fingerprints for audio fingerprinting. The SSC, which is known to be a robust audio fingerprint against various distortions, does not involve any temporal dynamics. Here, the temporal dynamics are defined as the difference between two neighboring SSCs. The robustness of the temporal dynamics against various distortions were compared to that of SSCs. The system using temporal dynamics showed similar performance to that using SSCs in various distortions except for time-scale modification and linear-speed change. This is to be expected since these distortions change the time correlation of an audio. In our experiment, the concatenation of SSCs and the temporal dynamics outperformed each of the individual fingerprints. This suggests that the SSCs and the temporal dynamics provide information which is mutually supplementary. Minho Jin, Chang Dong Yoo |
ICME | 2 |
| 2007 | A Syllable Lattice Approach to Speaker VerificationabstractThis paper proposes a syllable-lattice-based speaker verification algorithm for Mandarin Chinese input. For each speech utterance, a syllable lattice is generated with a speaker-independent large-vocabulary continuous speech recognition system in free syllable decoding. The verification decision is made based upon the likelihood ratio between a target-speaker model and a speaker-independent background model, computed on the decoded syllable lattice. The likelihood function is calculated efficiently in a forward algorithm by considering all paths in the lattice. The proposed algorithm was evaluated using a Mandarin Chinese database, where 1832 true and 26 250 impostor trials were recorded by 19 target speakers and 180 impostors. The average duration of each trial is 2 s long without silence. The target-speaker model was adapted from the speaker-independent background model using enrollment data of two minutes with silence. The proposed algorithm achieved an equal-error rate of 0.857% which is better than 1.21% of the hidden Markov model-based speaker verification algorithm without using syllable lattices. The equal-error rate was further reduced to 0.617% by incorporating the Goussian mixture model-universal background model algorithm with 2048 Gaussian kernels whose equal error rate is 0.990%. Minho Jin, Frank K. Soong, Chang Dong Yoo |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Reversible Image Watermarking Based on Integer-to-Integer Wavelet TransformabstractThis paper proposes a high capacity reversible image watermarking scheme based on integer-to-integer wavelet transforms. The proposed scheme divides an input image into nonoverlapping blocks and embeds a watermark into the high-frequency wavelet coefficients of each block. The conditions to avoid both underflow and overflow in the spatial domain are derived for an arbitrary wavelet and block size. The payload to be embedded includes not only messages but also side information used to reconstruct the exact original image. To minimize the mean-squared distortion between the original and the watermarked images given a payload, the watermark is adaptively embedded into the image. The experimental results show that the proposed scheme achieves higher embedding capacity while maintaining distortion at a lower level than the existing reversible watermarking schemes. Sunil Lee, Chang Dong Yoo, Ton Kalker |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2006 | A Novel Embedding Method For An Anti-Collusion Fingerprinting By Embedding Both A Code And An Orthogonal FingerprintabstractIn this paper, a fingerprint embedding method better-suited for the AND anti-collusion code (AND-ACC) is proposed. The proposed method embeds both a code and an orthogonal fingerprint using different basis vectors depending on the bit. Although the detection for the embedding method is complex, the performance of the fingerprinting system using proposed embedding method with the AND-ACC against average attack is improved compared with the AND-ACC fingerprinting scheme using code modulation embedding method. The system using the proposed embedding method is robust against the linear combination collusion attack (LCCA) whereas the system using the code modulation is not. Dalwon Jang, Chang Dong Yoo |
ICASSP (5) | 2 |
| 2006 | Syllable Lattice Based Re-Scoring For Speaker VerificationabstractThe Gaussian mixture based GMM-UBM approaches have shown good performance in speaker verification without using contextual information. In this paper, we exploit the information provided in the arcs of a decoded syllable lattice for speaker verification. The forward algorithm is used to summarize this information in the syllable lattice instead of the best decoded string. The performance is evaluated on a Mandarin Chinese database. With two minutes of target speaker's enrollment data, the proposed algorithm shows 1.03% of equal-error rate for short input utterances with an average duration of two seconds. By combining with the GMM-UBM, the system shows a 0.74% of equal-error rate. Minho Jin, Frank K. Soong, Chang Dong Yoo |
ICASSP (1) | 3 |
| 2006 | Video Fingerprinting Based on Centroids of Gradient OrientationsabstractFingerprints are feature vectors that can uniquely characterize the video signal. The goal of a video fingerprinting system is to judge whether two videos have the same contents by measuring distance between fingerprints extracted from the videos. In this paper, a novel video fingerprinting method based on the centroids of gradient orientations is proposed. The centroid of gradient orientations is chosen due to its reliability and robustness against common video processing steps. A threshold used to reliably determine a fingerprint match is theoretically derived, and its validity is experimentally verified. The experimental results show that the proposed fingerprint is not only pairwise independent but also robust against common video processing steps. Sunil Lee, Chang Dong Yoo |
ICASSP (2) | 2 |
| 2006 | Audio fingerprinting based on normalized spectral subband momentsabstractThe performance of a fingerprinting system, which is often measured in terms of reliability and robustness, is directly related to the features that the system uses. In this letter, we present a new audio-fingerprinting method based on the normalized spectral subband moments. A threshold used to reliably determine a fingerprint match is obtained by modeling the features as a stationary process. The robustness of the normalized moments was evaluated experimentally and compared with that of the spectral flatness measure. Among the considered subband features, the first-order normalized moment showed the best performance for fingerprinting. Jin S. Seo, Minho Jin, Sunil Lee, Dalwon Jang, Seungjae Lee 0001, Chang Dong Yoo |
IEEE Signal Process. Lett. | 6 |
| 2005 | An SVD-Based Watermarking Method for Image Content Authentication with Improved SecurityabstractFor image content authentication, a secure watermarking method using quantization-based embedding on the largest singular value (SV) is proposed. The block-wise quantization-based embedding can be vulnerable to vector quantization (VQ) attack and attacks associated with histogram analysis. To overcome these security problems, the proposed method places interdependency among image blocks and dithers the quantized value. By adjusting the threshold of the detector, a trade-off between the robustness to JPEG compression and the probability of misdetection can be made. The proposed method can detect a tampered area with high sensitivity. This is confirmed by experimental results and security analysis. Seungjae Lee 0001, Dalwon Jang, Chang Dong Yoo |
ICASSP (2) | 3 |
| 2005 | Audio fingerprinting based on normalized spectral subband centroidsabstractFor multimedia fingerprinting, it is crucial to extract relevant features that allow direct access to the distinguishing characteristics of a multimedia object. Features used for fingerprinting directly relate to the performance of the entire fingerprinting system. The paper proposes a novel audio fingerprinting method based on normalized spectral subband centroids. The spectral subband centroid is selected due to its resilience against equalization, compression, and noise addition. Both reliability and robustness issues in the fingerprinting system are addressed. Experimental results show that the proposed method is not only reliable, but also robust against various audio processing steps, including MP3 compression, equalization, random start, time-scale modification, and linear speed change. Jin S. Seo, Minho Jin, Sunil Lee, Dalwon Jang, Seungjae Lee 0001, Chang Dong Yoo |
ICASSP (3) | 6 |
| 2004 | Hybrid utterance verification based on n-best models and model derived from kulback-leibler divergenceabstractIn this paper, utterance verification based on hybrid scores obtained from three pairs of models is investigated. The three models considered are the on-line garbage model, the antiword function model and a model derived using Kullback-Leibler divergence. The performance of utterance verification algorithm using hypothesis testing depends on the accuracy of the estimate of the alternative hypothesis. The three models offer different perspectives in the probability estimation of the alternate hypothesis. Performance comparison between hybrid scores using different model pairs is made. In addition, performance improvement over conventional algorithm is experimentally verified. Minho Jin, Gyucheol Jang, Sungrack Yun, Chang Dong Yoo |
INTERSPEECH | 4 |
| 2004 | Blind separation of speech and sub-Gaussian signals in underdetermined caseabstractConventional blind source separation (BSS) algorithms are applicable when the number of sources equals to that of observations; however, they are inapplicable when the number of sources is larger than that of observations. Most underdetermined BSS algorithms have been developed based on an assumption that all sources have sparse distributions. These algorithms are applicable to separate speech signals with super-Gaussian distribution in the underdetermined case. However, they fail to separate the underdetermined mixtures of speech signals and sub-Gaussian signals. In this paper, a novel method for separating the underdetermined mixtures of sources with both superand sub-Gaussian distributions is proposed. In the proposed method, underdetermined BSS problem is converted to conventional BSS problem by generating hidden observations so that the probability of estimated sources is maximized. Simulation results show that the proposed method can separate the underdetermined mixtures of speech signals and sub-Gaussian signals. Sang-Gyun Kim, Chang Dong Yoo |
INTERSPEECH | 2 |
| 2004 | Localized image watermarking based on feature points of scale-space representation
Jin S. Seo, Chang Dong Yoo |
Pattern Recognit. | 2 |
| 2004 | A robust image fingerprinting system using the Radon transform
Jin S. Seo, Jaap Haitsma, Ton Kalker, Chang Dong Yoo |
Signal Process. Image Commun. | 4 |
| 2003 | Improvements in speaker adaptation using weighted trainingabstractRegardless of the distribution of the adaptation data in the testing environment, model-based adaptation methods that have so far been reported in the literature incorporate the adaptation data undiscriminately in reducing the mismatch between the training and testing environments. When the amount of data is small and the parameter tying is extensive, adaptation based on outlier data can be detrimental to the performance of the recognizer. The distribution of the adaptation data plays a critical role on the adaptation performance. In order to maximally improve the recognition rate in the testing environment using only a small amount of adaptation data, supervised weighted training is applied to the structural maximum a posterior (SMAP) algorithm. We evaluate the performance of the proposed weighted SMAP (WSMAP) and SMAP on TIDIGITS corpus. The proposed WSMAP has been found to perform better for a small amount of data. The general idea of incorporating the distribution of the adaptation data is applicable to other adaptation algorithms. Gyucheol Jang, Sooyoung Woo, Minho Jin, Chang Dong Yoo |
ICASSP (1) | 4 |
| 2003 | The incorporation of masking threshold to subspace speech enhancementabstractA subspace enhancement algorithm using the masking property is proposed. The proposed algorithm minimizes the signal distortion while constraining the energy of the residual noise below the human psychoacoustic masking threshold. This requires simple transformation of the masking threshold from the Fourier domain to the Karhunen-Loe/spl grave/ve (KL) domain. The proposed method incorporates subband whitening filters in the KL domain in order to deal with colored noise. Performance test results show that the proposed algorithm is superior to spectral subtraction and the subspace method suggested by Y. Ephraim and H.L. Van Trees (see IEEE Trans. Speech and Audio Processing, vol.3, no.4, p.251-66, 1995). Jong Uk Kim, Sang-Gyun Kim, Chang Dong Yoo |
ICASSP (1) | 3 |
| 2003 | A novel transcoding algorithm for AMR and EVRC speech codecs via direct parameter transformationabstractA novel transcoding algorithm for the adaptive multi rate (AMR) codec and the enhanced variable rate codec (EVRC) is proposed. In contrast to the conventional tandem transcoding algorithm, the proposed algorithm transcodes the parameters of one codec to the other without synthesizing the speech. The proposed algorithm decodes the parameters of source codec from the input bitstream, and based on frame classification and mode decision, it appropriately transforms the parameters of source codec to those of the target codec in the parametric domain. Finally, the transformed parameters are encoded into a bitstream that is decodable by the target codec. The parameters transcoded by the proposed algorithm are line-spectral pair (LSP), pitch delay, fixed codevector, codebook gains, and frame energy. Evaluation results show that while reducing both the computational complexity and delay by 50%, the proposed algorithm produces speech quality equivalent to that of produced by the tandem transcoding algorithm. The general idea is not restricted to the AMR and EVRC but is applicable to various other code-excited linear prediction (CELP) based codecs. Sunil Lee, Seongho Seo, Dalwon Jang, Chang Dong Yoo |
ICASSP (2) | 4 |
| 2003 | Affine transform resilient image fingerprintingabstractAffine transformations are a well-known robustness issue in many multimedia fingerprinting systems. Since it is quite easy with modem computers to apply affine transformations to audio, image and video content, there is an obvious necessity for affine trans-formation resilient fingerprinting. In this paper we present a new method for affine transformation resilient fingerprints that is based upon the auto-correlation of the Radon transform, the log map-ping and the Fourier transform. Besides robustness, we also ad-dress other issues such as security, database search efficiency and independence with perceptually different inputs. Experimental re-sults show that the proposed fingerprints are highly robust to affine transformations. 1. Jin S. Seo, Jaap Haitsma, Ton Kalker, Chang Dong Yoo |
ICASSP (3) | 4 |
| 2003 | Speaker adaptation based on confidence-weighted trainingabstractThis paper presents a novel method to enhance the performance of traditional speaker adaptation algorithm using discriminative adaptation procedure based on a novel confidence measure and non-linear weighting. Regardless of the distribution of the adaptation data, traditional model adaptation methods incorporate the adaptation data undiscriminatingly. When the data size is small and the parameter tying is extensive, adaptation based on outliers can be detrimental. A way to discriminate the contribution of each data in the adaptation is to incorporate a confidence measure based on likelihood. We evaluate and compare the performances of the proposed weighted SMAP (WSMAP) which controls the contribution of each data by sigmoid weighting using a novel confidence measure. The effectiveness of the proposed algorithm is experimentally verified by adapting native speaker models to nonnative speaker environment using TIDIGIT. 1. Gyucheol Jang, Minho Jin, Chang Dong Yoo |
INTERSPEECH | 3 |
| 2003 | A novel rate selection algorithm for transcoding CELP-type codec and SMVabstractIn this paper, we propose an efficient rate selection algorithm that can be used to transcode speech encoded by any code excited linear prediction (CELP)-type codec into a format compatible with selectable mode vocoder (SMV) via direct parameter transformation. The proposed algorithm performs rate selection using the CELP parameters. Simulation results show that while maintaining similar overall bit-rate compared to the rate selection algorithm of SMV, the proposed algorithm requires less computational load than that of SMV and does not degrade the quality of the transcoded speech. Dalwon Jang, Seongho Seo, Sunil Lee, Chang Dong Yoo |
INTERSPEECH | 4 |
| 2003 | A robust and sensitive word boundary decision algorithmabstractA robust and sensitive word boundary decision algorithm for automatic speech recognition (ASR) system is proposed. The algorithm uses a time-frequency feature to improve both robustness and sensitivity. The time-frequency features are passed through a bank of moving average filters for temporary decision of word boundary in each band. The decision results of each band are then passed through a median filter for the final decision. The adoption of time-frequency feature improves the sensitivity, while the median filtering improves the robustness. Proposed algorithm uses an adaptive threshold based on the signal-to-noise ratio (SNR) in each band which further improves the decision performance. Experimental result shows that the proposed algorithm outperforms the Q.Li et al’s robust algorithm. Jong Uk Kim, Sang-Gyun Kim, Chang Dong Yoo |
INTERSPEECH | 3 |
| 2003 | Accuracy improved double-talk detector based on state transition diagramabstractA double-talk detector (DTD) is generally used with an acoustic echo canceller (AEC) in pinpointing the region where far-end and near-end signal coexist. This region is called double-talk and during this region AEC usually freezes the adaptation. Decision variable used in DTD has a relatively longer transient time going from double-talk to single-talk than time going in opposite direction. Therefore, using a single threshold to pinpoint the location of double-talk region can be difficult. In this paper, a DTD based on a novel state transition diagram and a decision variable which requires minimal computational overhead is proposed to improve the accuracy of pinpointing the location. The use of different thresholds according to the state helps the DTD locate double-talk region more accurately. The proposed DTD algorithm is evaluated by obtaining a receiver operating characteristic (ROC) and is compared to that of Cho’s DTD. Sang-Gyun Kim, Jong Uk Kim, Chang Dong Yoo |
INTERSPEECH | 3 |
| 2003 | A novel transcoding algorithm for SMV and g.723.1 speech coders via direct parameter transformationabstractIn this paper, a novel transcoding algorithm for the Selectable Mode Vocoder (SMV) and the G.723.1 speech coder is proposed. In contrast to the conventional tandem transcoding algorithm, the proposed algorithm converts the parameters of one coder to the other without going through the decoding and encoding process. The proposed algorithm is composed of four parts: the parameter decoding, Line Spectral Pair (LSP) conversion, pitch period conversion and rate selection. The evaluation results show that the proposed algorithm achieves equivalent speech quality to that of tandem transcoding with reduced computational complexity and delay. Seongho Seo, Dalwon Jang, Sunil Lee, Chang Dong Yoo |
INTERSPEECH | 4 |
| 2002 | Acoustic echo cancellation based on m-channel IIR cosine-modulated filter bank
Sang-Gyun Kim, Chang Dong Yoo |
INTERSPEECH | 2 |
| 2002 | Subspace speech enhancement using subband whitening filterabstractA novel subspace approach for speech enhancement using a subband whitening filter is proposed. Previous subspace approaches for enhancement either assumed white noise or used a fullband pre-whitening filter before enhancement for colored noise. The previous approaches were successful only in reducing the upper bound of the signal distortion while reducing noise. However, the proposed method minimizes the overall signal distortion while reducing noise. Experimental results show that the proposed method attains higher segmental signal-to-noise ratio (seg SNR) than that attained by Ephraim et al. and also by the Wiener filter algorithm. In addition, the proposed algorithm requires less computational load than previous subspace approaches. Jong Uk Kim, Chang Dong Yoo |
INTERSPEECH | 2 |
| 2001 | Speech/noise-dominant decision for speech enhancementabstractA novel method to reduce additive non-stationary noise is proposed. The proposed method requires neither the statistical assumption about noise nor the estimate of the noise statistics from any pause regions. The enhancement is performed on a band-by-band basis for each time frame. Based on both the decision on whether a particular band in a frame is speech or noise dominant and the masking property of the human auditory system, an appropriate amount of noise is reduced using modified spectral subtraction. The proposed method was tested on various noisy conditions - car noise, F16 noise, white Gaussian noise, pink noise, tank noise and babble noise. On the basis of comparing segmental SNR with spectral subtraction proposed by Boll with pause detection for estimating noise, and visually inspecting the enhanced spectrograms and listening to the enhanced speech, the proposed method was found to effectively reduce various noise while minimizing distortion to speech. Sukhyun Yoon, Chang Dong Yoo |
INTERSPEECH | 2 |
| 1999 | Utilizing interband acoustical information for modeling stationary time-frequency regions of noisy speechabstractA novel enhancement system is developed that exploits the properties of stationary regions localized in both time and frequency. This system selects stationary time-frequency (TF) regions and adaptively enhances each region according to its local signal-to-noise ratio (LSNR) while utilizing both the acoustical knowledge of speech and the masking properties of the human auditory system. Each region is enhanced for maximum noise reduction while minimizing distortion. This paper evaluates the proposed system through informal listening tests and some objective measures. Chang Dong Yoo |
ICASSP | 1 |
| 1996 | Selective all-pole modeling of degraded speech using M-band decompositionabstractThis paper describes a speech enhancement system which exploits both time- and frequency-localized behavior. The local characteristics are obtained from stationary regions selected by M-band decomposition with an adaptive analysis window. The spectrum of each selected region is estimated with an all-pole model. In order to model only the spectral region of interest, selective linear prediction (SLP) is used. By modeling the local spectrum, either independently or dependently, with respect to other enhanced spectral regions, and adjusting the model order to the local characteristics, a balanced-tradeoff between noise reduction and speech distortion can be achieved. Chang Dong Yoo |
ICASSP | 1 |
| 1995 | Speech enhancement based on the generalized dual excitation model with adaptive analysis windowabstractIn this paper, we describe a generalized dual excitation (GDE) speech model that is more accurate in its characterization than the dual excitation (DE) model in that it takes into account pitch variations. This model, together with an analysis window whose length varies adaptively according to the changing characteristics of speech, forms the backbone of a new speech enhancement system. Informal comparisons of the GDE system with the traditional systems have shown a clear preference for the former. Chang Dong Yoo, Jae S. Lim |
ICASSP | 1 |