Jung-Woo Ha 0001

dblp:66/867-1 · also JungWoo Ha 0001, Jungwoo Ha 0001 · DBLP profile ↗
← Back
59ranked-venue papers
6as first author
32since 2021 · last 2025
0000-0002-7400-7681ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 5 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Deformable Graph Transformer
abstract
Transformer-based models have recently shown success in representation learning on graph-structured data beyond natural language processing and computer vision. However, the success is limited to small-scale graphs due to the drawbacks of full dot-product attention on graphs such as the quadratic complexity with respect to the number of nodes and message aggregation from enormous irrelevant nodes. To address these issues, we propose Deformable Graph Transformer (DGT) that performs sparse attention via dynamically selected relevant nodes for efficiently handling large-scale graphs with a linear complexity in the number of nodes. Specifically, our framework first constructs multiple node sequences with various criteria to consider both structural and semantic proximity. Then, combining with our learnable Katz Positional Encodings, the sparse attention is applied to the node sequences for learning node representations with a significantly reduced computational cost. Extensive experiments demonstrate that our DGT achieves superior performance on 7 graph benchmark datasets with 2.5 $\sim$∼ 449 times less computational cost compared to transformer-based graph models with full attention.
Jinyoung Park 0005, Seongjun Yun, Hyeon-Jin Park, Jaewoo Kang, Jisu Jeong, Jung-Woo Ha 0001, Hyunwoo J. Kim
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs
abstract
Large language models (LLMs) have made significant advancements in various natural language processing tasks, including question answering (QA) tasks. While incorporating new information with the retrieval of relevant passages is a promising way to improve QA with LLMs, the existing methods often require additional fine-tuning which becomes infeasible with recent LLMs. Augmenting retrieved passages via prompting has the potential to address this limitation, but this direction has been limitedly explored. To this end, we design a simple yet effective framework to enhance open-domain QA (ODQA) with LLMs, based on the summarized retrieval (SuRe). SuRe helps LLMs predict more accurate answers for a given question, which are well-supported by the summarized retrieval that could be viewed as an explicit rationale extracted from the retrieved passages. Specifically, SuRe first constructs summaries of the retrieved passages for each of the multiple answer candidates. Then, SuRe confirms the most plausible answer from the candidate set by evaluating the validity and ranking of the generated summaries. Experimental results on diverse ODQA benchmarks demonstrate the superiority of SuRe, with improvements of up to 4.6\% in exact match (EM) and 4.0\% in F1 score over standard prompting approaches. SuRe also can be integrated with a broad range of retrieval methods and LLMs. Finally, the generated summaries from SuRe show additional advantages to measure the importance of retrieved passages and serve as more preferred rationales by models and humans.
Jaehyung Kim 0001, Jaehyun Nam, Sangwoo Mo, Jongjin Park, Sang-Woo Lee 0001, Minjoon Seo, Jung-Woo Ha 0001, Jinwoo Shin
ICLR7
2024 Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
abstract
Large language models (LLMs) have shown remarkable performance in various natural language processing tasks. However, a primary constraint they face is the context limit, i.e., the maximum number of tokens they can process. Previous works have explored architectural changes and modifications in positional encoding to relax the constraint, but they often require expensive training or do not address the computational demands of self-attention. In this paper, we present Hierarchical cOntext MERging (HOMER), a new training-free scheme designed to overcome the limitations. HOMER uses a divide-and-conquer algorithm, dividing long inputs into manageable chunks. Each chunk is then processed collectively, employing a hierarchical strategy that merges adjacent chunks at progressive transformer layers. A token reduction technique precedes each merging, ensuring memory usage efficiency. We also propose an optimized computational order reducing the memory requirement to logarithmically scale with respect to input length, making it especially favorable for environments with tight memory restrictions. Our experiments demonstrate the proposed method's superior performance and memory efficiency, enabling the broader use of LLMs in contexts requiring extended context. Code is available at https://github.com/alinlab/HOMER.
Woomin Song, Seunghyuk Oh, Sangwoo Mo, Jaehyung Kim 0001, Sukmin Yun, Jung-Woo Ha 0001, Jinwoo Shin
ICLR6
2023 Scaling Law for Recommendation Models: Towards General-Purpose User Representations
abstract
Recent advancement of large-scale pretrained models such as BERT, GPT-3, CLIP, and Gopher, has shown astonishing achievements across various task domains. Unlike vision recognition and language models, studies on general-purpose user representation at scale still remain underexplored. Here we explore the possibility of general-purpose user representation learning by training a universal user encoder at large scales. We demonstrate that the scaling law is present in user representation learning areas, where the training error scales as a power-law with the amount of computation. Our Contrastive Learning User Encoder (CLUE), optimizes task-agnostic objectives, and the resulting user embeddings stretch our expectation of what is possible to do in various downstream tasks. CLUE also shows great transferability to other domains and companies, as performances on an online experiment shows significant improvements in Click-Through-Rate (CTR). Furthermore, we also investigate how the model performance is influenced by the scale factors, such as training data size, model capacity, sequence length, and batch size. Finally, we discuss the broader impacts of CLUE in general.
Kyuyong Shin, Hanock Kwak, Su Young Kim, Max Nihlén Ramström, Jisu Jeong, Jung-Woo Ha 0001
AAAI6
2023 SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine Collaboration
abstract
Hwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi, Byoungpil Kim, Gunhee Kim, Eun-Ju Lee, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hwaran Lee, Seokhee Hong 0002, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi 0001, Byoung Pil Kim, Gunhee Kim, Eun-Ju Lee 0001, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha 0001
ACL (1)13
2023 Query-Efficient Black-Box Red Teaming via Bayesian Optimization
abstract
Deokjae Lee, JunYeong Lee, Jung-Woo Ha, Jin-Hwa Kim, Sang-Woo Lee, Hwaran Lee, Hyun Oh Song. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Deokjae Lee, Jung-Woo Ha 0001, Jin-Hwa Kim, Sang-Woo Lee 0001, Hwaran Lee, Hyun Oh Song
ACL (1)3
2023 Pivotal Role of Language Modeling in Recommender Systems: Enriching Task-specific and Task-agnostic Representation Learning
abstract
Kyuyong Shin, Hanock Kwak, Wonjae Kim, Jisu Jeong, Seungjae Jung, Kyungmin Kim, Jung-Woo Ha, Sang-Woo Lee. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Kyuyong Shin, Hanock Kwak, Wonjae Kim, Jisu Jeong, Seungjae Jung, Jung-Woo Ha 0001, Sang-Woo Lee 0001
ACL (1)7
2023 Dense Text-to-Image Generation with Attention Modulation
abstract
Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image region. To address this, we propose DenseDiffusion, a training-free method that adapts a pre-trained text-to-image model to handle such dense captions while offering control over the scene layout. We first analyze the relationship between generated images’ layouts and the pre-trained model’s intermediate attention maps. Next, we develop an attention modulation method that guides objects to appear in specific regions according to layout guidance. Without requiring additional fine-tuning or datasets, we improve image generation performance given dense captions regarding both automatic and human evaluation scores. In addition, we achieve similar-quality visual results with models specifically trained with layout conditions. Code and data are available at https://github.com/naver-ai/DenseDiffusion.
Yunji Kim, Jiyoung Lee 0005, Jin-Hwa Kim, Jung-Woo Ha 0001, Jun-Yan Zhu
ICCV4
2023 Text-Conditioned Sampling Framework for Text-to-Image Generation with Masked Generative Models
abstract
Token-based masked generative models are gaining popularity for their fast inference time with parallel decoding. While recent token-based approaches achieve competitive performance to diffusion-based models, their generation performance is still suboptimal as they sample multiple tokens simultaneously without considering the dependence among them. We empirically investigate this problem and propose a learnable sampling model, Text-Conditioned Token Selection (TCTS), to select optimal tokens via localized supervision with text information. TCTS improves not only the image quality but also the semantic alignment of the generated images with the given texts. To further improve the image quality, we introduce a cohesive sampling strategy, Frequency Adaptive Sampling (FAS), to each group of tokens divided according to the self-attention maps. We validate the efficacy of TCTS combined with FAS with various generative tasks, demonstrating that it significantly outperforms the baselines in image-text alignment and image quality. Our text-conditioned sampling framework further reduces the original inference time by more than 50 % without modifying the original generative model.
Jaewoong Lee, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon, Yunji Kim, Jin-Hwa Kim, Jung-Woo Ha 0001, Sung Ju Hwang
ICCV7
2023 Rarity Score : A New Metric to Evaluate the Uncommonness of Synthesized Images
Jiyeon Han 0001, Hwanil Choi, Yunjey Choi, Jung-Woo Ha 0001, Jaesik Choi
ICLR5
2023 Online Boundary-Free Continual Learning by Scheduled Data Prior
Hyunseo Koh, Minhyuk Seo, Jihwan Bang, Hwanjun Song, Deokki Hong, Seulki Park, Jung-Woo Ha 0001
ICLR7
2023 Self-Supervised Set Representation Learning for Unsupervised Meta-Learning
Dong Bok Lee, Seanie Lee, Kenji Kawaguchi, Yunji Kim, Jihwan Bang, Jung-Woo Ha 0001, Sung Ju Hwang
ICLR6
2023 Switching Temporary Teachers for Semi-Supervised Semantic Segmentation
abstract
The teacher-student framework, prevalent in semi-supervised semantic segmentation, mainly employs the exponential moving average (EMA) to update a single teacher's weights based on the student's. However, EMA updates raise a problem in that the weights of the teacher and student are getting coupled, causing a potential performance bottleneck. Furthermore, this problem may become more severe when training with more complicated labels such as segmentation masks but with few annotated data. This paper introduces Dual Teacher, a simple yet effective approach that employs dual temporary teachers aiming to alleviate the coupling problem for the student. The temporary teachers work in shifts and are progressively improved, so consistently prevent the teacher and student from becoming excessively close. Specifically, the temporary teachers periodically take turns generating pseudo-labels to train a student model and maintain the distinct characteristics of the student model for each epoch. Consequently, Dual Teacher achieves competitive performance on the PASCAL VOC, Cityscapes, and ADE20K benchmarks with remarkably shorter training times than state-of-the-art methods. Moreover, we demonstrate that our approach is model-agnostic and compatible with both CNN- and Transformer-based models. Code is available at https://github.com/naver-ai/dual-teacher.
Jaemin Na, Jung-Woo Ha 0001, Hyung Jin Chang, Dongyoon Han, Wonjun Hwang
NeurIPS2
2023 Entropy regularization for weakly supervised object localization
Dongjun Hwang, Jung-Woo Ha 0001, Hyunjung Shim, Junsuk Choe
Pattern Recognit. Lett.2
2022 Continuous Decomposition of Granularity for Neural Paraphrase Generation
abstract
While Transformers have had significant success in paragraph generation, they treat sentences as linear sequences of tokens and often neglect their hierarchical information. Prior work has shown that decomposing the levels of granularity (e.g., word, phrase, or sentence) for input tokens has produced substantial improvements, suggesting the possibility of enhancing Transformers via more fine-grained modeling of granularity. In this work, we present continuous decomposition of granularity for neural paraphrase generation (C-DNPG): an advanced extension of multi-head self-attention with: 1) a granularity head that automatically infers the hierarchical structure of a sentence by neurally estimating the granularity level of each input token; and 2) two novel attention masks, namely, granularity resonance and granularity scope, to efficiently encode granularity into attention. Experiments on two benchmarks, including Quora question pairs and Twitter URLs have shown that C-DNPG outperforms baseline models by a significant margin. Qualitative analysis reveals that C-DNPG indeed captures fine-grained levels of granularity with effectiveness.
Xiaodong Gu 0002, Sang-Woo Lee 0001, Kang Min Yoo, Jung-Woo Ha 0001
COLING5
2022 Online Continual Learning on a Contaminated Data Stream with Blurry Task Boundaries
abstract
Learning under a continuously changing data distribution with incorrect labels is a desirable real-world problem yet challenging. A large body of continual learning (CL) methods, however, assumes data streams with clean labels, and online learning scenarios under noisy data streams are yet underexplored. We consider a more practical CL task setup of an online learning from blurry data stream with corrupted labels, where existing CL methods struggle. To address the task, we first argue the importance of both diversity and purity of examples in the episodic memory of continual learning models. To balance diversity and purity in the episodic memory, we propose a novel strategy to manage and use the memory by a unified approach of label noise aware diverse sampling and robust learning with semi-supervised learning. Our empirical validations on four real-world or synthetic noise datasets (CI-FAR10 and 100, mini-WebVision, and Food-101N) exhibit that our method significantly outperforms prior arts in this realistic and challenging continual learning scenario. Code and data splits are available in https://github.com/clovaai/puridiver.
Jihwan Bang, Hyunseo Koh, Seulki Park, Hwanjun Song, Jung-Woo Ha 0001
CVPR5
2022 Generator Knows What Discriminator Should Learn in Unconditional GANs
Gayoung Lee, Hyunsu Kim, Seonghyeon Kim, Jung-Woo Ha 0001, Yunjey Choi
ECCV (17)5
2022 K-centered Patch Sampling for Efficient Video Recognition
Seong Hyeon Park, Jihoon Tack, Byeongho Heo, Jung-Woo Ha 0001, Jinwoo Shin
ECCV (35)4
2022 Contrastive Fine-grained Class Clustering via Generative Adversarial Networks
Yunji Kim, Jung-Woo Ha 0001
ICLR2
2022 Online Continual Learning on Class Incremental Blurry Task Configuration with Anytime Inference
Hyunseo Koh, Dahyun Kim 0001, Jung-Woo Ha 0001
ICLR3
2022 Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks
Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Jung-Woo Ha 0001, Jinwoo Shin
ICLR6
2022 Dataset Condensation via Efficient Synthetic-Data Parameterization
abstract
The great success of machine learning with massive amounts of data comes at a price of huge computation costs and storage for training and tuning. Recent studies on dataset condensation attempt to reduce the dependence on such massive data by synthesizing a compact training dataset. However, the existing approaches have fundamental limitations in optimization due to the limited representability of synthetic datasets without considering any data regularity characteristics. To this end, we propose a novel condensation framework that generates multiple synthetic data with a limited storage budget via efficient parameterization considering data regularity. We further analyze the shortcomings of the existing gradient matching-based condensation methods and develop an effective optimization technique for improving the condensation of training data information. We propose a unified algorithm that drastically improves the quality of condensed data against the current state-of-the-art on CIFAR-10, ImageNet, and Speech Commands.
Seong Joon Oh, Sangdoo Yun, Hwanjun Song, Joonhyun Jeong, Jung-Woo Ha 0001, Hyun Oh Song
ICML7
2022 Time Is MattEr: Temporal Self-supervision for Video Transformers
abstract
Understanding temporal dynamics of video is an essential aspect of learning better video representations. Recently, transformer-based architectural designs have been extensively explored for video tasks due to their capability to capture long-term dependency of input sequences. However, we found that these Video Transformers are still biased to learn spatial dynamics rather than temporal ones, and debiasing the spurious correlation is critical for their performance. Based on the observations, we design simple yet effective self-supervised tasks for video models to learn temporal dynamics better. Specifically, for debiasing the spatial bias, our method learns the temporal order of video frames as extra self-supervision and enforces the randomly shuffled frames to have low-confidence outputs. Also, our method learns the temporal flow direction of video tokens among consecutive frames for enhancing the correlation toward temporal dynamics. Under various video action recognition tasks, we demonstrate the effectiveness of our method and its compatibility with state-of-the-art Video Transformers.
Sukmin Yun, Jaehyung Kim 0001, Dongyoon Han, Hwanjun Song, Jung-Woo Ha 0001, Jinwoo Shin
ICML5
2022 On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model
abstract
Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woomyoung Park, Jung-Woo Ha, Nako Sung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Seongjin Shin, Sang-Woo Lee 0001, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woo-Myoung Park, Jung-Woo Ha 0001, Nako Sung
NAACL-HLT10
2022 On Divergence Measures for Bayesian Pseudocoresets
abstract
A Bayesian pseudocoreset is a small synthetic dataset for which the posterior over parameters approximates that of the original dataset. While promising, the scalability of Bayesian pseudocoresets is not yet validated in large-scale problems such as image classification with deep neural networks. On the other hand, dataset distillation methods similarly construct a small dataset such that the optimization with the synthetic dataset converges to a solution similar to optimization with full data. Although dataset distillation has been empirically verified in large-scale settings, the framework is restricted to point estimates, and their adaptation to Bayesian inference has not been explored. This paper casts two representative dataset distillation algorithms as approximations to methods for constructing pseudocoresets by minimizing specific divergence measures: reverse KL divergence and Wasserstein distance. Furthermore, we provide a unifying view of such divergence measures in Bayesian pseudocoreset construction. Finally, we propose a novel Bayesian pseudocoreset algorithm based on minimizing forward KL divergence. Our empirical results demonstrate that the pseudocoresets constructed from these methods reflect the true posterior even in large-scale Bayesian inference problems.
Balhae Kim, Jungwon Choi, Seanie Lee, Yoonho Lee 0001, Jung-Woo Ha 0001, Juho Lee 0001
NeurIPS5
2021 DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank Utterances
abstract
Recent advances in pre-trained language models have significantly improved neural response generation. However, existing methods usually view the dialogue context as a linear sequence of tokens and learn to generate the next word through token-level self-attention. Such token-level encoding hinders the exploration of discourse-level coherence among utterances. This paper presents DialogBERT, a novel conversational response generation model that enhances previous PLM-based dialogue models. DialogBERT employs a hierarchical Transformer architecture. To efficiently capture the discourse-level coherence among utterances, we propose two training objectives, including masked utterance regression and distributed utterance order ranking in analogy to the original BERT training. Experiments on three multi-turn conversation datasets show that our approach remarkably outperforms three baselines, such as BART and DialoGPT, in terms of quantitative evaluation. The human evaluation suggests that DialogBERT generates more coherent, informative, and human-like responses than the baselines with significant margins.
Xiaodong Gu 0002, Kang Min Yoo, Jung-Woo Ha 0001
AAAI3
2021 Rainbow Memory: Continual Learning With a Memory of Diverse Samples
abstract
Continual learning is a realistic learning scenario for AI models. Prevalent scenario of continual learning, however, assumes disjoint sets of classes as tasks and is less realistic rather artificial. Instead, we focus on ‘blurry’ task boundary; where tasks shares classes and is more realistic and practical. To address such task, we argue the importance of diversity of samples in an episodic memory. To enhance the sample diversity in the memory, we propose a novel memory management strategy based on per-sample classification uncertainty and data augmentation, named Rainbow Memory (RM). With extensive empirical validations on MNIST, CIFAR10, CIFAR100, and ImageNet datasets, we show that the proposed method significantly improves the accuracy in blurry continual learning setups, outperforming state of the arts by large margins despite its simplicity. Code and data splits will be available in https://github.com/clovaai/rainbow-memory.
Jihwan Bang, Heesu Kim, Young Joon Yoo, Jung-Woo Ha 0001
CVPR4
2021 What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers
abstract
Boseop Kim, HyoungSeok Kim, Sang-Woo Lee, Gichang Lee, Donghyun Kwak, Jeon Dong Hyeon, Sunghyun Park, Sungju Kim, Seonhoon Kim, Dongpil Seo, Heungsub Lee, Minyoung Jeong, Sungjae Lee, Minsub Kim, Suk Hyun Ko, Seokhun Kim, Taeyong Park, Jinuk Kim, Soyoung Kang, Na-Hyeon Ryu, Kang Min Yoo, Minsuk Chang, Soobin Suh, Sookyo In, Jinseong Park, Kyungduk Kim, Hiun Kim, Jisu Jeong, Yong Goo Yeo, Donghoon Ham, Dongju Park, Min Young Lee, Jaewook Kang, Inho Kang, Jung-Woo Ha, Woomyoung Park, Nako Sung. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Boseop Kim, HyoungSeok Kim, Sang-Woo Lee 0001, Gichang Lee, Donghyun Kwak, Dong Hyeon Jeon, Sunghyun Park 0005, Sungju Kim, Seonhoon Kim, Dongpil Seo, Heungsub Lee, Minyoung Jeong, Sungjae Lee 0002, Minsub Kim, SukHyun Ko, Seokhun Kim, Taeyong Park 0003, Soyoung Kang, Na-Hyeon Ryu, Kang Min Yoo, Minsuk Chang, Soobin Suh, Sookyo In, Kyungduk Kim, Hiun Kim, Jisu Jeong, Yong Goo Yeo, Donghoon Ham, Dongju Park, Min Young Lee, Jaewook Kang, Inho Kang, Jung-Woo Ha 0001, Woo-Myoung Park, Nako Sung
EMNLP (1)35
2021 St-Bert: Cross-Modal Language Model Pre-Training for End-to-End Spoken Language Understanding
abstract
Language model pre-training has shown promising results in various downstream tasks. In this context, we introduce a cross-modal pre-trained language model, called Speech-Text BERT (ST-BERT), to tackle end-to-end spoken language understanding (E2E SLU) tasks. Taking phoneme posterior and subword-level text as an input, ST-BERT learns a contextualized cross-modal alignment via our two proposed pre-training tasks: Cross-modal Masked Language Modeling (CM-MLM) and Cross-modal Conditioned Language Modeling (CM-CLM). Experimental results on three benchmarks present that our approach is effective for various SLU datasets and shows a surprisingly marginal performance degradation even when 1% of the training data are available. Also, our method shows further SLU performance gain via domain-adaptive pre-training with domain-specific speech-text pair data.
Gyuwan Kim, Sang-Woo Lee 0001, Jung-Woo Ha 0001
ICASSP4
2021 AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, Jung-Woo Ha 0001
ICLR8
2021 Metropolis-Hastings Data Augmentation for Graph Neural Networks
abstract
Graph Neural Networks (GNNs) often suffer from weak-generalization due to sparsely labeled data despite their promising results on various graph-based tasks. Data augmentation is a prevalent remedy to improve the generalization ability of models in many domains. However, due to the non-Euclidean nature of data space and the dependencies between samples, designing effective augmentation on graphs is challenging. In this paper, we propose a novel framework Metropolis-Hastings Data Augmentation (MH-Aug) that draws augmented graphs from an explicit target distribution for semi-supervised learning. MH-Aug produces a sequence of augmented graphs from the target distribution enables flexible control of the strength and diversity of augmentation. Since the direct sampling from the complex target distribution is challenging, we adopt the Metropolis-Hastings algorithm to obtain the augmented samples. We also propose a simple and effective semi-supervised learning strategy with generated samples from MH-Aug. Our extensive experiments demonstrate that MH-Aug can generate a sequence of samples according to the target distribution to significantly improve the performance of GNNs.
Hyeon-Jin Park, Seunghun Lee 0001, Sihyeon Kim, Jinyoung Park 0005, Jisu Jeong, Jung-Woo Ha 0001, Hyunwoo J. Kim
NeurIPS7
2021 Region-based dropout with attention prior for weakly supervised object localization
Junsuk Choe, Dongyoon Han, Sangdoo Yun, Jung-Woo Ha 0001, Seong Joon Oh, Hyunjung Shim
Pattern Recognit.4
2020 StarGAN v2: Diverse Image Synthesis for Multiple Domains
abstract
A good image-to-image translation model should learn a mapping between different visual domains while satisfying the following properties: 1) diversity of generated images and 2) scalability over multiple domains. Existing methods address either of the issues, having limited diversity or multiple models for all domains. We propose StarGAN v2, a single framework that tackles both and shows significantly improved results over the baselines. Experiments on CelebA-HQ and a new animal faces dataset (AFHQ) validate our superiority in terms of visual quality, diversity, and scalability. To better assess image-to-image translation models, we release AFHQ, high-quality animal faces with large inter- and intra-domain differences. The code, pretrained models, and dataset are available at https://github.com/clovaai/stargan-v2.
Yunjey Choi, Youngjung Uh, Jaejun Yoo 0001, Jung-Woo Ha 0001
CVPR4
2020 Context-Aware Answer Extraction in Question Answering
abstract
Extractive QA models have shown very promising performance in predicting the correct answer to a question for a given passage.However, they sometimes result in predicting the correct answer text but in a context irrelevant to the given question.This discrepancy becomes especially important as the number of occurrences of the answer text in a passage increases.To resolve this issue, we propose BLANC (BLock AttentioN for Context prediction) based on two main ideas: context prediction as an auxiliary task in multi-task learning manner, and a block attention method that learns the context prediction task.With experiments on reading comprehension, we show that BLANC outperforms the state-ofthe-art QA models, and the performance gap increases as the number of answer text occurrences increases.We also conduct an experiment of training the models using SQuAD and predicting the supporting facts on HotpotQA and show that BLANC outperforms all baseline models in this zero-shot setting.
Yeon Seonwoo, Jung-Woo Ha 0001, Alice Oh
EMNLP (1)3
2020 ClovaCall: Korean Goal-Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact Centers
abstract
Automatic speech recognition (ASR) via call is essential for various applications, including AI for contact center (AICC) services. Despite the advancement of ASR, however, most publicly available call-based speech corpora such as Switchboard are old-fashioned. Also, most existing call corpora are in English and mainly focus on open domain dialog or general scenarios such as audiobooks. Here we introduce a new large-scale Korean call-based speech corpus under a goal-oriented dialog scenario from more than 11,000 people, i.e., ClovaCall corpus. ClovaCall includes approximately 60,000 pairs of a short sentence and its corresponding spoken utterance in a restaurant reservation domain. We validate the effectiveness of our dataset with intensive experiments using two standard ASR models. Furthermore, we release our ClovaCall dataset and baseline source codes to be available via https://github.com/ClovaAI/ClovaCall. Copyright © 2020 ISCA
Jung-Woo Ha 0001, Kihyun Nam, Sang-Woo Lee 0001, Sohee Yang, Hyunhoon Jung, Hyeji Kim, Eunmi Kim, Soojin Kim, Hyun Ah Kim, Kyoungtae Doh, Chan Kyu Lee, Nako Sung, Sunghun Kim 0001
INTERSPEECH1
2020 Self-supervised Auxiliary Learning with Meta-paths for Heterogeneous Graphs
abstract
Graph neural networks have shown superior performance in a wide range of applications providing a powerful representation of graph-structured data. Recent works show that the representation can be further improved by auxiliary tasks. However, the auxiliary tasks for heterogeneous graphs, which contain rich semantic information with various types of nodes and edges, have less explored in the literature. In this paper, to learn graph neural networks on heterogeneous graphs we propose a novel self-supervised auxiliary learning method using meta paths, which are composite relations of multiple edge types. Our proposed method is learning to learn a primary task by predicting meta-paths as auxiliary tasks. This can be viewed as a type of meta-learning. The proposed method can identify an effective combination of auxiliary tasks and automatically balance them to improve the primary task. Our methods can be applied to any graph neural networks in a plug-in manner without manual labeling or additional data. The experiments demonstrate that the proposed method consistently improves the performance of link prediction and node classification on heterogeneous graphs.
Dasol Hwang, Jinyoung Park 0005, Sunyoung Kwon, Jung-Woo Ha 0001, Hyunwoo J. Kim
NeurIPS5
2019 Paraphrase Diversification Using Counterfactual Debiasing
abstract
The problem of generating a set of diverse paraphrase sentences while (1) not compromising the original meaning of the original sentence, and (2) imposing diversity in various semantic aspects, such as a lexical or syntactic structure, is examined. Existing work on paraphrase generation has focused more on the former, and the latter was trained as a fixed style transfer, such as transferring from positive to negative sentiments, even at the cost of losing semantics. In this work, we consider style transfer as a means of imposing diversity, with a paraphrasing correctness constraint that the target sentence must remain a paraphrase of the original sentence. However, our goal is to maximize the diversity for a set of k generated paraphrases, denoted as the diversified paraphrase (DP) problem. Our key contribution is deciding the style guidance at generation towards the direction of increasing the diversity of output with respect to those generated previously. As pre-materializing training data for all style decisions is impractical, we train with biased data, but with debiasing guidance. Compared to state-of-the-art methods, our proposed model can generate more diverse and yet semantically consistent paraphrase sentences. That is, our model, trained with the MSCOCO dataset, achieves the highest embedding scores, .94/.95/.86, similar to state-of-the-art results, but with a lower mBLEU score (more diverse) by 8.73%.
Sunghyun Park 0005, Seung-won Hwang, Fuxiang Chen, Jaegul Choo, Jung-Woo Ha 0001, Sunghun Kim 0001, Jinyeong Yim
AAAI5
2019 NL2pSQL: Generating Pseudo-SQL Queries from Under-Specified Natural Language Questions
abstract
Fuxiang Chen, Seung-won Hwang, Jaegul Choo, Jung-Woo Ha, Sunghun Kim. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Fuxiang Chen, Seung-won Hwang, Jaegul Choo, Jung-Woo Ha 0001, Sunghun Kim 0001
EMNLP/IJCNLP (1)4
2019 Photorealistic Style Transfer via Wavelet Transforms
abstract
Recent style transfer models have provided promising artistic results. However, given a photograph as a reference style, existing methods are limited by spatial distortions or unrealistic artifacts, which should not happen in real photographs. We introduce a theoretically sound correction to the network architecture that remarkably enhances photorealism and faithfully transfers the style. The key ingredient of our method is wavelet transforms that naturally fits in deep networks. We propose a wavelet corrected transfer based on whitening and coloring transforms (WCT2) that allows features to preserve their structural information and statistical properties of VGG feature space during stylization. This is the first and the only end-to-end model that can stylize a 1024x1024 resolution image in 4.7 seconds, giving a pleasing and photorealistic quality without any post-processing. Last but not least, our model provides a stable video stylization without temporal constraints. Our code, generated images, pre-trained models and supplementary documents are all available at https://github.com/ClovaAI/WCT2.
Jaejun Yoo 0001, Youngjung Uh, Sanghyuk Chun, Byeongkyu Kang, Jung-Woo Ha 0001
ICCV5
2019 Phase-Aware Speech Enhancement with Deep Complex U-Net
Hyeong-Seok Choi, Jaesung Huh, Adrian Kim, Jung-Woo Ha 0001, Kyogu Lee
ICLR (Poster)5
2019 DialogWAE: Multimodal Response Generation with Conditional Wasserstein Auto-Encoder
Xiaodong Gu 0002, Kyunghyun Cho, Jung-Woo Ha 0001, Sunghun Kim 0001
ICLR (Poster)3
2019 Large-Scale Answerer in Questioner's Mind for Visual Dialog Question Generation
Sang-Woo Lee 0001, Sohee Yang, Jaejun Yoo 0001, Jung-Woo Ha 0001
ICLR (Poster)5
2018 Interpretable Prediction of Vascular Diseases from Electronic Health Records via Deep Attention Networks
abstract
Precise prediction of severe diseases resulting in mortality is one of the main issues in medical fields. Even if pathological and radiological measurements provide competitive precision, they usually require large costs of time and expense to obtain and analyze the data for prediction. Recently, end-to-end approaches based on deep neural networks have been proposed, however, they still suffer from the low classification performance and difficulties of interpretation. In this study, we propose a novel disease prediction method, EHAN (EHR History-based prediction using Attention Network), based on the recurrent neural network (RNN) and attention mechanism. The proposed method incorporates (1) a bidirectional gated recurrent units (GRU) for automated sequential modeling, (2) attention mechanism for improving long-term dependence modeling, (3) RNN-based gradient-weighted class activation mapping (Grad-CAM) to visualize the class specific attention-weights. We conducted the experiments to predict the occurrence of risky disease containing cardiovascular and cerebrovascular diseases from more than 40,000 hypertension patients' electronic health records (EHR). The results showed that the proposed method outperformed the state-of-the-art model with respect to the various performance metrics. Furthermore, we confirmed that the proposed visualizing methods can be used to assist data-driven discovery.
Seunghyun Park 0001, You Jin Kim, Jeong-Whun Kim, Jin Joo Park, Borim Ryu, Jung-Woo Ha 0001
BIBE6
2018 StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation
abstract
Recent studies have shown remarkable success in image-to-image translation for two domains. However, existing approaches have limited scalability and robustness in handling more than two domains, since different models should be built independently for every pair of image domains. To address this limitation, we propose StarGAN, a novel and scalable approach that can perform image-to-image translations for multiple domains using only a single model. Such a unified model architecture of StarGAN allows simultaneous training of multiple datasets with different domains within a single network. This leads to StarGAN's superior quality of translated images compared to existing models as well as the novel capability of flexibly translating an input image to any desired target domain. We empirically demonstrate the effectiveness of our approach on a facial attribute transfer and a facial expression synthesis tasks.
Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha 0001, Sunghun Kim 0001, Jaegul Choo
CVPR4
2017 Dual Attention Networks for Multimodal Reasoning and Matching
abstract
We propose Dual Attention Networks (DANs) which jointly leverage visual and textual attention mechanisms to capture fine-grained interplay between vision and language. DANs attend to specific regions in images and words in text through multiple steps and gather essential information from both modalities. Based on this framework, we introduce two types of DANs for multimodal reasoning and matching, respectively. The reasoning model allows visual and textual attentions to steer each other during collaborative inference, which is useful for tasks such as Visual Question Answering (VQA). In addition, the matching model exploits the two attention mechanisms to estimate the similarity between images and sentences by focusing on their shared semantics. Our extensive experiments validate the effectiveness of DANs in combining vision and language, achieving the state-of-the-art performance on public benchmarks for VQA and image-text matching.
Hyeonseob Nam, Jung-Woo Ha 0001, Jeonghee Kim
CVPR2
2017 Hadamard Product for Low-rank Bilinear Pooling
Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha 0001, Byoung-Tak Zhang
ICLR (Poster)5
2017 Overcoming Catastrophic Forgetting by Incremental Moment Matching
abstract
Catastrophic forgetting is a problem of neural networks that loses the information of the first task after training the second task. Here, we propose a method, i.e. incremental moment matching (IMM), to resolve this problem. IMM incrementally matches the moment of the posterior distribution of the neural network which is trained on the first and the second task, respectively. To make the search space of posterior parameter smooth, the IMM procedure is complemented by various transfer learning techniques including weight transfer, L2-norm of the old and the new parameter, and a variant of dropout with the old parameter. We analyze our approach on a variety of datasets including the MNIST, CIFAR-10, Caltech-UCSD-Birds, and Lifelog datasets. The experimental results show that IMM achieves state-of-the-art performance by balancing the information between an old and a new network.
Sang-Woo Lee 0001, Jin-Hwa Kim, Jaehyun Jun, Jung-Woo Ha 0001, Byoung-Tak Zhang
NIPS4
2017 Dual-memory neural networks for modeling cognitive activities of humans via wearable sensors
Sang-Woo Lee 0001, Chung-Yeon Lee, Donghyun Kwak, Jung-Woo Ha 0001, Jeonghee Kim, Byoung-Tak Zhang
Neural Networks4
2016 Large-Scale Item Categorization in e-Commerce Using Multiple Recurrent Neural Networks
abstract
Precise item categorization is a key issue in e-commerce domains. However, it still remains a challenging problem due to data size, category skewness, and noisy metadata. Here, we demonstrate a successful report on a deep learning-based item categorization method, i.e., deep categorization network (DeepCN), in an e-commerce website. DeepCN is an end-to-end model using multiple recurrent neural networks (RNNs) dedicated to metadata attributes for generating features from text metadata and fully connected layers for classifying item categories from the generated features. The categorization errors are propagated back through the fully connected layers to the RNNs for weight update in the learning process. This deep learning-based approach allows diverse attributes to be integrated into a common representation, thus overcoming sparsity and scalability problems. We evaluate DeepCN on large-scale real-world data including more than 94 million items with approximately 4,100 leaf categories from a Korean e-commerce website. Experiment results show our method improves the categorization accuracy compared to the model using single RNN as well as a standard classification model using unigram-based bag-of-words. Furthermore, we investigate how much the model parameters and the used attributes influence categorization performances.
Jung-Woo Ha 0001, Hyuna Pyo, Jeonghee Kim
KDD1
2016 Multimodal Residual Learning for Visual QA
abstract
Deep neural networks continue to advance the state-of-the-art of image recognition tasks with various methods. However, applications of these methods to multimodality remain limited. We present Multimodal Residual Networks (MRN) for the multimodal residual learning of visual question-answering, which extends the idea of the deep residual learning. Unlike the deep residual learning, MRN effectively learns the joint representation from visual and language information. The main idea is to use element-wise multiplication for the joint residual mappings exploiting the residual learning of the attentional models in recent studies. Various alternative models introduced by multimodality are explored based on our study. We achieve the state-of-the-art results on the Visual QA dataset for both Open-Ended and Multiple-Choice tasks. Moreover, we introduce a novel method to visualize the attention effect of the joint representations for each learning block using back-propagation algorithm, even though the visual features are collapsed without spatial information.
Jin-Hwa Kim, Sang-Woo Lee 0001, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha 0001, Byoung-Tak Zhang
NIPS6
2015 Automated Construction of Visual-Linguistic Knowledge via Concept Learning from Cartoon Videos
abstract
Learning mutually-grounded vision-language knowledge is a foundational task for cognitive systems and human-level artificial intelligence. Most of knowledge-learning techniques are focused on single modal representations in a static environment with a fixed set of data. Here, we explore an ecologically more-plausible setting by using a stream of cartoon videos to build vision-language concept hierarchies continuously. This approach is motivated by the literature on cognitive development in early childhood. We present the model of deep concept hierarchy (DCH) that enables the progressive abstraction of concept knowledge in multiple levels. We develop a stochastic method for graph construction, i.e. a graph Monte Carlo algorithm, to search efficiently the huge compositional space of the vision-language concepts. The concept hierarchies are built incrementally and can handle concept drift, allowing for being deployed in lifelong learning environments. Using a series of approximately 200 episodes of educational cartoon videos we demonstrate the emergence and evolution of the concept hierarchies as the video stories unfold. We also present the application of the deep concept hierarchies for context-dependent translation between vision and language, i.e. the transcription of a visual scene into text and the generation of visual imagery from text.
Jung-Woo Ha 0001, Byoung-Tak Zhang
AAAI1
2014 Bayesian evolutionary hypergraph learning for predicting cancer clinical outcomes
Soo-Jin Kim, Jung-Woo Ha 0001, Byoung-Tak Zhang
J. Biomed. Informatics2
2013 Evolutionary concept learning from cartoon videos by multimodal hypernetworks
abstract
Concepts have been widely used for categorizing and representing knowledge in artificial intelligence. Previous researches on concept learning have focused on unimodal data, usually on linguistic domains in a static environment. Concept learning from multimodal stream data, such as videos, remains a challenge due to their dynamic change and high-dimensionality. Here we propose an evolutionary method that simulates the process of human concept learning from multimodal video streams. Two key ideas on evolutionary concept learning are representing concepts in a large collection (population) of hyperedges or a hypergraph and to incrementally learning from video streams based on an evolutionary approach. The hypergraph is learned "evolutionarily" by repeating the generation and selection process of hyperedge concepts from the video data. The advantage of this evolutionary learning process is that the population-based distributed coding allows flexible and robust trace of the change of concept relations as the video story unfolds. We evaluate the proposed method on a suite of children's cartoon videos for 517 minutes of total playing time. Experimental results show that the proposed method effectively represents visual-textual concept relations and our evolutionary concept learning method effectively models the conceptual change as an evolutionary process. We also investigate the structure properties of the constructed concept networks.
Beom-Jin Lee, Jung-Woo Ha 0001, Byoung-Tak Zhang
IEEE Congress on Evolutionary Computation2
2012 Sparse Population Code Models of Word Learning in Concept Drift
Byoung-Tak Zhang, Jung-Woo Ha 0001, Myunggu Kang
CogSci2
2012 Text-to-image retrieval based on incremental association via multimodal hypernetworks
abstract
Text-to-image retrieval is to retrieve the images associated with the textual queries. A text-to-image retrieval model requires an incremental learning method for its practical use since the multimodal data grow up dramatically. Here we propose an incremental text-to-image retrieval method using a multimodal association model. The association model is based on a hypernetwork (HN) where a vertex corresponds to a textual word or a visual patch and a hyperedge represents a higher-order multimodal association. Using the HN incrementally learned by a sequential Bayesian sampling, in the multimodal hypernetwork-based text-to-image retrieval, a given text query is crossmodally expanded to the visual query and then similar images are retrieved to the expanded visual query. We evaluated the proposed method using 3,000 images with textual description from Flickr.com. The experimental results present that the proposed method achieves very competitive retrieval performances compared to a baseline method. Moreover, we demonstrate that our method provides robust text-to-image retrieval results for the increasing data.
Jung-Woo Ha 0001, Beom-Jin Lee, Byoung-Tak Zhang
SMC1
2011 Mutual information-based evolution of hypernetworks for brain data analysis
abstract
Cortical analysis becomes increasingly important for brain research and clinical diagnosis. This problem involves a combinatorial search to find the essential modules among a large number of brain regions. Despite several statistical approaches, cortical analysis remains a formidable challenge due to high dimensionality and sparsity of data. Here we describe an evolutionary method for finding significant modules from cortical data. The method uses a hypernetwork which is encoded as a population of hyperedges, where hyperedges represent building blocks or potential modules. We develop an efficient method for evolving the hypernetwork using mutual information to generate essential hyperedges. We evaluate the method on predicting intelligence quotient (IQ) levels and finding potential significant modules on IQ from brain MRI data consisting of 62 healthy adults with over 80,000 measured points (variables). The experimental results show that our information-theoretic evolutionary hypernetworks improve the classification accuracy by 5-15%. Moreover, it extracts significant cortical modules that distinguish high IQ from low IQ groups.
Eun-Sol Kim, Jung-Woo Ha 0001, Wi Hoon Jung, Joon Hwan Jang, Jun Soo Kwon, Byoung-Tak Zhang
IEEE Congress on Evolutionary Computation2
2010 Evolutionary layered hypernetworks for identifying microRNA-mRNA regulatory modules
abstract
Exploring micro RNA (miRNA) and mRNA regulatory interactions may give new insights into diverse biological phenomena. While elucidating complex miRNA-mRNA interactions has been studied with experimental and computational approaches, it is still difficult to infer miRNA-mRNA regulatory modules. Here we present a novel method for identifying functional miRNA-mRNA modules from heterogeneous expression data. The proposed approach is layered hypernetworks consisting of two layers which are the layer of modality-dependent hypernetworks and of an integrating hypernetwork. The layered hypernetwork model is suitable for detecting relationships between heterogeneous modalities. Applied to the analysis of miRNA and mRNA expression profiles on multiple human cancers, the proposed model identifies oncogenic miRNA-mRNA regulatory modules. The experimental results show that our method provides a competitive performance to support vector machines, and outperforms other standard machine learning algorithms. The biological significance of the discovered miRNA-mRNA modules were validated by literature reviews.
Soo-Jin Kim, Jung-Woo Ha 0001, Bado Lee, Byoung-Tak Zhang
IEEE Congress on Evolutionary Computation2
2010 Layered Hypernetwork Models for Cross-Modal Associative Text and Image Keyword Generation in Multimodal Information Retrieval
Jung-Woo Ha 0001, Byoung-Hee Kim, Bado Lee, Byoung-Tak Zhang
PRICAI1
2009 Gender classification with cortical thickness measurement from magnetic resonance imaging by using a feature selection method based on evolutionary hypernetworks
abstract
Hypernetworks are a weighted hypergraph where evolutionary methods are learning the model structure and parameters. The evolutionary methods enable the hypernetwork model to conserve significant features implicitly during the learning process. In this study, we propose a novel feature selection method based on occurrence frequencies of attributes in hyperedges by analyzing the structure of a hypernetwork. We also apply the evolutionary hypernetwork with the proposed feature selection method to the gender classification based on cortical thickness measurement on healthy young adults from Magnetic Resonance Imaging (MRI). The experimental results show that the proposed selection method improves the classification accuracy by approximately 20%. Also, a comparative study on four classification algorithms and three feature selection methods shows that the hypernetwork model with the proposed feature selection method achieves a competitive classification performance.
Jung-Woo Ha 0001, Joon Hwan Jang, Do-Hyung Kang, Wi Hoon Jung, Jun Soo Kwon, Byoung-Tak Zhang
FUZZ-IEEE1