EDBT 2026 Demo / reviewers in the wild / expert
Gunhee Kim
dblp:45/115
· DBLP profile ↗
133ranked-venue papers
24as first author
67since 2021 · last 2026
0000-0002-9543-7453ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 122 · 19 first-author · 65 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 10 first-author · 23 since 2021Systems, architecture and hardware · 7 · 5 first-authorDatabases, data management, data science and information retrieval · 7 · 4 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gaussian Blending: Rethinking Alpha Blending in 3D Gaussian SplattingabstractThe recent introduction of 3D Gaussian Splatting (3DGS) has significantly advanced novel view synthesis. Several studies have further improved the rendering quality of 3DGS, yet they still exhibit noticeable visual discrepancies when synthesizing views at sampling rates unseen during training. Specifically, they suffer from (i) erosion-induced blurring artifacts when zooming in and (ii) dilation-induced staircase artifacts when zooming out. We speculate that these artifacts arise from the fundamental limitation of the alpha blending adopted in 3DGS methods. Instead of the conventional alpha blending that computes alpha and transmittance as scalar quantities over a pixel, we propose to replace it with our novel Gaussian Blending that treats alpha and transmittance as spatially varying distributions. Thus, transmittances can be updated considering the spatial distribution of alpha values across the pixel area, allowing nearby background splats to contribute to the final rendering. Our Gaussian Blending maintains real-time rendering speed and requires no additional memory cost, while being easily integrated as a drop-in replacement into existing 3DGS-based or other NVS frameworks. Extensive experiments demonstrate that Gaussian Blending effectively captures fine details at various sampling rates unseen during training, consistently outperforming existing novel view synthesis models across both unseen and seen sampling rates. Junseo Koo, Jinseo Jeong, Gunhee Kim |
AAAI | 3 |
| 2026 | MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question AnsweringabstractSource attribution aims to enhance the reliability of AI-generated answers by including references for each statement, helping users validate the provided answers. However, existing work has primarily focused on text-only scenario and largely overlooked the role of multimodality. We introduce MAVIS, the first benchmark designed to evaluate multimodal source attribution systems that understand user intent behind visual questions, retrieve multimodal evidence, and generate long-form answers with citations. Our dataset comprises 157K visual QA instances, where each answer is annotated with fact-level citations referring to multimodal documents. We develop fine-grained automatic metrics along three dimensions of informativeness, groundedness, and fluency, and demonstrate their strong correlation with human judgments. Our key findings are threefold: (1) LVLMs with multimodal RAG generate more informative and fluent answers than unimodal RAG, but they exhibit weaker groundedness for image documents than for text documents, a gap amplified in multimodal settings. (2) Given the same multimodal documents, there is a trade-off between informativeness and groundedness across different prompting methods. (3) Our proposed method highlights mitigating contextual bias in interpreting image documents as a crucial direction for future research. Seokwon Song, Gunhee Kim |
AAAI | 3 |
| 2026 | Towards Scene-Aware Video-to-Spatial Audio Generation
Jaeyeon Kim, Heeseung Yun, Gunhee Kim |
Int. J. Comput. Vis. | 3 |
| 2025 | Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text UpdatesabstractWhile pre-trained multimodal representations (e.g., CLIP) have shown impressive capabilities, they exhibit significant compositional vulnerabilities leading to counterintuitive judgments.We introduce Multimodal Adversarial Compositionality (MAC), a benchmark that leverages large language models (LLMs) to generate deceptive text samples to exploit these vulnerabilities across different modalities and evaluates them through both sample-wise attack success rate and group-wise entropy-based diversity.To improve zero-shot methods, we propose a self-training approach that leverages rejectionsampling fine-tuning with diversity-promoting filtering, which enhances both attack success rate and sample diversity.Using smaller language models like Llama-3.1-8B, our approach demonstrates superior performance in revealing compositional vulnerabilities across various multimodal representations, including images, videos, and audios. Jaewoo Ahn, Heeseung Yun, Dayoon Ko, Gunhee Kim |
ACL (1) | 4 |
| 2025 | LPOI: Listwise Preference Optimization for Vision Language ModelsabstractAligning large VLMs with human preferences is a challenging task, as methods like RLHF and DPO often overfit to textual information or exacerbate hallucinations.Although augmenting negative image samples partially addresses these pitfalls, no prior work has employed listwise preference optimization for VLMs, due to the complexity and cost of constructing listwise image samples.In this work, we propose LPOI, the first object-aware listwise preference optimization developed for reducing hallucinations in VLMs.LPOI identifies and masks a critical object in the image, and then interpolates the masked region between the positive and negative images to form a sequence of incrementally more complete images.The model is trained to rank these images in ascending order of object visibility, effectively reducing hallucinations while retaining visual fidelity.LPOI requires no extra annotations beyond standard pairwise preference data, as it automatically constructs the ranked lists through object masking and interpolation.Comprehensive experiments on MMHalBench, AMBER, and Object HalBench confirm that LPOI outperforms existing preference optimization methods in reducing hallucinations and enhancing VLM performance.We make the code available at https: //github.com/fatemehpesaran310/lpoi. Fatemeh Pesaran Zadeh, Yoojin Oh, Gunhee Kim |
ACL (1) | 3 |
| 2025 | ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data StreamsabstractThe rapid growth of video-text data presents challenges in storage and computation during training. Online learning, which processes streaming data in real-time, offers a promising solution to these issues while also allowing swift adaptations in scenarios demanding real-time responsiveness. One strategy to enhance the efficiency and effectiveness of learning involves identifying and prioritizing data that enhances performance on target downstream tasks. We propose Relevance and Specificity-based online filtering framework (ReSpec) that selects data based on four criteria: (i) modality alignment for clean data, (ii) task relevance for target focused data, (iii) specificity for informative and detailed data, and (iv) efficiency for low-latency processing. Relevance is determined by the probabilistic alignment of incoming data with downstream tasks, while specificity employs the distance to a root embedding representing the least specific data as an efficient proxy for informativeness. By establishing reference points from target task data, ReSpec filters incoming data in real-time, eliminating the need for extensive storage and compute. Evaluating on large-scale datasets WebVid2M and VideoCC3M, ReSpec attains state-of-the-art performance on five zero-shot video retrieval tasks, using as little as 5% of the data while incurring minimal compute. The source code is available at https://github.com/cdjkim/ReSpec. Chris Dongjoo Kim, Jihwan Moon 0002, Sangwoo Moon 0001, Heeseung Yun, Sihaeng Lee, Aniruddha Kembhavi, Soonyoung Lee, Gunhee Kim, Sangho Lee 0008 |
CVPR | 8 |
| 2025 | HalLoc: Token-level Localization of Hallucinations for Vision Language ModelsabstractHallucinations pose a significant challenge to the reliability of large vision-language models, making their detection essential for ensuring accuracy in critical applications. Current detection methods often rely on computationally intensive models, leading to high latency and resource demands. Their definitive outcomes also fail to account for real-world scenarios where the line between hallucinated and truthful information is unclear. To address these issues, we propose HalLoc, a dataset designed for efficient, probabilistic hallucination detection. It features 150K token-level annotated samples, including hallucination types, across Visual Question Answering (VQA), instruction-following, and image captioning tasks. This dataset facilitates the development of models that detect hallucinations with graded confidence, enabling more informed user interactions. Additionally, we introduce a baseline model trained on HalLoc, offering low-overhead, concurrent hallucination detection during generation. The model can be seamlessly integrated into existing VLMs, improving reliability while preserving efficiency. The prospect of a robust plug-and-play hallucination detection module opens new avenues for enhancing the trustworthiness of vision-language models in real-world applications. The HalLoc dataset and code are publicly available at: https://github.com/dbsltm/cvpr25_halloc. Eunkyu Park, Minyeong Kim 0003, Gunhee Kim |
CVPR | 3 |
| 2025 | FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure GamesabstractGUI agents powered by LLMs show promise in interacting with diverse digital environments.Among these, video games offer a valuable testbed due to their varied interfaces, with adventure games posing additional challenges through complex, narrative-driven interactions.Existing game benchmarks, however, lack diversity and rarely evaluate agents on completing entire storylines.To address this, we introduce FlashAdventure, a benchmark of 34 Flashbased adventure games designed to test full story arc completion and tackle the observationbehavior gap: the challenge of remembering and acting on earlier gameplay information.We also propose CUA-as-a-Judge, an automated gameplay evaluator, and COAST, an agentic framework leveraging long-term clue memory to better plan and solve sequential tasks.Experiments show current GUI agents struggle with full story arcs, while COAST improves milestone completion by bridging the observationbehavior gap.Nonetheless, a marked discrepancy between humans and best-performing agents warrants continued research efforts to narrow this divide. * Equal contribution. †Work done during an internship at KRAFTON. Flash-Based Adventure GamesInput GUI Agent (Operator) Gameplay Jaewoo Ahn, Junseo Kim, Heeseung Yun, Jaehyeon Son, Dongmin Park, Jaewoong Cho, Gunhee Kim |
EMNLP | 7 |
| 2025 | Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible SpeechabstractSpoken dialogue systems increasingly employ large language models (LLMs) to leverage their advanced reasoning capabilities.However, direct application of LLMs in spoken communication often yield suboptimal results due to mismatches between optimal textual and verbal delivery.While existing approaches adapt LLMs to produce speech-friendly outputs, their impact on reasoning performance remains underexplored.In this work, we propose THINK-VERBALIZE-SPEAK, a framework that decouples reasoning from spoken delivery to preserve the full reasoning capacity of LLMs.Central to our method is verbalizing, an intermediate step that translates thoughts into natural, speech-ready text.We also introduce REVERT, a latency-efficient verbalizer based on incremental and asynchronous summarization.Experiments across multiple benchmarks show that our method enhances speech naturalness and conciseness with minimal impact on reasoning.The project page with the dataset and the source code is available at https: //yhytoto12.github.io/TVS-ReVerT. Tony Woo, Sehun Lee, Kang-Wook Kim 0002, Gunhee Kim |
EMNLP | 4 |
| 2025 | ChartCap: Mitigating Hallucination of Dense Chart Captioning
Junyoung Lim, Jaewoo Ahn, Gunhee Kim |
ICCV | 3 |
| 2025 | FedMeNF: Privacy-Preserving Federated Meta-Learning for Neural FieldsabstractNeural fields provide a memory-efficient representation of data, which can effectively handle diverse modalities and large-scale data. However, learning to map neural fields often requires large amounts of training data and computations, which can be limited to resource-constrained edge devices. One approach to tackle this limitation is to leverage Federated Meta-Learning (FML), but traditional FML approaches suffer from privacy leakage. To address these issues, we introduce a novel FML approach called FedMeNF. FedMeNF utilizes a new privacy-preserving loss function that regulates privacy leakage in the local meta-optimization. This enables the local meta-learner to optimize quickly and efficiently without retaining the client's private data. Our experiments demonstrate that FedMeNF achieves fast optimization speed and robust reconstruction performance, even with few-shot or non-IID data across diverse data modalities, while preserving client data privacy. Junhyeog Yun, Minui Hong, Gunhee Kim |
ICCV | 3 |
| 2025 | ViSAGe: Video-to-Spatial Audio GenerationabstractSpatial audio is essential for enhancing the immersiveness of audio-visual experiences, yet its production typically demands complex recording systems and specialized expertise. In this work, we address a novel problem of generating first-order ambisonics, a widely used spatial audio format, directly from silent videos. To support this task, we introduce YT-Ambigen, a dataset comprising 102K 5-second YouTube video clips paired with corresponding first-order ambisonics. We also propose new evaluation metrics to assess the spatial aspect of generated audio based on audio energy maps and saliency metrics. Furthermore, we present Video-to-Spatial Audio Generation (ViSAGe), an end-to-end framework that generates first-order ambisonics from silent video frames by leveraging CLIP visual features, autoregressive neural audio codec modeling with both directional and visual guidance. Experimental results demonstrate that ViSAGe produces plausible and coherent first-order ambisonics, outperforming two-stage approaches consisting of video-to-audio generation and audio spatialization. Qualitative examples further illustrate that ViSAGe generates temporally aligned high-quality spatial audio that adapts to viewpoint changes. Jaeyeon Kim, Heeseung Yun, Gunhee Kim |
ICLR | 3 |
| 2025 | Distilling Reinforcement Learning Algorithms for In-Context Model-Based PlanningabstractRecent studies have shown that Transformers can perform in-context reinforcement learning (RL) by imitating existing RL algorithms, enabling sample-efficient adaptation to unseen tasks without parameter updates. However, these models also inherit the suboptimal behaviors of the RL algorithms they imitate. This issue primarily arises due to the gradual update rule employed by those algorithms. Model-based planning offers a promising solution to this limitation by allowing the models to simulate potential outcomes before taking action, providing an additional mechanism to deviate from the suboptimal behavior. Rather than learning a separate dynamics model, we propose Distillation for In-Context Planning (DICP), an in-context model-based RL framework where Transformers simultaneously learn environment dynamics and improve policy in-context. We evaluate DICP across a range of discrete and continuous environments, including Darkroom variants and Meta-World. Our results show that DICP achieves state-of-the-art performance while requiring significantly fewer environment interactions than baselines, which include both model-free counterparts and existing meta-RL methods. Jaehyeon Son, Soochan Lee, Gunhee Kim |
ICLR | 3 |
| 2025 | Meta-Continual Learning of Neural FieldsabstractNeural Fields (NF) have gained prominence as a versatile framework for complex data representation. This work unveils a new problem setting termed Meta-Continual Learning of Neural Fields (MCL-NF) and introduces a novel strategy that employs a modular architecture combined with optimization-based meta-learning. Focused on overcoming the limitations of existing methods for continual learning of neural fields, such as catastrophic forgetting and slow convergence, our strategy achieves high-quality reconstruction with significantly improved learning speed. We further introduce Fisher Information Maximization loss for neural radiance fields (FIM-NeRF), which maximizes information gains at the sample level to enhance learning generalization, with proved convergence guarantee and generalization bound. We perform extensive evaluations across image, audio, video reconstruction, and view synthesis tasks on six diverse datasets, demonstrating our method’s superiority in reconstruction quality and speed over existing MCL and CL-NF approaches. Notably, our approach attains rapid adaptation of neural fields for city-scale NeRF rendering with reduced parameter requirement. Seungyoon Woo, Junhyeog Yun, Gunhee Kim |
ICLR | 3 |
| 2025 | How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary ObjectsabstractMotion synthesis for diverse object categories holds great potential for 3D content creation but remains underexplored due to two key challenges: (1) the lack of comprehensive motion datasets that include a wide range of high-quality motions and annotations, and (2) the absence of methods capable of handling heterogeneous skeletal templates from diverse objects.
To address these challenges, we contribute the following:
First, we augment the Truebones Zoo dataset—a high-quality animal motion dataset covering over 70 species—by annotating it with detailed text descriptions, making it suitable for text-based motion synthesis.
Second, we introduce rig augmentation techniques that generate diverse motion data while preserving consistent dynamics, enabling models to adapt to various skeletal configurations.
Finally, we redesign existing motion diffusion models to dynamically adapt to arbitrary skeletal templates, enabling motion synthesis for a diverse range of objects with varying structures.
Experiments show that our method learns to generate high-fidelity motions from textual descriptions for diverse and even unseen objects, setting a strong foundation for motion synthesis across diverse object categories and skeletal templates.
Qualitative results are available on this [link](https://t2m4lvo.github.io). Wonkwang Lee, Jongwon Jeong, Taehong Moon, Hyeon-Jong Kim, Jaehyeon Kim, Gunhee Kim, Byeong-Uk Lee |
ICML | 6 |
| 2025 | Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language ModelsabstractSehun Lee, Kang-wook Kim, Gunhee Kim. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sehun Lee, Kang-Wook Kim 0004, Gunhee Kim |
NAACL (Long Papers) | 3 |
| 2025 | Is a Peeled Apple Still Red? Evaluating LLMs' Ability for Conceptual Combination with Property TypeabstractSeokwon Song, Taehyun Lee, Jaewoo Ahn, Jae Hyuk Sung, Gunhee Kim. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Seokwon Song, Taehyun Lee, Jaewoo Ahn, Jae Hyuk Sung, Gunhee Kim |
NAACL (Long Papers) | 5 |
| 2025 | Gaze Beyond the Frame: Forecasting Egocentric 3D Visual SpanabstractPeople continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors.
While research in egocentric user and scene understanding has focused primarily on motion and contact-based interaction, forecasting human visual perception itself remains less explored despite its fundamental role in guiding human actions and its implications for AR/VR and assistive technologies.
We address the challenge of egocentric 3D visual span forecasting, predicting where a person's visual perception will focus next within their three-dimensional environment.
To this end, we propose EgoSpanLift, a novel method that transforms egocentric visual span forecasting from 2D image planes to 3D scenes.
EgoSpanLift converts SLAM-derived keypoints into gaze-compatible geometry and extracts volumetric visual span regions.
We further combine EgoSpanLift with 3D U-Net and unidirectional transformers, enabling spatio-temporal fusion to efficiently predict future visual span in the 3D grid.
In addition, we curate a comprehensive benchmark from raw egocentric multisensory data, creating a testbed with 364.6K samples for 3D visual span forecasting.
Our approach outperforms competitive baselines for egocentric gaze anticipation and 3D localization, while achieving comparable results even when projected back onto 2D image planes without additional 2D-specific training. Heeseung Yun, Joonil Na, Jaeyeon Kim, Calvin Murdock, Gunhee Kim |
NeurIPS | 5 |
| 2025 | When Meta-Learning Meets Online and Continual Learning: A SurveyabstractOver the past decade, deep neural networks have demonstrated significant success using the training scheme that involves mini-batch stochastic gradient descent on extensive datasets. Expanding upon this accomplishment, there has been a surge in research exploring the application of neural networks in other learning scenarios. One notable framework that has garnered significant attention is meta-learning. Often described as "learning to learn," meta-learning is a data-driven approach to optimize the learning algorithm. Other branches of interest are continual learning and online learning, both of which involve incrementally updating a model with streaming data. While these frameworks were initially developed independently, recent works have started investigating their combinations, proposing novel problem settings and learning algorithms. However, due to the elevated complexity and lack of unified terminology, discerning differences between the learning frameworks can be challenging even for experienced researchers. To facilitate a clear understanding, this paper provides a comprehensive survey that organizes various problem settings using consistent terminology and formal descriptions. By offering an overview of these learning paradigms, our work aims to foster further advancements in this promising area of research. Jaehyeon Son, Soochan Lee, Gunhee Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?abstractIn the real world, knowledge is constantly evolving, which can render existing knowledge-based datasets outdated.This unreliability highlights the critical need for continuous updates to ensure both accuracy and relevance in knowledgeintensive tasks.To address this, we propose GrowOVER-QA and GrowOVER-Dialogue, dynamic open-domain QA and dialogue benchmarks that undergo a continuous cycle of updates, keeping pace with the rapid evolution of knowledge.Our research indicates that retrieval-augmented language models (RaLMs) struggle with knowledge that has not been trained on or recently updated.Consequently, we introduce a novel retrieval-interactive language model framework, where the language model evaluates and reflects on its answers for further re-retrieval.Our exhaustive experiments demonstrate that our training-free framework significantly improves upon existing methods, performing comparably to or even surpassing continuously trained language models. Dayoon Ko, Hahyeon Choi, Gunhee Kim |
ACL (1) | 4 |
| 2024 | Who Wrote this Code? Watermarking for Code GenerationabstractTaehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, Gunhee Kim. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Taehyun Lee, Seokhee Hong 0002, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, Gunhee Kim |
ACL (1) | 8 |
| 2024 | ESR-NeRF: Emissive Source Reconstruction Using LDR Multi-View ImagesabstractExisting NeRF-based inverse rendering methods suppose that scenes are exclusively illuminated by distant light sources, neglecting the potential influence of emissive sources within a scene. In this work, we confront this limitation using LDR multi-view images captured with emissive sources turned on and off. Two key issues must be addressed: 1) ambiguity arising from the limited dynamic range along with unknown lighting details, and 2) the expensive computational cost in volume rendering to back-trace the paths leading to final object colors. We present a novel approach, ESR-NeRF, leveraging neural networks as learnable functions to represent ray-traced fields. By training networks to satisfy light transport segments, we regulate outgoing radiances, progressively identifying emissive sources while being aware of reflection areas. The results on scenes encompassing emissive sources with various properties demonstrate the superiority of ESR-NeRF in qualitative and quantitative ways. Our approach also extends its applicability to the scenes devoid of emissive sources, achieving lower CD metrics on the DTU dataset. Jinseo Jeong, Junseo Koo, Qimeng Zhang, Gunhee Kim |
CVPR | 4 |
| 2024 | Bi-directional Contextual Attention for 3D Dense Captioning
Minjung Kim 0001, Hyung Suk Lim, Soonyoung Lee, Bumsoo Kim 0005, Gunhee Kim |
ECCV (18) | 5 |
| 2024 | Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar 0003, Jacob Donley, Gunhee Kim, Vamsi K. Ithapu, Calvin Murdock |
ECCV (24) | 7 |
| 2024 | DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAGabstractIn the rapidly evolving landscape of language, resolving new linguistic expressions in continuously updating knowledge bases remains a formidable challenge.This challenge becomes critical in retrieval-augmented generation (RAG) with knowledge bases, as emerging expressions hinder the retrieval of relevant documents, leading to generator hallucinations.To address this issue, we introduce a novel task aimed at resolving emerging mentions to dynamic entities and present DYNAM-ICER benchmark.Our benchmark includes dynamic entity mention resolution and entitycentric knowledge-intensive QA task, evaluating entity linking and RAG model's adaptability to new expressions, respectively.We discovered that current entity linking models struggle to link these new expressions to entities.Therefore, we propose a temporal segmented clustering method with continual adaptation, effectively managing the temporal dynamics of evolving entities and emerging mentions.Extensive experiments demonstrate that our method outperforms existing baselines, enhancing RAG model performance on QA task with resolved mentions. Dayoon Ko, Gunhee Kim |
EMNLP | 3 |
| 2024 | Text2Chart31: Instruction Tuning for Chart Generation with Automatic FeedbackabstractLarge language models (LLMs) have demonstrated strong capabilities across various language tasks, notably through instruction-tuning methods.However, LLMs face challenges in visualizing complex, real-world data through charts and plots.Firstly, existing datasets rarely cover a full range of chart types, such as 3D, volumetric, and gridded charts.Secondly, supervised fine-tuning methods do not fully leverage the intricate relationships within rich datasets, including text, code, and figures.To address these challenges, we propose a hierarchical pipeline and a new dataset for chart generation.Our dataset, Text2Chart31, includes 31 unique plot types referring to the Matplotlib library, with 11.1K tuples of descriptions, code, data tables, and plots.Moreover, we introduce a reinforcement learningbased instruction tuning technique for chart generation tasks without requiring human feedback.Our experiments show that this approach significantly enhances the model performance, enabling smaller models to outperform larger open-source models and be comparable to state-of-the-art proprietary models in data visualization tasks.We make the code and dataset available at https://github. com/fatemehpesaran310/Text2Chart31. Fatemeh Pesaran Zadeh, Jin-Hwa Kim, Gunhee Kim |
EMNLP | 4 |
| 2024 | Compositional Conservatism: A Transductive Approach in Offline Reinforcement LearningabstractOffline reinforcement learning (RL) is a compelling framework for learning optimal policies from past experiences without additional interaction with the environment. Nevertheless, offline RL inevitably faces the problem of distributional shifts, where the states and actions encountered during policy execution may not be in the training dataset distribution. A common solution involves incorporating conservatism into the policy or the value function to safeguard against uncertainties and unknowns. In this work, we focus on achieving the same objectives of conservatism but from a different perspective. We propose COmpositional COnservatism with Anchor-seeking (COCOA) for offline RL, an approach that pursues conservatism in a _compositional_ manner on top of the transductive reparameterization (Netanyahu et al., 2023), which decomposes the input variable (the state in our case) into an anchor and its difference from the original input. Our COCOA seeks both in-distribution anchors and differences by utilizing the learned reverse dynamics model, encouraging conservatism in the compositional input space for the policy or value function. Such compositional conservatism is independent of and agnostic to the prevalent _behavioral_ conservatism in offline RL. We apply COCOA to four state-of-the-art offline RL algorithms and evaluate them on the D4RL benchmark, where COCOA generally improves the performance of each algorithm. The code is available at https://github.com/runamu/compositional-conservatism. Yeda Song, Gunhee Kim |
ICLR | 3 |
| 2024 | Learning to Continually Learn with the Bayesian PrincipleabstractIn the present era of deep learning, continual learning research is mainly focused on mitigating forgetting when training a neural network with stochastic gradient descent on a non-stationary stream of data. On the other hand, in the more classical literature of statistical machine learning, many models have sequential Bayesian update rules that yield the same learning outcome as the batch training, i.e., they are completely immune to catastrophic forgetting. However, they are often overly simple to model complex real-world data. In this work, we adopt the meta-learning paradigm to combine the strong representational power of neural networks and simple statistical models’ robustness to forgetting. In our novel meta-continual learning framework, continual learning takes place only in statistical models via ideal sequential Bayesian update rules, while neural networks are meta-learned to bridge the raw data and the statistical models. Since the neural networks remain fixed during continual learning, they are protected from catastrophic forgetting. This approach not only achieves significantly improved performance but also exhibits excellent scalability. Since our approach is domain-agnostic and model-agnostic, it can be applied to a wide range of problems and easily integrated with existing model architectures. Soochan Lee, Hyeonseong Jeon, Jaehyeon Son, Gunhee Kim |
ICML | 4 |
| 2024 | FedAvP: Augment Local Data via Shared Policy in Federated LearningabstractFederated Learning (FL) allows multiple clients to collaboratively train models without directly sharing their private data. While various data augmentation techniques have been actively studied in the FL environment, most of these methods share input-level or feature-level data information over communication, posing potential privacy leakage. In response to this challenge, we introduce a federated data augmentation algorithm named FedAvP that shares only the augmentation policies, not the data-related information.
For data security and efficient policy search, we interpret the policy loss as a meta update loss in standard FL algorithms and utilize the first-order gradient information to further enhance privacy and reduce communication costs. Moreover, we propose a meta-learning method to search for adaptive personalized policies tailored to heterogeneous clients. Our approach outperforms existing best performing augmentation policy search methods and federated data augmentation methods, in the benchmarks for heterogeneous FL. Minui Hong, Junhyeog Yun, Insu Jeon, Gunhee Kim |
NeurIPS | 4 |
| 2024 | Sample Selection via Contrastive Fragmentation for Noisy Label RegressionabstractAs with many other problems, real-world regression is plagued by the presence of noisy labels, an inevitable issue that demands our attention.
Fortunately, much real-world data often exhibits an intrinsic property of continuously ordered correlations between labels and features, where data points with similar labels are also represented with closely related features.
In response, we propose a novel approach named ConFrag, where we collectively model the regression data by transforming them into disjoint yet contrasting fragmentation pairs.
This enables the training of more distinctive representations, enhancing the ability to select clean samples.
Our ConFrag framework leverages a mixture of neighboring fragments to discern noisy labels through neighborhood agreement among expert feature extractors.
We extensively perform experiments on four newly curated benchmark datasets of diverse domains, including age prediction, price prediction, and music production year estimation.
We also introduce a metric called Error Residual Ratio (ERR) to better account for varying degrees of label noise.
Our approach consistently outperforms fourteen state-of-the-art baselines, being robust against symmetric and random Gaussian label noise. Chris Dongjoo Kim, Sangwoo Moon 0001, Jihwan Moon 0002, Dongyeon Woo, Gunhee Kim |
NeurIPS | 5 |
| 2023 | MPCHAT: Towards Multimodal Persona-Grounded ConversationabstractIn order to build self-consistent personalized dialogue agents, previous research has mostly focused on textual persona that delivers personal facts or personalities.However, to fully describe the multi-faceted nature of persona, image modality can help better reveal the speaker's personal characteristics and experiences in episodic memory (Rubin et al., 2003;Conway, 2009).In this work, we extend persona-based dialogue to the multimodal domain and make two main contributions.First, we present the first multimodal persona-based dialogue dataset named MPCHAT, which extends persona with both text and images to contain episodic memories.Second, we empirically show that incorporating multimodal persona, as measured by three proposed multimodal persona-grounded dialogue tasks (i.e., next response prediction, grounding persona prediction, and speaker identification), leads to statistically significant performance improvements across all tasks.Thus, our work highlights that multimodal persona is crucial for improving multimodal dialogue comprehension, and our MPCHAT serves as a high-quality resource for this research. Jaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee Kim |
ACL (1) | 4 |
| 2023 | SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine CollaborationabstractHwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi, Byoungpil Kim, Gunhee Kim, Eun-Ju Lee, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hwaran Lee, Seokhee Hong 0002, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi 0001, Byoung Pil Kim, Gunhee Kim, Eun-Ju Lee 0001, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha 0001 |
ACL (1) | 8 |
| 2023 | Fusing Pre-Trained Language Models with Multimodal Prompts through Reinforcement LearningabstractLanguage models are capable of commonsense reasoning: while domain-specific models can learn from explicit knowledge (e.g. commonsense graphs [6] ethical norms [25]), and larger models like GPT-3 [7] mani-fest broad commonsense reasoning capacity. Can their knowledge be extended to multimodal inputs such as images and audio without paired domain data? In this work, we propose‡ESPER (Extending Sensory PErception with Reinforcement learning) which enables text-only pretrained models to address multimodal tasks such as visual commonsense reasoning. Our key novelty is to use rein-forcement learning to align multimodal inputs to language model generations without direct supervision: for example, our reward optimization relies only on cosine similarity derived from CLIP [52] and requires no additional paired (image, text) data. Experiments demonstrate that ESPER outperforms baselines and prior work on a variety of multimodal text generation tasks ranging from captioning to commonsense reasoning; these include a new benchmark we collect and release, the ESP dataset, which tasks models with generating the text of several different domains for each image. Our code and data are publicly released at https://github.com/JiwanChung/esper. Youngjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel, Ximing Lu, Rowan Zellers, Prithviraj Ammanabrolu, Ronan Le Bras 0001, Gunhee Kim, Yejin Choi 0001 |
CVPR | 10 |
| 2023 | SODA: Million-scale Dialogue Distillation with Social Commonsense ContextualizationabstractHyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Hyunwoo Kim 0002, Jack Hessel, Peter West, Ximing Lu, Youngjae Yu, Ronan Le Bras 0001, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi 0001 |
EMNLP | 10 |
| 2023 | FANToM: A Benchmark for Stress-testing Machine Theory of Mind in InteractionsabstractTheory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity.We introduce FANTOM, a new benchmark designed to stress-test ToM within information-asymmetric conversational contexts via question answering.Our benchmark draws upon important theoretical requisites from psychology and necessary empirical considerations when evaluating large language models (LLMs).In particular, we formulate multiple types of questions that demand the same underlying reasoning to identify illusory or false sense of ToM capabilities in LLMs.We show that FANTOM is challenging for state-of-the-art LLMs, which perform significantly worse than humans even with chainof-thought reasoning or fine-tuning.1 Linda: Yeah, I got a golden retriever.She's so adorable.David: What's her favorite food?Kailey: Hey guys, I' Hyunwoo Kim 0002, Melanie Sclar, Ronan Le Bras 0001, Gunhee Kim, Yejin Choi 0001, Maarten Sap |
EMNLP | 5 |
| 2023 | Can Language Models Laugh at YouTube Short-form Videos?abstractAs short-form funny videos on social networks are gaining popularity, it becomes demanding for AI models to understand them for better communication with humans.Unfortunately, previous video humor datasets target specific domains such as speeches or sitcoms, and mostly focus on verbal cues.We curate a usergenerated dataset of 10K multimodal funny videos from YouTube, called ExFunTube.Using a video filtering pipeline with GPT-3.5, we verify both verbal and visual elements contributing to humor.After filtering, we annotate each video with timestamps and text explanations for funny moments.Our ExFunTube is unique over existing datasets in that our videos cover a wide range of domains with various types of humor that necessitate a multimodal understanding of the content.Also, we develop a zero-shot video-to-text prompting to maximize video humor understanding of large language models (LLMs).With three different evaluation methods using automatic scores, rationale quality experiments, and human evaluations, we show that our prompting significantly improves LLMs' ability for humor explanation. Dayoon Ko, Sangho Lee 0008, Gunhee Kim |
EMNLP | 3 |
| 2023 | mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesabstractThe growing number of multimodal online discussions necessitates automatic summarization to save time and reduce content overload.However, existing summarization datasets are not suitable for this purpose, as they either do not cover discussions, multiple modalities, or both.To this end, we present MREDDITSUM, the first multimodal discussion summarization dataset.It consists of 3,033 discussion threads where a post solicits advice regarding an issue described with an image and text, and respective comments express diverse opinions.We annotate each thread with a human-written summary that captures both the essential information from the text, as well as the details available only in the image.Experiments show that popular summarization models-GPT-3.5,BART, and T5-consistently improve in performance when visual information is incorporated.We also introduce a novel method, cluster-based multi-stage summarization, that outperforms existing baselines and serves as a competitive baseline for future work. * Equal contribution. Keighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park, Gunhee Kim |
EMNLP | 5 |
| 2023 | EP2P-Loc: End-to-End 3D Point to 2D Pixel Localization for Large-Scale Visual LocalizationabstractVisual localization is the task of estimating a 6-DoF camera pose of a query image within a provided 3D reference map. Thanks to recent advances in various 3D sensors, 3D point clouds are becoming a more accurate and affordable option for building the reference map, but research to match the points of 3D point clouds with pixels in 2D images for visual localization remains challenging. Existing approaches that jointly learn 2D-3D feature matching suffer from low inliers due to representational differences between the two modalities, and the methods that bypass this problem into classification have an issue of poor refinement. In this work, we propose EP2P-Loc, a novel large-scale visual localization method that mitigates such appearance discrepancy and enables end-to-end training for pose estimation. To increase the number of inliers, we propose a simple algorithm to remove invisible 3D points in the image, and find all 2D-3D correspondences without keypoint detection. To reduce memory usage and search complexity, we take a coarse-to-fine approach where we extract patch-level features from 2D images, then perform 2D patch classification on each 3D point, and obtain the exact corresponding 2D pixel coordinates through positional encoding. Finally, for the first time in this task, we employ a differentiable PnP for end-to-end training. In the experiments on newly curated large-scale indoor and outdoor benchmarks based on 2D-3D-S and KITTI, we show that our method achieves the state-of-the-art performance compared to existing visual localization and image-to-point cloud registration methods. Minjung Kim 0001, Junseo Koo, Gunhee Kim |
ICCV | 3 |
| 2023 | Dense 2D-3D Indoor Prediction with Sound via Aligned Cross-Modal DistillationabstractSound can convey significant information for spatial reasoning in our daily lives. To endow deep networks with such ability, we address the challenge of dense indoor prediction with sound in both 2D and 3D via cross-modal knowledge distillation. In this work, we propose a Spatial Alignment via Matching (SAM) distillation framework that elicits local correspondence between the two modalities in vision-to-audio knowledge transfer. SAM integrates audio features with visually coherent learnable spatial embeddings to resolve inconsistencies in multiple layers of a student model. Our approach does not rely on a specific input representation, allowing for flexibility in the input shapes or dimensions without performance degradation. With a newly curated benchmark named Dense Auditory Prediction of Surroundings (DAPS), we are the first to tackle dense indoor prediction of omnidirectional surroundings in both 2D and 3D with audio observations. Specifically, for audio-based depth estimation, semantic segmentation, and challenging 3D scene reconstruction, the proposed distillation framework consistently achieves state-of-the-art performance across various metrics and backbone architectures. Heeseung Yun, Joonil Na, Gunhee Kim |
ICCV | 3 |
| 2023 | Federated Learning via Meta-Variational DropoutabstractFederated Learning (FL) aims to train a global inference model from remotely distributed clients, gaining popularity due to its benefit of improving data privacy. However, traditional FL often faces challenges in practical applications, including model overfitting and divergent local models due to limited and non-IID data among clients. To address these issues, we introduce a novel Bayesian meta-learning approach called meta-variational dropout (MetaVD). MetaVD learns to predict client-dependent dropout rates via a shared hypernetwork, enabling effective model personalization of FL algorithms in limited non-IID data settings. We also emphasize the posterior adaptation view of meta-learning and the posterior aggregation view of Bayesian FL via the conditional dropout posterior. We conducted extensive experiments on various sparse and non-IID FL datasets. MetaVD demonstrated excellent classification accuracy and uncertainty calibration performance, especially for out-of-distribution (OOD) clients. MetaVD compresses the local model parameters needed for each client, mitigating model overfitting and reducing communication costs. Code is available at https://github.com/insujeon/MetaVD. Insu Jeon, Minui Hong, Junhyeog Yun, Gunhee Kim |
NeurIPS | 4 |
| 2023 | Benchmark of Machine Learning Force Fields for Semiconductor Simulations: Datasets, Metrics, and Comparative AnalysisabstractAs semiconductor devices become miniaturized and their structures become more complex, there is a growing need for large-scale atomic-level simulations as a less costly alternative to the trial-and-error approach during development.Although machine learning force fields (MLFFs) can meet the accuracy and scale requirements for such simulations, there are no open-access benchmarks for semiconductor materials.Hence, this study presents a comprehensive benchmark suite that consists of two semiconductor material datasets and ten MLFF models with six evaluation metrics. We select two important semiconductor thin-film materials silicon nitride and hafnium oxide, and generate their datasets using computationally expensive density functional theory simulations under various scenarios at a cost of 2.6k GPU days.Additionally, we provide a variety of architectures as baselines: descriptor-based fully connected neural networks and graph neural networks with rotational invariant or equivariant features.We assess not only the accuracy of energy and force predictions but also five additional simulation indicators to determine the practical applicability of MLFF models in molecular dynamics simulations.To facilitate further research, our benchmark suite is available at https://github.com/SAITPublic/MLFF-Framework. Geonu Kim, Byunggook Na, Gunhee Kim, Hyuntae Cho, Seungjin Kang, Hee Sun Lee, Saerom Choi, Heejae Kim, Yongdeok Kim |
NeurIPS | 3 |
| 2023 | Recasting Continual Learning as Sequence ModelingabstractIn this work, we aim to establish a strong connection between two significant bodies of machine learning research: continual learning and sequence modeling.
That is, we propose to formulate continual learning as a sequence modeling problem, allowing advanced sequence models to be utilized for continual learning.
Under this formulation, the continual learning process becomes the forward pass of a sequence model.
By adopting the meta-continual learning (MCL) framework, we can train the sequence model at the meta-level, on multiple continual learning episodes.
As a specific example of our new formulation, we demonstrate the application of Transformers and their efficient variants as MCL methods.
Our experiments on seven benchmarks, covering both classification and regression, show that sequence models can be an attractive solution for general MCL. Soochan Lee, Jaehyeon Son, Gunhee Kim |
NeurIPS | 3 |
| 2022 | On Convergence of Lookahead in Smooth GamesabstractA key challenge in smooth games is that there is no general guarantee for gradient methods to converge to an equilibrium. Recently, Chavdarova et al. (2021) reported a promising empirical observation that Lookahead (Zhang et al., 2019) significantly improves GAN training. While promising, few theoretical guarantees has been studied for Lookahead in smooth games. In this work, we establish the first convergence guarantees of Lookahead for smooth games. We present a spectral analysis and provide a geometric explanation of how and when it actually improves the convergence around a stationary point. Based on the analysis, we derive sufficient conditions for Lookahead to stabilize or accelerate the local convergence in smooth games. Our study reveals that Lookahead provides a general mechanism for stabilization and acceleration in smooth games. Junsoo Ha, Gunhee Kim |
AISTATS | 2 |
| 2022 | Panoramic Vision Transformer for Saliency Detection in 360$^\circ $ Videos
Heeseung Yun, Sehun Lee, Gunhee Kim |
ECCV (35) | 3 |
| 2022 | ProsocialDialog: A Prosocial Backbone for Conversational AgentsabstractMost existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them.To address this issue, we introduce PROSOCIALDIALOG, the first large-scale multi-turn dialogue dataset to teach conversational agents to respond to problematic content following social norms.Covering diverse unethical, problematic, biased, and toxic situations, PROSOCIALDIALOG contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-ofthumb, RoTs).Created via a human-AI collaborative framework, PROSOCIALDIALOG consists of 58K dialogues, with 331K utterances, 160K unique RoTs, and 497K dialogue safety labels accompanied by free-form rationales.With this dataset, we introduce a dialogue safety detection module, Canary, capable of generating RoTs given conversational context, and a socially-informed dialogue agent, Prost.Empirical results show that Prost generates more socially acceptable dialogues compared to other state-of-the-art language and dialogue models in both in-domain and out-of-domain settings.Additionally, Canary effectively guides off-the-shelf language models to generate significantly more prosocial responses.Our work highlights the promise and importance of creating and steering conversational AI to be socially responsible. Hyunwoo Kim 0002, Youngjae Yu, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi 0001, Maarten Sap |
EMNLP | 6 |
| 2022 | Neural Variational Dropout Processes
Insu Jeon, Youngjin Park, Gunhee Kim |
ICLR | 3 |
| 2022 | Lipschitz-constrained Unsupervised Skill Discovery
Seohong Park, Jaekyeom Kim, Honglak Lee, Gunhee Kim |
ICLR | 5 |
| 2022 | Constrained GPI for Zero-Shot Transfer in Reinforcement LearningabstractFor zero-shot transfer in reinforcement learning where the reward function varies between different tasks, the successor features framework has been one of the popular approaches. However, in this framework, the transfer to new target tasks with generalized policy improvement (GPI) relies on only the source successor features [5] or additional successor features obtained from the function approximators’ generalization to novel inputs [11]. The goal of this work is to improve the transfer by more tightly bounding the value approximation errors of successor features on the new target tasks. Given a set of source tasks with their successor features, we present lower and upper bounds on the optimal values for novel task vectors that are expressible as linear combinations of source task vectors. Based on the bounds, we propose constrained GPI as a simple test-time approach that can improve transfer by constraining action-value approximation errors on new target tasks. Through experiments in the Scavenger and Reacher environment with state observations as well as the DeepMind Lab environment with visual observations, we show that the proposed constrained GPI significantly outperforms the prior GPI’s transfer performance. Our code and additional information are available at https://jaekyeom.github.io/projects/cgpi/. Jaekyeom Kim, Seohong Park, Gunhee Kim |
NeurIPS | 3 |
| 2022 | Guest Editorial Introduction to the Special Section on Video and LanguageabstractComputer Vision (CV) and Natural Language Processing (NLP) are two most fundamental disciplines under a broad area of artificial intelligence (AI). CV is regarded as a field of research that explores the techniques to teach computers to see and understand digital content such as images and videos. NLP is a branch of linguistics that enables computers to process, interpret, and even generate human language. With the rise and development of deep learning over the past decade, there has been a steady momentum of innovation and breakthroughs that convincingly push the limits and improve the state-of-the-art of both vision and language modeling. An interesting observation is that the research in the two areas starts to interact, with a significant growth in both the volume of publications and extensive applications. Meanwhile, many previous experiences have shown that this can naturally build up the circle of human intelligence. Tao Mei 0001, Jason J. Corso, Gunhee Kim, Jiebo Luo 0001, Chunhua Shen, Hanwang Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Dual Compositional Learning in Interactive Image RetrievalabstractWe present an approach named Dual Composition Network (DCNet) for interactive image retrieval that searches for the best target image for a natural language query and a reference image. To accomplish this task, existing methods have focused on learning a composite representation of the reference image and the text query to be as close to the embedding of the target image as possible. We refer this approach as Composition Network. In this work, we propose to close the loop with Correction Network that models the difference between the reference and target image in the embedding space and matches it with the embedding of the text query. That is, we consider two cyclic directional mappings for triplets of (reference image, text query, target image) by using both Composition Network and Correction Network. We also propose a joint training loss that can further improve the robustness of multimodal representation learning. We evaluate the proposed model on three benchmark datasets for multimodal retrieval: Fashion-IQ, Shoes, and Fashion200K. Our experiments show that our DCNet achieves new state-of-the-art performance on all three datasets, and the addition of Correction Network consistently improves multiple existing methods that are solely based on Composition Network. Moreover, an ensemble of our model won the first place in Fashion-IQ 2020 challenge held in a CVPR 2020 workshop. Jongseok Kim 0002, Youngjae Yu, Hoeseong Kim, Gunhee Kim |
AAAI | 4 |
| 2021 | IB-GAN: Disentangled Representation Learning with Information Bottleneck Generative Adversarial NetworksabstractWe propose a new GAN-based unsupervised model for disentangled representation learning. The new model is discovered in an attempt to utilize the Information Bottleneck (IB) framework to the optimization of GAN, thereby named IB-GAN. The architecture of IB-GAN is partially similar to that of InfoGAN but has a critical difference; an intermediate layer of the generator is leveraged to constrain the mutual information between the input and the generated output. The intermediate stochastic layer can serve as a learnable latent distribution that is trained with the generator jointly in an end-to-end fashion. As a result, the generator of IB-GAN can harness the latent space in a disentangled and interpretable manner. With the experiments on dSprites and Color-dSprites dataset, we demonstrate that IB-GAN achieves competitive disentanglement scores to those of state-of-the-art β-VAEs and outperforms InfoGAN. Moreover, the visual quality and the diversity of samples generated by IB-GAN are often better than those by β-VAEs and Info-GAN in terms of FID score on CelebA and 3D Chairs dataset. Insu Jeon, Wonkwang Lee, Myeongjang Pyeon, Gunhee Kim |
AAAI | 4 |
| 2021 | StyleMix: Separating Content and Style for Enhanced Data AugmentationabstractIn spite of the great success of deep neural networks for many challenging classification tasks, the learned networks are vulnerable to overfitting and adversarial attacks. Recently, mixup based augmentation methods have been actively studied as one practical remedy for these drawbacks. However, these approaches do not distinguish between the content and style features of the image, but mix or cut-and-paste the images. We propose StyleMix and StyleCutMix as the first mixup method that separately manipulates the content and style information of input image pairs. By carefully mixing up the content and style of images, we can create more abundant and robust samples, which eventually enhance the generalization of model training. We also develop an automatic scheme to decide the degree of style mixing according to the pair’s class distance, to prevent messy mixed images from too differently styled pairs. Our experiments on CIFAR-10, CIFAR-100 and ImageNet datasets show that StyleMix achieves better or comparable performance to state of the art mixup methods and learns more robust classifiers to adversarial attacks. Minui Hong, Gunhee Kim |
CVPR | 3 |
| 2021 | Transitional Adaptation of Pretrained Models for Visual StorytellingabstractPrevious models for vision-to-language generation tasks usually pretrain a visual encoder and a language generator in the respective domains and jointly finetune them with the target task. However, this direct transfer practice may suffer from the discord between visual specificity and language fluency since they are often separately trained from large corpora of visual and text data with no common ground. In this work, we claim that a transitional adaptation task is required between pretraining and finetuning to harmonize the visual encoder and the language model for challenging downstream target tasks like visual storytelling. We propose a novel approach named Transitional Adaptation of Pre-trained Model (TAPM) that adapts the multi-modal modules to each other with a simpler alignment task between visual inputs only with no need for text labels. Through extensive experiments, we show that the adaptation step significantly improves the performance of multiple language models for sequential video and image captioning tasks. We achieve new state-of-the-art performance on both language metrics and human evaluation in the multi-sentence description task of LSMDC 2019 [50] and the image storytelling task of VIST [18]. Our experiments reveal that this improvement in caption quality does not depend on the specific choice of language models. Youngjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim 0002, Gunhee Kim |
CVPR | 5 |
| 2021 | Perspective-taking and Pragmatics for Generating Empathetic Responses Focused on Emotion CausesabstractEmpathy is a complex cognitive ability based on the reasoning of others' affective states.In order to better understand others and express stronger empathy in dialogues, we argue that two issues must be tackled at the same time: (i) identifying which word is the cause for the other's emotion from his or her utterance and (ii) reflecting those specific words in the response generation.However, previous approaches for recognizing emotion cause words in text require sub-utterance level annotations, which can be demanding.Taking inspiration from social cognition, we leverage a generative estimator to infer emotion cause words from utterances with no word-level label.Also, we introduce a novel method based on pragmatics to make dialogue models focus on targeted words in the input during generation.Our method is applicable to any dialogue models with no additional training on the fly.We show our approach improves multiple best performing dialogue agents on generating more focused empathetic responses in terms of both automatic and human evaluation. Hyunwoo Kim 0002, Byeongchang Kim 0002, Gunhee Kim |
EMNLP (1) | 3 |
| 2021 | Continual Learning on Noisy Data Streams via Self-Purified ReplayabstractContinually learning in the real world must overcome many challenges, among which noisy labels are a common and inevitable issue. In this work, we present a replay-based continual learning framework that simultaneously addresses both catastrophic forgetting and noisy labels for the first time. Our solution is based on two observations; (i) forgetting can be mitigated even with noisy labels via self-supervised learning, and (ii) the purity of the replay buffer is crucial. Building on this regard, we propose two key components of our method: (i) a self-supervised replay technique named Self-Replay which can circumvent erroneous training signals arising from noisy labeled data, and (ii) the Self-Centered filter that maintains a purified replay buffer via centrality-based stochastic graph ensembles. The empirical results on MNIST, CIFAR-10, CIFAR-100, and WebVision with real-world noise demonstrate that our framework can maintain a highly pure replay buffer amidst noisy streamed data while greatly outperforming the combinations of the state-of-the-art continual learning and noisy label learning methods. Chris Dongjoo Kim, Jinseo Jeong, Sangwoo Moon 0001, Gunhee Kim |
ICCV | 4 |
| 2021 | Viewpoint-Agnostic Change Captioning with Cycle ConsistencyabstractChange captioning is the task of identifying the change and describing it with a concise caption. Despite recent advancements, filtering out insignificant changes still remains as a challenge. Namely, images from different camera perspectives can cause issues; a mere change in viewpoint should be disregarded while still capturing the actual changes. In order to tackle this problem, we present a new Viewpoint-Agnostic change captioning network with Cycle Consistency (VACC) that requires only one image each for the before and after scene, without depending on any other information. We achieve this by devising a new difference encoder module which can encode viewpoint information and model the difference more effectively. In addition, we propose a cycle consistency module that can potentially improve the performance of any change captioning networks in general by matching the composite feature of the generated caption and before image with the after image feature. We evaluate the performance of our proposed model across three datasets for change captioning, including a novel dataset we introduce here that contains images with changes under extreme viewpoint shifts. Through our experiments, we show the excellence of our method with respect to the CIDEr, BLEU-4, METEOR and SPICE scores. Moreover, we demonstrate that attaching our proposed cycle consistency module yields a performance boost for existing change captioning networks, even with varying image encoding mechanisms. Hoeseong Kim, Hyungseok Lee, Hyunsung Park, Gunhee Kim |
ICCV | 5 |
| 2021 | ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation LearningabstractThe natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the ever-growing amount of online videos an attractive source of training data. However, large portions of online videos contain irrelevant audio-visual signals because of edited/overdubbed audio, and models trained on such uncurated videos have shown to learn suboptimal representations. Therefore, existing self-supervised approaches rely on datasets with predetermined taxonomies of semantic concepts, where there is a high chance of audio-visual correspondence. Unfortunately, constructing such datasets require labor intensive manual annotation and/or verification, which severely limits the utility of online videos for large-scale learning. In this work, we present an automatic dataset curation approach based on subset optimization where the objective is to maximize the mutual information between audio and visual channels in videos. We demonstrate that our approach finds videos with high audio-visual correspondence and show that self-supervised models trained on our data achieve competitive performances compared to models trained on existing manually curated datasets. The most significant benefit of our approach is scalability: We release ACAV100M that contains 100 million videos with high audio-visual correspondence, ideal for self-supervised video representation learning. Sangho Lee 0008, Jiwan Chung, Youngjae Yu, Gunhee Kim, Thomas M. Breuel, Gal Chechik, Yale Song |
ICCV | 4 |
| 2021 | Pano-AVQA: Grounded Audio-Visual Question Answering on 360° Videosabstract360° videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond predetermined normal field of views and displays distinctive spatial relations on a sphere. However, previous benchmark tasks for panoramic videos are still limited to evaluate the semantic understanding of audio-visual relationships or spherical spatial property in surroundings. We propose a novel benchmark named Pano-AVQA as a large-scale grounded audio-visual question answering dataset on panoramic videos. Using 5.4K 360° video clips harvested online, we collect two types of novel question-answer pairs with bounding-box grounding: spherical spatial relation QAs and audio-visual relation QAs. We train several transformer-based models from Pano-AVQA, where the results suggest that our proposed spherical spatial embeddings and multimodal training objectives fairly contribute to a better semantic understanding of the panoramic surroundings on the dataset. Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 0005, Gunhee Kim |
ICCV | 5 |
| 2021 | Contextual Label Transformation For Scene Graph Generation
Sungeun Kim, Gunhee Kim |
ICIP | 3 |
| 2021 | Drop-Bottleneck: Learning Discrete Compressed Representation for Noise-Robust Exploration
Jaekyeom Kim, Minjung Kim 0001, Dongyeon Woo, Gunhee Kim |
ICLR | 4 |
| 2021 | Parameter Efficient Multimodal Transformers for Video Representation Learning
Sangho Lee 0008, Youngjae Yu, Gunhee Kim, Thomas M. Breuel, Jan Kautz, Yale Song |
ICLR | 3 |
| 2021 | SEDONA: Search for Decoupled Neural Networks toward Greedy Block-wise Learning
Myeongjang Pyeon, Jihwan Moon 0002, Taeyoung Hahn, Gunhee Kim |
ICLR | 4 |
| 2021 | Self-Supervised Learning of Compressed Video Representations
Youngjae Yu, Sangho Lee 0008, Gunhee Kim, Yale Song |
ICLR | 3 |
| 2021 | Unsupervised Skill Discovery with Bottleneck Option LearningabstractHaving the ability to acquire inherent skills from environments without any external rewards or supervision like humans is an important problem. We propose a novel unsupervised skill discovery method named Information Bottleneck Option Learning (IBOL). On top of the linearization of environments that promotes more various and distant state transitions, IBOL enables the discovery of diverse skills. It provides the abstraction of the skills learned with the information bottleneck framework for the options with improved stability and encouraged disentanglement. We empirically demonstrate that IBOL outperforms multiple state-of-the-art unsupervised skill discovery methods on the information-theoretic evaluations and downstream tasks in MuJoCo environments, including Ant, HalfCheetah, Hopper and D’Kitty. Our code is available at https://vision.snu.ac.kr/projects/ibol. Jaekyeom Kim, Seohong Park, Gunhee Kim |
ICML | 3 |
| 2021 | Unsupervised Representation Learning via Neural Activation CodingabstractWe present neural activation coding (NAC) as a novel approach for learning deep representations from unlabeled data for downstream applications. We argue that the deep encoder should maximize its nonlinear expressivity on the data for downstream predictors to take full advantage of its representation power. To this end, NAC maximizes the mutual information between activation patterns of the encoder and the data over a noisy communication channel. We show that learning for a noise-robust activation code increases the number of distinct linear regions of ReLU encoders, hence the maximum nonlinear expressivity. More interestingly, NAC learns both continuous and discrete representations of data, which we respectively evaluate on two downstream tasks: (i) linear classification on CIFAR-10 and ImageNet-1K and (ii) nearest neighbor retrieval on CIFAR-10 and FLICKR-25K. Empirical results show that NAC attains better or comparable performance on both tasks over recent baselines including SimCLR and DistillHash. In addition, NAC pretraining provides significant benefits to the training of deep generative models. Our code is available at https://github.com/yookoon/nac. Yookoon Park, Sangho Lee 0008, Gunhee Kim, David M. Blei |
ICML | 3 |
| 2021 | How Robust are Fact Checking Systems on Colloquial Claims?abstractByeongchang Kim, Hyunwoo Kim, Seokhee Hong, Gunhee Kim. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Byeongchang Kim 0002, Hyunwoo Kim 0002, Seokhee Hong 0002, Gunhee Kim |
NAACL-HLT | 4 |
| 2021 | Time Discretization-Invariant Safe Action Repetition for Policy Gradient MethodsabstractIn reinforcement learning, continuous time is often discretized by a time scale $\delta$, to which the resulting performance is known to be highly sensitive. In this work, we seek to find a $\delta$-invariant algorithm for policy gradient (PG) methods, which performs well regardless of the value of $\delta$. We first identify the underlying reasons that cause PG methods to fail as $\delta \to 0$, proving that the variance of the PG estimator can diverge to infinity in stochastic environments under a certain assumption of stochasticity. While durative actions or action repetition can be employed to have $\delta$-invariance, previous action repetition methods cannot immediately react to unexpected situations in stochastic environments. We thus propose a novel $\delta$-invariant method named Safe Action Repetition (SAR) applicable to any existing PG algorithm. SAR can handle the stochasticity of environments by adaptively reacting to changes in states during action repetition. We empirically show that our method is not only $\delta$-invariant but also robust to stochasticity, outperforming previous $\delta$-invariant approaches on eight MuJoCo environments with both deterministic and stochastic settings. Our code is available at https://vision.snu.ac.kr/projects/sar. Seohong Park, Jaekyeom Kim, Gunhee Kim |
NeurIPS | 3 |
| 2020 | Rethinking Class Activation Mapping for Weakly Supervised Object Localization
Wonho Bae, Junhyug Noh, Gunhee Kim |
ECCV (15) | 3 |
| 2020 | Imbalanced Continual Learning with Partitioning Reservoir Sampling
Chris Dongjoo Kim, Jinseo Jeong, Gunhee Kim |
ECCV (13) | 3 |
| 2020 | Model-Agnostic Boundary-Adversarial Sampling for Test-Time Generalization in Few-Shot Learning
Jaekyeom Kim, Hyoungseok Kim, Gunhee Kim |
ECCV (1) | 3 |
| 2020 | Character Grounding and Re-identification in Story of Videos and Text Descriptions
Youngjae Yu, Jongseok Kim 0002, Heeseung Yun, Jiwan Chung, Gunhee Kim |
ECCV (5) | 5 |
| 2020 | Will I Sound Like Me? Improving Persona Consistency in Dialogues through Pragmatic Self-ConsciousnessabstractWe explore the task of improving persona consistency of dialogue agents.Recent models tackling consistency often train with additional Natural Language Inference (NLI) labels or attach trained extra modules to the generative agent for maintaining consistency.However, such additional labels and training can be demanding.Also, we find even the bestperforming persona-based agents are insensitive to contradictory words.Inspired by social cognition and pragmatics, we endow existing dialogue agents with public self-consciousness on the fly through an imaginary listener.Our approach, based on the Rational Speech Acts framework (Frank and Goodman, 2012), can enforce dialogue agents to refrain from uttering contradiction.We further extend the framework by learning the distractor selection, which has been usually done manually or randomly.Results on Dialogue NLI (Welleck et al., 2019) and PersonaChat (Zhang et al., 2018) dataset show that our approach reduces contradiction and improves consistency of existing dialogue models.Moreover, we show that it can be generalized to improve contextconsistency beyond persona in dialogues. Hyunwoo Kim 0002, Byeongchang Kim 0002, Gunhee Kim |
EMNLP (1) | 3 |
| 2020 | Sequential Latent Knowledge Selection for Knowledge-Grounded Dialogue
Byeongchang Kim 0002, Jaewoo Ahn, Gunhee Kim |
ICLR | 3 |
| 2020 | A Neural Dirichlet Process Mixture Model for Task-Free Continual Learning
Soochan Lee, Junsoo Ha, Dongsu Zhang, Gunhee Kim |
ICLR | 4 |
| 2019 | Multi-Task Self-Supervised Object Detection via Recycling of Bounding Box AnnotationsabstractIn spite of recent enormous success of deep convolutional networks in object detection, they require a large amount of bounding box annotations, which are often time-consuming and error-prone to obtain. To make better use of given limited labels, we propose a novel object detection approach that takes advantage of both multi-task learning (MTL) and self-supervised learning (SSL). We propose a set of auxiliary tasks that help improve the accuracy of object detection. They create their own labels by recycling the bounding box labels (i.e. annotations of the main task) in an SSL manner, and are jointly trained with the object detection model in an MTL way. Our approach is integrable with any region proposal based detection models. We empirically validate that our approach effectively improves detection performance on various architectures and datasets. We test two state-of-the-art region proposal object detectors, including Faster R-CNN and R-FCN, with three CNN backbones of ResNet-101, Inception-ResNet-v2, and MobileNet on two benchmark datasets of PASCAL VOC and COCO. Joonil Na, Gunhee Kim |
CVPR | 3 |
| 2019 | Better to Follow, Follow to Be Better: Towards Precise Supervision of Feature Super-Resolution for Small Object DetectionabstractIn spite of recent success of proposal-based CNN models for object detection, it is still difficult to detect small objects due to the limited and distorted information that small region of interests (RoI) contain. One way to alleviate this issue is to enhance the features of small RoIs using a super-resolution (SR) technique. We investigate how to improve feature-level super-resolution especially for small object detection, and discover its performance can be significantly improved by (i) utilizing proper high-resolution target features as supervision signals for training of a SR model and (ii) matching the relative receptive fields of training pairs of input low-resolution features and target high-resolution features. We propose a novel feature-level super-resolution approach that not only correctly addresses these two desiderata but also is integrable with any proposal-based detectors with feature pooling. In our experiments, our approach significantly improves the performance of Faster R-CNN on three benchmarks of Tsinghua-Tencent 100K, PASCAL VOC and MS COCO. The improvement for small objects is remarkably large, and encouragingly, those for medium and large objects are nontrivial too. As a result, we achieve new state-of-the-art performance on Tsinghua-Tencent 100K and highly competitive results on both PASCAL VOC and MS COCO. Junhyug Noh, Wonho Bae, Jinhwan Seo, Gunhee Kim |
ICCV | 5 |
| 2019 | Automating System Configuration of Distributed Machine LearningabstractThe performance of distributed machine learning systems is dependent on their system configuration. However, configuring the system for optimal performance is challenging and time consuming even for experts due to the diverse runtime factors such as workloads or the system environment. We present cost-based optimization to automatically find a good system configuration for parameter server (PS) machine learning (ML) frameworks. We design and implement Cruise that applies the optimization technique to tune distributed PS ML execution automatically. Evaluation results on three ML applications verify that Cruise automates the system configuration of the applications to achieve good performance with minor reconfiguration costs. Woo-Yeon Lee, Markus Weimer, Byung-Gon Chun, Yunseong Lee, Joo Seong Jeong, Gyeong-In Yu, Hojin Park, Beomyeol Jeon, Won Wook Song, Gunhee Kim |
ICDCS | 12 |
| 2019 | Harmonizing Maximum Likelihood with GANs for Multimodal Conditional Generation
Soochan Lee, Junsoo Ha, Gunhee Kim |
ICLR (Poster) | 3 |
| 2019 | Discovery of Natural Language Concepts in Individual Units of CNNs
Seil Na, Yo Joong Choe, Gunhee Kim |
ICLR (Poster) | 4 |
| 2019 | Curiosity-Bottleneck: Exploration By Distilling Task-Specific NoveltyabstractExploration based on state novelty has brought great success in challenging reinforcement learning problems with sparse rewards. However, existing novelty-based strategies become inefficient in real-world problems where observation contains not only task-dependent state novelty of our interest but also task-irrelevant information that should be ignored. We introduce an information- theoretic exploration strategy named Curiosity-Bottleneck that distills task-relevant information from observation. Based on the information bottleneck principle, our exploration bonus is quantified as the compressiveness of observation with respect to the learned representation of a compressive value network. With extensive experiments on static image classification, grid-world and three hard-exploration Atari games, we show that Curiosity-Bottleneck learns an effective exploration strategy by robustly measuring the state novelty in distractive environments where state-of-the-art exploration methods often degenerate. Wontae Nam, Hyunwoo Kim 0002, Gunhee Kim |
ICML | 5 |
| 2019 | Variational Laplace AutoencodersabstractVariational autoencoders employ an amortized inference model to approximate the posterior of latent variables. However, such amortized variational inference faces two challenges: (1) the limited posterior expressiveness of fully-factorized Gaussian assumption and (2) the amortization error of the inference model. We present a novel approach that addresses both challenges. First, we focus on ReLU networks with Gaussian output and illustrate their connection to probabilistic PCA. Building on this observation, we derive an iterative algorithm that finds the mode of the posterior and apply fullcovariance Gaussian posterior approximation centered on the mode. Subsequently, we present a general framework named Variational Laplace Autoencoders (VLAEs) for training deep generative models. Based on the Laplace approximation of the latent variable posterior, VLAEs enhance the expressiveness of the posterior while reducing the amortization error. Empirical results on MNIST, Omniglot, Fashion-MNIST, SVHN and CIFAR10 show that the proposed approach significantly outperforms other recent amortized or iterative methods on the ReLU networks. Yookoon S. Park, Chris Dongjoo Kim, Gunhee Kim |
ICML | 3 |
| 2019 | Self-Routing Capsule NetworksabstractCapsule networks have recently gained a great deal of interest as a new architecture of neural networks that can be more robust to input perturbations than similar-sized CNNs. Capsule networks have two major distinctions from the conventional CNNs: (i) each layer consists of a set of capsules that specialize in disjoint regions of the feature space and (ii) the routing-by-agreement coordinates connections between adjacent capsule layers. Although the routing-by-agreement is capable of filtering out noisy predictions of capsules by dynamically adjusting their influences, its unsupervised clustering nature causes two weaknesses: (i) high computational complexity and (ii) cluster assumption that may not hold in presence of heavy input noise. In this work, we propose a novel and surprisingly simple routing strategy called self-routing where each capsule is routed independently by its subordinate routing network. Therefore, the agreement between capsules is not required anymore but both poses and activations of upper-level capsules are obtained in a way similar to Mixture-of-Experts. Our experiments on CIFAR-10, SVHN and SmallNORB show that the self-routing performs more robustly against white-box adversarial attacks and affine transformations, requiring less computation. Taeyoung Hahn, Myeongjang Pyeon, Gunhee Kim |
NeurIPS | 3 |
| 2019 | POL360: A Universal Mobile VR Motion Controller using Polarized LightabstractWe introduce POL360: the first universal VR motion controller that leverages the principle of light polarization. POL360 enables a user who holds it and wears a VR headset to see their hand motion in a virtual world via its accurate 6-DOF position tracking. Compared to other techniques for VR positioning, POL360 has several advantages as follows. (1) Mobile compatibility: Neither additional computing resource like a PC/console nor any complicated pre-installation is required in the environment. Only necessary device is a VR headset with an IR LED module as a light source to which a thin-film linear polarizer is attached. (2) On-device computing: Our POL360’s computation for positioning is completed on the microprocessor in the device. Thus, it does not require additional computing resource of a VR headset. (3) Competitive accuracy and update rate: In spite of POL360’s superior mobile compatibility and affordability, POL360 attains competitive performance of accuracy and fast update rates. That is, it achieves the subcentimeter accuracy of positioning and the tracking rate higher than 60 Hz. In this paper, we derive the mathematical formulation of 6-DOF positioning using light polarization for the first time and implement a POL360 prototype that can directly operate with any commercial VR headset systems. In order to demonstrate POL360’s performance and usability, we carry out thorough quantitative evaluation and a user study and develop three game demos as use cases. Hyouk Jang, Juheon Choi, Gunhee Kim |
VRST | 3 |
| 2019 | Video Question Answering with Spatio-Temporal Reasoning
Yunseok Jang 0001, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Gunhee Kim |
Int. J. Comput. Vis. | 6 |
| 2019 | Towards Personalized Image Captioning via Multimodal Memory NetworksabstractWe address personalized image captioning, which generates a descriptive sentence for a user's image, accounting for prior knowledge such as her active vocabularies or writing style in her previous documents. As applications of personalized image captioning, we solve two post automation tasks in social networks: hashtag prediction and post generation. The hashtag prediction predicts a list of hashtags for an image, while the post generation creates a natural post text consisting of normal words, emojis, and even hashtags. We propose a novel personalized captioning model named Context Sequence Memory Network (CSMN). Its unique updates over existing memory networks include (i) exploiting memory as a repository for multiple types of context information, (ii) appending previously generated words into memory to capture long-term information, and (iii) adopting CNN memory structure to jointly represent nearby ordered memory slots for better context understanding. For evaluation, we collect a new dataset InstaPIC-1.1M, comprising 1.1M Instagram posts from 6.3K users. We further use the benchmark YFCC100M dataset to validate the generality of our approach. With quantitative evaluation and user studies via Amazon Mechanical Turk, we show that the three novel features of the CSMN help enhance the performance of personalized image captioning over state-of-the-art captioning models. Cesc C. Park, Byeongchang Kim 0002, Gunhee Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | A Deep Ranking Model for Spatio-Temporal Highlight Detection From a 360◦ VideoabstractWe address the problem of highlight detection from a 360◦ video by summarizing it both spatially and temporally. Given a long 360◦ video, we spatially select pleasantly-looking normal field-of-view (NFOV) segments from unlimited field of views (FOV) of the 360◦ video, and temporally summarize it into a concise and informative highlight as a selected subset of subshots. We propose a novel deep ranking model named as Composition View Score (CVS) model, which produces a spherical score map of composition per video segment, and determines which view is suitable for highlight via a sliding window kernel at inference. To evaluate the proposed framework, we perform experiments on the Pano2Vid benchmark dataset (Su, Jayaraman, and Grauman 2016) and our newly collected 360◦ video highlight dataset from YouTube and Vimeo. Through evaluation using both quantitative summarization metrics and user studies via Amazon Mechanical Turk, we demonstrate that our approach outperforms several state-of-the-art highlight detection methods.We also show that our model is 16 times faster at inference than AutoCam (Su, Jayaraman, and Grauman 2016), which is one of the first summarization algorithms of 360◦ videos. Youngjae Yu, Sangho Lee 0008, Joonil Na, Jaeyun Kang, Gunhee Kim |
AAAI | 5 |
| 2018 | A Memory Network Approach for Story-Based Temporal Summarization of 360° VideosabstractWe address the problem of story-based temporal summarization of long 360° videos. We propose a novel memory network model named Past-Future Memory Network (PFMN), in which we first compute the scores of 81 normal field of view (NFOV) region proposals cropped from the input 360° video, and then recover a latent, collective summary using the network with two external memories that store the embeddings of previously selected subshots and future candidate subshots. Our major contributions are twofold. First, our work is the first to address story-based temporal summarization of 360° videos. Second, our model is the first attempt to leverage memory networks for video summarization tasks. For evaluation, we perform three sets of experiments. First, we investigate the view selection capability of our model on the Pano2Vid dataset [42]. Second, we evaluate the temporal summarization with a newly collected 360° video dataset. Finally, we experiment our model's performance in another domain, with image-based storytelling VIST dataset [22]. We verify that our model achieves state-of-the-art performance on all the tasks. Sangho Lee 0008, Jinyoung Sung, Youngjae Yu, Gunhee Kim |
CVPR | 4 |
| 2018 | Improving Occlusion and Hard Negative Handling for Single-Stage Pedestrian DetectorsabstractWe propose methods of addressing two critical issues of pedestrian detection: (i) occlusion of target objects as false negative failure, and (ii) confusion with hard negative examples like vertical structures as false positive failure. Our solutions to these two problems are general and flexible enough to be applicable to any single-stage detection models. We implement our methods into four state-of-the-art single-stage models, including SqueezeDet+ [22], YOLOv2 [17], SSD [12], and DSSD [8]. We empirically validate that our approach indeed improves the performance of those four models on Caltech pedestrian [4] and CityPersons dataset [25]. Moreover, in some heavy occlusion settings, our approach achieves the best reported performance. Specifically, our two solutions are as follows. For better occlusion handling, we update the output tensors of single-stage models so that they include the prediction of part confidence scores, from which we compute a final occlusion-aware detection score. For reducing confusion with hard negative examples, we introduce average grid classifiers as post-refinement classifiers, trainable in an end-to-end fashion with little memory and time overhead (e.g. increase of 1-5 MB in memory and 1-2 ms in inference time). Junhyug Noh, Soochan Lee, Gunhee Kim |
CVPR | 4 |
| 2018 | A Joint Sequence Fusion Model for Video Question Answering and Retrieval
Youngjae Yu, Jongseok Kim 0002, Gunhee Kim |
ECCV (7) | 3 |
| 2018 | Memorization Precedes Generation: Learning Unsupervised GANs with Memory Networks
Minjung Kim 0001, Gunhee Kim |
ICLR (Poster) | 3 |
| 2018 | Video Prediction with Appearance and Motion ConditionsabstractVideo prediction aims to generate realistic future frames by learning dynamic visual patterns. One fundamental challenge is to deal with future uncertainty: How should a model behave when there are multiple correct, equally probable future? We propose an Appearance-Motion Conditional GAN to address this challenge. We provide appearance and motion information as conditions that specify how the future may look like, reducing the level of uncertainty. Our model consists of a generator, two discriminators taking charge of appearance and motion pathways, and a perceptual ranking module that encourages videos of similar conditions to look similar. To train our model, we develop a novel conditioning scheme that consists of different combinations of appearance and motion conditions. We evaluate our model using facial expression and human action datasets and report favorable results compared to existing methods. Yunseok Jang 0001, Gunhee Kim, Yale Song |
ICML | 2 |
| 2018 | A Hierarchical Latent Structure for Variational Conversation ModelingabstractYookoon Park, Jaemin Cho, Gunhee Kim. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Yookoon Park, Jaemin Cho 0001, Gunhee Kim |
NAACL-HLT | 3 |
| 2018 | Retrieval of Sentence Sequences for an Image Stream via Coherence Recurrent Convolutional NetworksabstractWe propose an approach for retrieving a sequence of natural sentences for an image stream. Since general users often take a series of pictures on their experiences, much online visual information exists in the form of image streams, for which it would better take into consideration of the whole image stream to produce natural language descriptions. While almost all previous studies have dealt with the relation between a single image and a single natural sentence, our work extends both input and output dimension to a sequence of images and a sequence of sentences. For retrieving a coherent flow of multiple sentences for a photo stream, we propose a multimodal neural architecture called coherence recurrent convolutional network (CRCN), which consists of convolutional neural networks, bidirectional long short-term memory (LSTM) networks, and an entity-based local coherence model. Our approach directly learns from vast user-generated resource of blog posts as text-image parallel training data. We collect more than 22 K unique blog posts with 170 K associated images for the travel topics of NYC, Disneyland , Australia, and Hawaii. We demonstrate that our approach outperforms other state-of-the-art image captioning methods for text sequence generation, using both quantitative measures and user studies via Amazon Mechanical Turk. Cesc C. Park, Gunhee Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Detection and Recognition of Text Embedded in Online Images via Neural Context ModelsabstractWe address the problem of detecting and recognizing the text embedded in online images that are circulated over the Web. Our idea is to leverage context information for both text detection and recognition. For detection, we use local image context around the text region, based on that the text often sequentially appear in online images. For recognition, we exploit the metadata associated with the input online image, including tags, comments, and title, which are used as a topic prior for the word candidates in the image. To infuse such two sets of context information, we propose a contextual text spotting network (CTSN). We perform comparative evaluation with five state-of-the-art text spotting methods on newly collected Instagram and Flickr datasets. We show that our approach that benefits from context information is more successful for text spotting in online images. Chulmoo Kang, Gunhee Kim, Suk I. Yoo |
AAAI | 2 |
| 2017 | TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question AnsweringabstractVision and language understanding has emerged as a subject undergoing intense study in Artificial Intelligence. Among many tasks in this line of research, visual question answering (VQA) has been one of the most successful ones, where the goal is to learn a model that understands visual content at region-level details and finds their associations with pairs of questions and answers in the natural language form. Despite the rapid progress in the past few years, most existing work in VQA have focused primarily on images. In this paper, we focus on extending VQA to the video domain and contribute to the literature in three important ways. First, we propose three new tasks designed specifically for video VQA, which require spatio-temporal reasoning from videos to answer questions correctly. Next, we introduce a new large-scale dataset for video VQA named TGIF-QA that extends existing VQA work with our new tasks. Finally, we propose a dual-LSTM based approach with both spatial and temporal attention, and show its effectiveness over conventional VQA techniques through empirical evaluations. Yunseok Jang 0001, Yale Song, Youngjae Yu, Gunhee Kim |
CVPR | 5 |
| 2017 | Attend to You: Personalized Image Captioning with Context Sequence Memory NetworksabstractWe address personalization issues of image captioning, which have not been discussed yet in previous research. For a query image, we aim to generate a descriptive sentence, accounting for prior knowledge such as the users active vocabularies in previous documents. As applications of personalized image captioning, we tackle two post automation tasks: hashtag prediction and post generation, on our newly collected Instagram dataset, consisting of 1.1M posts from 6.3K users. We propose a novel captioning model named Context Sequence Memory Network (CSMN). Its unique updates over previous memory network models include (i) exploiting memory as a repository for multiple types of context information, (ii) appending previously generated words into memory to capture long-term information without suffering from the vanishing gradient problem, and (iii) adopting CNN memory structure to jointly represent nearby ordered memory slots for better context understanding. With quantitative evaluation and user studies via Amazon Mechanical Turk, we show the effectiveness of the three novel features of CSMN and its performance enhancement for personalized image captioning over state-of-the-art captioning models. Cesc C. Park, Byeongchang Kim 0002, Gunhee Kim |
CVPR | 3 |
| 2017 | Supervising Neural Attention Models for Video Captioning by Human Gaze DataabstractThe attention mechanisms in deep neural networks are inspired by humans attention that sequentially focuses on the most relevant parts of the information over time to generate prediction output. The attention parameters in those models are implicitly trained in an end-to-end manner, yet there have been few trials to explicitly incorporate human gaze tracking to supervise the attention models. In this paper, we investigate whether attention models can benefit from explicit human gaze labels, especially for the task of video captioning. We collect a new dataset called VAS, consisting of movie clips, and corresponding multiple descriptive sentences along with human gaze tracking data. We propose a video captioning model named Gaze Encoding Attention Network (GEAN) that can leverage gaze tracking information to provide the spatial and temporal attention for sentence generation. Through evaluation of language similarity metrics and human assessment via Amazon mechanical Turk, we demonstrate that spatial attentions guided by human gaze data indeed improve the performance of multiple captioning methods. Moreover, we show that the proposed approach achieves the state-of-the-art performance for both gaze prediction and video captioning not only in our VAS dataset but also in standard datasets (e.g. LSMDC [24] and Hollywood2 [18]). Youngjae Yu, Yeonhwa Kim, Kyung Yoo, Gunhee Kim |
CVPR | 6 |
| 2017 | End-to-End Concept Word Detection for Video Captioning, Retrieval, and Question AnsweringabstractWe propose a high-level concept word detector that can be integrated with any video-to-language models. It takes a video as input and generates a list of concept words as useful semantic priors for language generation models. The proposed word detector has two important properties. First, it does not require any external knowledge sources for training. Second, the proposed word detector is trainable in an end-to-end manner jointly with any video-to-language models. To effectively exploit the detected words, we also develop a semantic attention mechanism that selectively focuses on the detected concept words and fuse them with the word encoding and decoding in the language model. In order to demonstrate that the proposed approach indeed improves the performance of multiple video-to-language tasks, we participate in all the four tasks of LSMDC 2016 [18]. Our approach has won three of them, including fill-in-the-blank, multiple-choice test, and movie retrieval. Youngjae Yu, Hyungjin Ko, Gunhee Kim |
CVPR | 4 |
| 2017 | A Read-Write Memory Network for Movie Story UnderstandingabstractWe propose a novel memory network model named Read-Write Memory Network (RWMN) to perform question and answering tasks for large-scale, multimodal movie story understanding. The key focus of our RWMN model is to design the read network and the write network that consist of multiple convolutional layers, which enable memory read and write operations to have high capacity and flexibility. While existing memory-augmented network models treat each memory slot as an independent block, our use of multi-layered CNNs allows the model to read and write sequential memory cells as chunks, which is more reasonable to represent a sequential story because adjacent memory blocks often have strong correlations. For evaluation, we apply our model to all the six tasks of the MovieQA benchmark [24], and achieve the best accuracies on several tasks, especially on the visual QA task. Our model shows a potential to better understand not only the content in the story, but also more abstract information, such as relationships between characters and the reasons for their actions. Seil Na, Sangho Lee 0008, Jisung Kim, Gunhee Kim |
ICCV | 4 |
| 2017 | SplitNet: Learning to Semantically Split Deep Networks for Parameter Reduction and Model ParallelizationabstractWe propose a novel deep neural network that is both lightweight and effectively structured for model parallelization. Our network, which we name as SplitNet, automatically learns to split the network weights into either a set or a hierarchy of multiple groups that use disjoint sets of features, by learning both the class-to-group and feature-to-group assignment matrices along with the network weights. This produces a tree-structured network that involves no connection between branched subtrees of semantically disparate class groups. SplitNet thus greatly reduces the number of parameters and requires significantly less computations, and is also embarrassingly model parallelizable at test time, since the network evaluation for each subnetwork is completely independent except for the shared lower layer weights that can be duplicated over multiple processors. We validate our method with two deep network models (ResNet and AlexNet) on two different datasets (CIFAR-100 and ILSVRC 2012) for image classification, on which our method obtains networks with significantly reduced number of parameters while achieving comparable or superior classification accuracies over original full deep networks, and accelerated test speed with multiple GPUs. Juyong Kim 0002, Yookoon Park, Gunhee Kim, Sung Ju Hwang |
ICML | 3 |
| 2016 | Taxonomy-Regularized Semantic Deep Convolutional Neural Networks
Wonjoon Goo, Juyong Kim 0002, Gunhee Kim, Sung Ju Hwang |
ECCV (2) | 3 |
| 2016 | A calibration method for optical see-through head-mounted displays with a depth cameraabstractWe propose a fast and accurate calibration method for the optical see-through (OST) head-mounted displays (HMD), taking advantage of a low-cost time-of-flight depth-camera. Recently, affordable OST-HMDs and depth-cameras are widely appearing in the commercial market. In order to correctly reflect the user experience into the calibration process, our method demands a user wearing the HMD to repeatedly point at rendered virtual circles with their fingertips. From the repeated calibration data, we perform two stages of full calibration and simplified calibration, to compute key calibration parameters. The full calibration is required when the depth-camera is first installed to the HMD, and afterwards only the simplified calibration is performed whenever a user wears it again. Our experimental results show that the full and simplified calibration can be achieved with 10 and 5 user's repetitions (theoretically 3 and 2 at minimum), which are significantly less than about 20 of the stereo-SPAAM, one of the most popular existing calibration techniques. We also demonstrate that the 3D position errors of our calibration become much quickly smaller than those of the state-of-the-art method. Hanseul Jun, Gunhee Kim |
VR | 2 |
| 2015 | Ranking and retrieval of image sequences from multiple paragraph queriesabstractWe propose a method to rank and retrieve image sequences from a natural language text query, consisting of multiple sentences or paragraphs. One of the method's key applications is to visualize visitors' text-only reviews on TRIPADVISOR or YELP, by automatically retrieving the most illustrative image sequences. While most previous work has dealt with the relations between a natural language sentence and an image or a video, our work extends to the relations between paragraphs and image sequences. Our approach leverages the vast user-generated resource of blog posts and photo streams on the Web. We use blog posts as text-image parallel training data that co-locate informative text with representative images that are carefully selected by users. We exploit large-scale photo streams to augment the image samples for retrieval. We design a latent structural SVM framework to learn the semantic relevance relations between text and image sequences. We present both quantitative and qualitative results on the newly created DISNEYLAND dataset. Gunhee Kim, Seungwhan Moon, Leonid Sigal |
CVPR | 1 |
| 2015 | Joint photo stream and blog post summarization and explorationabstractWe propose an approach that utilizes large collections of photo streams and blog posts, two of the most prevalent sources of data on the Web, for joint story-based summarization and exploration. Blogs consist of sequences of images and associated text; they portray events and experiences with concise sentences and representative images. We leverage blogs to help achieve story-based semantic summarization of collections of photo streams. In the opposite direction, blog posts can be enhanced with sets of photo streams by showing interpolations between consecutive images in the blogs. We formulate the problem of joint alignment from blogs to photo streams and photo stream summarization in a unified latent ranking SVM framework. We alternate between solving the two coupled latent SVM problems, by first fixing the summarization and solving for the alignment from blog images to photo streams and vice versa. On a newly collected large-scale Disneyland dataset of 10K blogs (120K associated images) and 6K photo streams (540K images), we demonstrate that blog posts and photo streams are mutually beneficial for summarization, exploration, semantic knowledge transfer, and photo interpolation. Gunhee Kim, Seungwhan Moon, Leonid Sigal |
CVPR | 1 |
| 2015 | Storyline Representation of Egocentric Videos with an Applications to Story-Based SearchabstractEgocentric videos are a valuable source of information as a daily log of our lives. However, large fraction of egocentric video content is typically irrelevant and boring to re-watch. It is an agonizing task, for example, to manually search for the moment when your daughter first met Mickey Mouse from hours-long egocentric videos taken at Disneyland. Although many summarization methods have been successfully proposed to create concise representations of videos, in practice, the value of the subshots to users may change according to their immediate preference/mood, thus summaries with fixed criteria may not fully satisfy users' various search intents. To address this, we propose a storyline representation that expresses an egocentric video as a set of jointly inferred, through MRF inference, story elements comprising of actors, locations, supporting objects and events, depicted on a timeline. We construct such a storyline with very limited annotation data (a list of map locations and weak knowledge of what events may be possible at each location), by bootstrapping the process with data obtained through focused Web image and video searches. Our representation promotes story-based search with queries in the form of AND-OR graphs, which span any subset of story elements and their spatio-temporal composition. We show effectiveness of our approach on a set of unconstrained YouTube egocentric videos of visits to Disneyland. Gunhee Kim, Leonid Sigal |
ICCV | 2 |
| 2015 | Discovering Collective Narratives of Theme Parks from Large Collections of Visitors' Photo StreamsabstractWe present an approach for generating pictorial storylines from large collections of online photo streams shared by visitors to theme parks (e.g. Disneyland), along with publicly available information such as visitor's maps. The story graph visualizes various events and activities recurring across visitors' photo sets, in the form of hierarchically branching narrative structure associated with attractions and districts in theme parks. We first estimate story elements of each photo stream, including the detection of faces and supporting objects, and attraction-based localization. We then create spatio-temporal story graphs via an inference of sparse time-varying directed graphs. Through quantitative evaluation and crowdsourcing-based user studies via Amazon Mechanical Turk, we show that the story graphs serve as a more convenient mid-level data structure to perform photo-based recommendation tasks than other alternatives. We also present storybook-like demo examples regarding exploration, recommendation, and temporal analysis, which may be most beneficial uses of the story graphs to visitors. Gunhee Kim, Leonid Sigal |
KDD | 1 |
| 2015 | Dynamic Topic Modeling for Monitoring Market Competition from Online Text and Image DataabstractWe propose a dynamic topic model for monitoring temporal evolution of market competition by jointly leveraging tweets and their associated images. For a market of interest (e.g. luxury goods), we aim at automatically detecting the latent topics (e.g. bags, clothes, luxurious) that are competitively shared by multiple brands (e.g. Burberry, Prada, and Chanel), and tracking temporal evolution of the brands' stakes over the shared topics. One of key applications of our work is social media monitoring that can provide companies with temporal summaries of highly overlapped or discriminative topics with their major competitors. We design our model to correctly address three major challenges: multiview representation of text and images, modeling of competitiveness of multiple brands over shared topics, and tracking their temporal evolution. As far as we know, no previous model can satisfy all the three challenges. For evaluation, we analyze about 10 millions of tweets and 8 millions of associated images of the 23 brands in the two categories of luxury and beer. Through experiments, we show that the proposed approach is more successful than other candidate methods for the topic modeling of competition. We also quantitatively demonstrate the generalization power of the proposed method for three prediction tasks. Hao Zhang 0025, Gunhee Kim, Eric P. Xing |
KDD | 2 |
| 2015 | Expressing an Image Stream with a Sequence of Natural SentencesabstractWe propose an approach for generating a sequence of natural sentences for an image stream. Since general users usually take a series of pictures on their special moments, much online visual information exists in the form of image streams, for which it would better take into consideration of the whole set to generate natural language descriptions. While almost all previous studies have dealt with the relation between a single image and a single natural sentence, our work extends both input and output dimension to a sequence of images and a sequence of sentences. To this end, we design a novel architecture called coherent recurrent convolutional network (CRCN), which consists of convolutional networks, bidirectional recurrent networks, and entity-based local coherence model. Our approach directly learns from vast user-generated resource of blog posts as text-image parallel training data. We demonstrate that our approach outperforms other state-of-the-art candidate methods, using both quantitative measures (e.g. BLEU and top-K recall) and user studies via Amazon Mechanical Turk. Cesc C. Park, Gunhee Kim |
NIPS | 2 |
| 2014 | Joint Summarization of Large-Scale Collections of Web Images and Videos for Storyline ReconstructionabstractIn this paper, we address the problem of jointly summarizing large sets of Flickr images and YouTube videos. Starting from the intuition that the characteristics of the two media types are different yet complementary, we develop a fast and easily-parallelizable approach for creating not only high-quality video summaries but also novel structural summaries of online images as storyline graphs. The storyline graphs can illustrate various events or activities associated with the topic in a form of a branching network. The video summarization is achieved by diversity ranking on the similarity graphs between images and video frames. The reconstruction of storyline graphs is formulated as the inference of sparse time-varying directed graphs from a set of photo streams with assistance of videos. For evaluation, we collect the datasets of 20 outdoor activities, consisting of 2.7M Flickr images and 16K YouTube videos. Due to the large-scale nature of our problem, we evaluate our algorithm via crowdsourcing using Amazon Mechanical Turk. In our experiments, we demonstrate that the proposed joint summarization approach outperforms other baselines and our own methods using videos or images only. Gunhee Kim, Leonid Sigal, Eric P. Xing |
CVPR | 1 |
| 2014 | Reconstructing Storyline Graphs for Image Recommendation from Web Community PhotosabstractIn this paper, we investigate an approach for reconstructing storyline graphs from large-scale collections of Internet images, and optionally other side information such as friendship graphs. The storyline graphs can be an effective summary that visualizes various branching narrative structure of events or activities recurring across the input photo sets of a topic class. In order to explore further the usefulness of the storyline graphs, we leverage them to perform the image sequential prediction tasks, from which photo recommendation applications can benefit. We formulate the storyline reconstruction problem as an inference of sparse time-varying directed graphs, and develop an optimization algorithm that successfully addresses a number of key challenges of Web-scale problems, including global optimality, linear complexity, and easy parallelization. With experiments on more than 3.3 millions of images of 24 classes and user studies via Amazon Mechanical Turk, we show that the proposed algorithm improves other candidate methods for both storyline reconstruction and image prediction tasks. Gunhee Kim, Eric P. Xing |
CVPR | 1 |
| 2014 | Visualizing brand associations from web community photosabstractBrand Associations, one of central concepts in marketing, describe customers' top-of-mind attitudes or feelings toward a brand. Thus, this consumer-driven brand equity often attains the grounds for purchasing products or services of the brand. Traditionally, brand associations are measured by analyzing the text data from consumers' responses to the survey or their online conversation logs. In this paper, we propose to go beyond text data and leverage large-scale online photo collections contributed by the general public, which have not been explored so far. As a first technical step toward the study of photo-based brand associations, we aim to jointly achieve the following two visualization tasks in a mutually-rewarding way: (i) detecting and visualizing core visual concepts associated with brands, and (ii) localizing the regions of brand in the images. With experiments on about five millions of images of 48 brands crawled from five popular online photo sharing sites, we demonstrate that our approach can discover complementary views on the brand associations that are hardly mined from the text data. We also quantitatively show that our approach outperforms other candidate methods on the both visualization tasks. Gunhee Kim, Eric P. Xing |
WSDM | 1 |
| 2014 | QuMinS: Fast and scalable querying, mining and summarizing multi-modal databases
Robson L. F. Cordeiro, Fan Guo 0006, Donna S. Haverkamp, James H. Horne, Ellen K. Hughes, Gunhee Kim, Luciana A. S. Romani, Priscila P. Coltri, Tamires T. Souza, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
Inf. Sci. | 6 |
| 2013 | Jointly Aligning and Segmenting Multiple Web Photo Streams for the Inference of Collective Photo StorylinesabstractWith an explosion of popularity of online photo sharing, we can trivially collect a huge number of photo streams for any interesting topics such as scuba diving as an outdoor recreational activity class. Obviously, the retrieved photo streams are neither aligned nor calibrated since they are taken in different temporal, spatial, and personal perspectives. However, at the same time, they are likely to share common storylines that consist of sequences of events and activities frequently recurred within the topic. In this paper, as a first technical step to detect such collective storylines, we propose an approach to jointly aligning and segmenting uncalibrated multiple photo streams. The alignment task discovers the matched images between different photo streams, and the image segmentation task parses each image into multiple meaningful regions to facilitate the image understanding. We close a loop between the two tasks so that solving one task helps enhance the performance of the other in a mutually rewarding way. To this end, we design a scalable message-passing based optimization framework to jointly achieve both tasks for the whole input image set at once. With evaluation on the new Flickr dataset of 15 outdoor activities that consist of 1.5 millions of images of 13 thousands of photo streams, our empirical results show that the proposed algorithms are more successful than other candidate methods for both tasks. Gunhee Kim, Eric P. Xing |
CVPR | 1 |
| 2013 | Time-sensitive web image ranking and retrieval via dynamic multi-task regressionabstractIn this paper, we investigate a time-sensitive image retrieval problem, in which given a query keyword, a query time point, and optionally user information, we retrieve the most relevant and temporally suitable images from the database. Inspired by recently emerging interests on query dynamics in information retrieval research, our time-sensitive image retrieval algorithm can infer users' implicit search intent better and provide more engaging and diverse search results according to temporal trends of Web user photos. We model observed image streams as instances of multivariate point processes represented by several different descriptors, and develop a regularized multi-task regression framework that automatically selects and learns stochastic parametric models to solve the relations between image occurrence probabilities and various temporal factors that influence them. Using Flickr datasets of more than seven million images of 30 topics, our experimental results show that the proposed algorithm is more successful in time-sensitive image retrieval than other candidate methods, including ranking SVM, a PageRank-based image ranking, and a generative temporal topic model. Gunhee Kim, Eric P. Xing |
WSDM | 1 |
| 2012 | On multiple foreground cosegmentationabstractIn this paper, we address a challenging image segmentation problem called multiple foreground cosegmentation (MFC), which concerns a realistic scenario in general Webuser photo sets where a finite number of K foregrounds of interest repeatedly occur cross the entire photo set, but only an unknown subset of them is presented in each image. This contrasts the classical cosegmentation problem dealt with by most existing algorithms, which assume a much simpler but less realistic setting where the same set of foregrounds recurs in every image. We propose a novel optimization method for MFC, which makes no assumption on foreground configurations and does not suffer from the aforementioned limitation, while still leverages all the benefits of having co-occurring or (partially) recurring contents across images. Our method builds on an iterative scheme that alternates between a foreground modeling module and a region assignment module, both highly efficient and scalable. In particular, our approach is flexible enough to integrate any advanced region classifiers for foreground modeling, and our region assignment employs a combinatorial auction framework that enjoys several intuitively good properties such as optimality guarantee and linear complexity. We show the superior performance of our method in both segmentation quality and scalability in comparison with other state-of-the-art techniques on a newly introduced FlickrMFC dataset and the standard ImageNet dataset. Gunhee Kim, Eric P. Xing |
CVPR | 1 |
| 2012 | Web image prediction using multivariate point processesabstractIn this paper, we investigate a problem of predicting what images are likely to appear on the Web at a future time point, given a query word and a database of historical image streams that potentiates learning of uploading patterns of previous user images and associated metadata. We address such a Web image prediction problem at both a collective group level and an individual user level. We develop a predictive framework based on the multivariate point process, which employs a stochastic parametric model to solve the relations between image occurrence and the covariates that influence it, in a flexible, scalable, and globally optimal way. Using Flickr datasets of more than ten million images of 40 topics, our empirical results show that the proposed algorithm is more successful in predicting unseen Web images than other candidate methods, including forecasting on semantic meanings only, a PageRank-based image retrieval, and a generative author-time topic model. Gunhee Kim, Li Fei-Fei 0001, Eric P. Xing |
KDD | 1 |
| 2011 | Distributed cosegmentation via submodular optimization on anisotropic diffusionabstractThe saliency of regions or objects in an image can be significantly boosted if they recur in multiple images. Leveraging this idea, cosegmentation jointly segments common regions from multiple images. In this paper, we propose CoSand, a distributed cosegmentation approach for a highly variable large-scale image collection. The segmentation task is modeled by temperature maximization on anisotropic heat diffusion, of which the temperature maximization with finite K heat sources corresponds to a K-way segmentation that maximizes the segmentation confidence of every pixel in an image. We show that our method takes advantage of a strong theoretic property in that the temperature under linear anisotropic diffusion is a submodular function; therefore, a greedy algorithm guarantees at least a constant factor approximation to the optimal solution for temperature maximization. Our theoretic result is successfully applied to scalable cosegmentation as well as diversity ranking and single-image segmentation. We evaluate CoSand on MSRC and ImageNet datasets, and show its competence both in competitive performance over previous work, and in much superior scalability. Gunhee Kim, Eric P. Xing, Li Fei-Fei 0001, Takeo Kanade |
ICCV | 1 |
| 2010 | Modeling and Analysis of Dynamic Behaviors of Web Image Collections
Gunhee Kim, Eric P. Xing, Antonio Torralba 0001 |
ECCV (5) | 1 |
| 2010 | QMAS: Querying, Mining and Summarization of Multi-modal DatabasesabstractGiven a large collection of images, very few of which have labels, how can we guess the labels of the remaining majority, and how can we spot those images that need brand new labels, different from the existing ones? Current automatic labeling techniques usually scale super linearly with the data size, and/or they fail when only a tiny amount of labeled data is provided. In this paper, we propose QMAS (Querying, Mining And Summarization of Multi-modal Databases), a fast solution to the following problems: (i) low-labor labeling (L3) – given a collection of images, very few of which are labeled with keywords, find the most suitable labels for the remaining ones, and (ii) mining and attention routing – in the same setting, find clusters, the top-NO outlier images, and the top-NR representative images. We report experiments on real satellite images, two large sets (1.5GB and 2.25GB) of proprietary images and a smaller set (17MB) of public images. We show that QMAS scales linearly with the data size, being up to 40 times faster than top competitors (GCap), obtaining better or equal accuracy. In contrast to other methods, QMAS does low-labor labeling (L3), that is, it works even with tiny initial label sets. It also solves both presented problems and spots tiles that potentially require new labels. Robson L. F. Cordeiro, Fan Guo 0006, Donna S. Haverkamp, James H. Horne, Ellen K. Hughes, Gunhee Kim, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
ICDM | 6 |
| 2009 | Object Recognition with 3D ModelsabstractWe propose techniques for designing and training of pose-invariant object recognition systems using realistic 3d computer graphics models. We look at the relation between the size of the training set and the classification accuracy for a basic recognition task and provide a method for estimating the degree of difficulty of detecting an object. We show how to sample, align, and cluster images of objects on the view sphere. We address the problem of training on large, highly redundant data and propose a novel active learning method which generates compact training sets and compact classifiers. © 2009. The copyright of this document resides with its authors. Bernd Heisele, Gunhee Kim, Andrew Meyer |
BMVC | 2 |
| 2009 | Context-aware communication support system with pictographic cardsabstractWe present a context-aware pictographic display system that facilitates the search for communication cards that bear some relation to the location and goal of the user. The system consists of a server and a mobile device: the server searches for relevant cards on the basis of the context, and then prioritizes them in terms of their relatedness; the mobile device displays pictographic cards according to the priority. This system can help people who have speech-language difficulties by reducing the searching time and difficulties, when the user wants to find the pictographic cards. Gunhee Kim, Jukyung Park, Manchul Han, Se Hyung Park, Sungdo Ha |
Mobile HCI | 1 |
| 2009 | Unsupervised Detection of Regions of Interest Using Iterative Link AnalysisabstractThis paper proposes a fast and scalable alternating optimization technique to detect regions of interest (ROIs) in cluttered Web images without labels. The proposed approach discovers highly probable regions of object instances by iteratively repeating the following two functions: (1) choose the exemplar set (i.e. small number of high ranked reference ROIs) across the dataset and (2) refine the ROIs of each image with respect to the exemplar set. These two subproblems are formulated as ranking in two different similarity networks of ROI hypotheses by link analysis. The experiments with the PASCAL 06 dataset show that our unsupervised localization performance is better than one of state-of-the-art techniques and comparable to supervised methods. Also, we test the scalability of our approach with five objects in Flickr dataset consisting of more than 200,000 images. Gunhee Kim, Antonio Torralba 0001 |
NIPS | 1 |
| 2008 | Unsupervised modeling of object categories using link analysis techniquesabstractWe propose an approach for learning visual models of object categories in an unsupervised manner in which we first build a large-scale complex network which captures the interactions of all unit visual features across the entire training set and we infer information, such as which features are in which categories, directly from the graph by using link analysis techniques. The link analysis techniques are based on well-established graph mining techniques used in diverse applications such as WWW, bioinformatics, and social networks. The techniques operate directly on the patterns of connections between features in the graph rather than on statistical properties, e.g., from clustering in feature space. We argue that the resulting techniques are simpler, and we show that they perform similarly or better compared to state of the art techniques on common data sets. We also show results on more challenging data sets than those that have been used in prior work on unsupervised modeling. Gunhee Kim, Christos Faloutsos, Martial Hebert |
CVPR | 1 |
| 2008 | Segmentation of Salient Regions in Outdoor Scenes Using Imagery and 3-D DataabstractThis paper describes a segmentation method for extracting salient regions in outdoor scenes using both 3-D laser scans and imagery information. Our approach is a bottom- up attentive process without any high-level priors, models, or learning. As a mid-level vision task, it is not only robust against noise and outliers but it also provides valuable information for other high-level tasks in the form of optimal segments and their ranked saliency. In this paper, we propose a new saliency definition for 3-D point clouds and we incorporate it with saliency features from color information. Gunhee Kim, Daniel F. Huber, Martial Hebert |
WACV | 1 |
| 2007 | Navigation Behavior Selection Using Generalized Stochastic Petri Nets for a Service RobotabstractAppropriate design and control of behaviors of mobile robots are important for their successful autonomous navigation in a real dynamic environment. This paper proposes a formal selection framework of multiple navigation behaviors for a service robot. In the presented approach, modeling, analysis, and performance evaluation are carried out based on generalized stochastic Petri nets (GSPNs). By adopting a probabilistic approach, the proposed framework helps the robot to select the most desirable navigation behavior in run time according to environmental conditions. Moreover, after mission completion, the robot evaluates its prior navigation performance from accumulated data, and automatically uses the results to improve its future operations. Also, GSPNs have several advantages over direct use of other modeling formalisms such as finite state automata (FSA) or Markov processes (MPs). We conduct experiments on real guidance tasks with visitors by implementing the framework in the guide robotJinnyat the National Science Museum of Korea. The results show that the proposed strategy is useful for a robot's selection of an appropriate navigation behavior in a dynamic environment. Gunhee Kim, Woojin Chung |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2005 | A Selection Framework of Multiple Navigation Primitives Using Generalized Stochastic Petri NetsabstractThis paper proposes a selection framework of multiple navigation primitives for a service robot using Generalized Stochastic Petri Nets (GSPN’s). By adopting probabilistic approach, our framework helps the robot to select the most desirable navigation primitive in run time through the performance estimation according to environmental conditions. Moreover, after a mission, the robot evaluates prior navigation performance from accumulated data, and uses the results for the improvement of future operations. Modeling, analysis, and performance evaluation are conducted on firm mathematical foundation. Also, GSPN’s have several advantages over classic automata or direct use of Markov Process. We conducted simulations of the model derived from our experience of practical installations. The results showed that the framework is useful for primitive selection and performance analysis. Gunhee Kim, Woojin Chung |
ICRA | 1 |
| 2005 | Experimental research of navigation behavior selection using generalized stochastic Petri nets (GSPN) for a tour-guide robotabstractThis paper proposes a formal selection framework of multiple navigation behaviors for a service robot. In our approach, modeling, analysis, and performance evaluation are carried out based on the generalized stochastic Petri nets (GSPN). By adopting probabilistic approach, our framework helps the robot to select the most desirable navigation behavior in run time according to environmental conditions. Moreover, after a mission, the robot evaluates prior navigation performance from accumulated data, and uses the results for the improvement of future operations. Also, GSPN has several advantages over classic automata or direct use of Markov process. The basic ideas of the framework were introduced in our previous work (Gunhee Kim et al., 2005). Thus, this paper focuses on experimental verification by implementing the framework into the guide robot Jinny. We conduct the experiments about real guidance tasks with visitors in the National Science Museum of Korea. The results show that the proposed strategy is useful to select an appropriate navigation behavior in a dynamic space. Gunhee Kim, Woojin Chung, Sung-Kee Park |
IROS | 1 |
| 2004 | An Effective Adaptation of Encryption on MPEG-4 Video Streams for Digital Rights Management in an Ubiquitous Computing Environment
Gunhee Kim, Dongkyoo Shin, Dongil Shin |
EUC | 1 |
| 2004 | Design of a Middleware and HIML (Human Interaction Markup Language) for Context Aware Services in a Ubiquitous Computing Environment
Gunhee Kim, Dongkyoo Shin, Dongil Shin |
EUC | 1 |
| 2004 | Integrated Navigation System for Indoor Service Robots in Large-scale EnvironmentsabstractThis paper describes an integrated navigation strategy for the autonomous service robot in large-scale indoor environments. It includes architecture of navigation system, the development of crucial navigation algorithms like map, path planning, and localization, and planning scheme such as error/fault handling. Major advantages of proposed navigation are as follows: 1) A range sensor based generalized scheme of navigation without modification of the environment. 2) Intelligent navigation-related components. 3) Framework supporting the selection of multiple behaviors and error/fault handling schemes. A experimental result shows the feasibility of proposed navigation system. The result of this research has been successfully applied to our three service robots in a variety of task domains including a delivery, a patrol, a guide, and a floor cleaning task. Woojin Chung, Gunhee Kim, Chong-Won Lee |
ICRA | 2 |
| 2004 | Implementation of Multi-functional Service Robots using Tripodal Schematic Control ArchitectureabstractThis paper describes the implementation of multi-functional service robots using the Tripodal schematic control architecture. Our strategy has two major advantages. First, the proposed architecture supports Petri net based formal description of tasks and error/fault handling schemes. Second, it provides intuitive and straightforward guidelines system for integration issues. We show reusability and scalability of the proposed architecture by giving two examples of our experience. First, we explain how to add a newly developed cleaning function to our robot system. Second, we introduce the implementation process of a newly developed guide robot Jinny. Most of the modules developed for former robots are used directly in the Jinny system. Experimental results clearly showed that the developed strategy is efficient and easy-to-use. Gunhee Kim, Woojin Chung, Chong-Won Lee |
ICRA | 1 |
| 2004 | The autonomous tour-guide robot JinnyabstractThis paper explains a new tour-guide robot Jinny. The Jinny is developed by focusing on human robot interaction and autonomous navigation. In order to achieve reliable and safe navigation performance, an integrated navigation strategy is established based on the analysis of a robot's states and the decision making process of robot behaviors. According to the condition of environments, the robot can select its motion algorithm among four types of navigation strategy. Also, we emphasized the manageability of a robot's knowledge base for human friendly interactions. The robot's knowledge base can be extended or modified intuitively enough to be managed by non-experts. In order to show the feasibility and effectiveness of our system, we also present experimental results of the navigation system and some experiences on practical installations. Gunhee Kim, Woojin Chung, Kyung-Rock Kim, Sangmok Han, Richard H. Shinn |
IROS | 1 |
| 2003 | Tripodal schematic design of the control architecture for the service robot PSRabstractThis paper describes a control architecture design and a system integration strategy for the autonomous service robot PSR (Public Service Robot). The PSR is under development at the KIST (Korea Institute of Science and Technology) for service tasks in public spaces such as office buildings and hospitals. The proposed control architecture is designed by tripodal frameworks, which are layered functionality diagram, class diagram, and configuration diagram. The tripodal schematic design clearly points out the way of integrating various hardware and software components. The developed strategy is implemented on the PSR and successfully tested. Gunhee Kim, Woojin Chung, Chong-Won Lee |
ICRA | 1 |