Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kyunghoon Bae

dblp:276/0021 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 28% Generative modeling · 18% Trustworthy machine learning · 12%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
pre-training
0.912025
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents · CVPR 2025
Computer vision › Vision and language
vision-language model
0.912025
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents · CVPR 2025
Machine learning › Transfer learning and domain adaptation
cross-task transfer
0.812024
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks · EMNLP 2024
Natural language and speech › Language models and text generation
in-context learning
0.812024
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents · NeurIPS 2024
Natural language and speech › Language models and text generation
instruction tuning
0.812024
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks · EMNLP 2024
Natural language and speech › Language models and text generation
LLM agents
0.812024
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents · NeurIPS 2024
Machine learning › Reinforcement learning › curriculum reinforcement learning
task selection
0.812024
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks · EMNLP 2024
Machine learning › Generative modeling
autoregressive model
0.612022
L-Verse: Bidirectional Generation Between Image and Text · CVPR 2022
Machine learning › Generative modeling › multimodal generation
bidirectional generation
0.612022
L-Verse: Bidirectional Generation Between Image and Text · CVPR 2022
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
convolutional neural network interpretation
0.512021
Explaining Convolutional Neural Networks through Attribution-Based Input Sampling and Block-Wise Feature Aggregation · AAAI 2021
Machine learning › Trustworthy machine learning
interpretability
0.512021
Explaining Convolutional Neural Networks through Attribution-Based Input Sampling and Block-Wise Feature Aggregation · AAAI 2021
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.212024
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks · EMNLP 2024
Machine learning › Generative modeling
variational autoencoder
0.212022
L-Verse: Bidirectional Generation Between Image and Text · CVPR 2022
Machine learning › Generative modeling › variational autoencoder
vector-quantized variational autoencoder
0.212022
L-Verse: Bidirectional Generation Between Image and Text · CVPR 2022
Visualization and visual analytics › explainable AI › explainable machine learning
explanation visualization
0.112021
Explaining Convolutional Neural Networks through Attribution-Based Input Sampling and Block-Wise Feature Aggregation · AAAI 2021

Methods — techniques the papers use, named apart from their topics

class activation mapping · 1.0block-wise feature aggregation · 1.0attribution-based input sampling · 1.0action identification · 0.9UI element detection · 0.9OCR · 0.9instruction tuning · 0.8in-context learning · 0.8vector quantized variational autoencoder · 0.6autoregressive transformer · 0.6
YearPublicationVenuePosition
2025 Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
abstract
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents from YouTube), a large-scale dataset of 313K annotated frames from 20K instructional videos capturing diverse real-world mobile OS navigation across multiple platforms. Models that include MONDAY in their pretraining phases demonstrate robust cross-platform generalization capabilities, consistently outperforming models trained on existing single OS datasets while achieving an average performance gain of 18.11%p on an unseen mobile OS platform. To enable continuous dataset expansion as mobile platforms evolve, we present an automated framework that leverages publicly available video content to create comprehensive task datasets without manual annotation. Our framework comprises robust OCR-based scene detection (95.04% F1-score), near-perfect UI element detection (99.87% hit ratio), and novel multi-step action identification to extract reliable action sequences across diverse interface configurations. We contribute both the MONDAY dataset and our automated collection framework to facilitate future research in mobile OS navigation, available at https://github.com/runamu/monday.
Yunseok Jang 0001, Yeda Song, Sungryull Sohn, Lajanugen Logeswaran, Tiange Luo, Dong-Ki Kim, Kyunghoon Bae, Honglak Lee
CVPR7
2024 Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks
abstract
Instruction tuning has been proven effective in enhancing zero-shot generalization across various tasks and in improving the performance of specific tasks.For task-specific improvements, strategically selecting and training on related tasks that provide meaningful supervision is crucial, as this approach enhances efficiency and prevents performance degradation from learning irrelevant tasks.In this light, we introduce a simple yet effective task selection method that leverages instruction information alone to identify relevant tasks, optimizing instruction tuning for specific tasks.Our method is significantly more efficient than traditional approaches, which require complex measurements of pairwise transferability between tasks or the creation of data samples for the target task.Additionally, by aligning the model with the unique instructional template style of the meta-dataset, we enhance its ability to granularly discern relevant tasks, leading to improved overall performance.Experimental results demonstrate that training on a small set of tasks, chosen solely based on the instructions, results in substantial improvements in performance on benchmarks such as P3, Big-Bench, NIV2, and Big-Bench Hard.Significantly, these improvements surpass those achieved by prior task selection methods, highlighting the superiority of our approach.1 * Equal contribution. 1 Code, model checkpoints, and data resources are available at https://github.com/CHLee0801/INSTA.
Janghoon Han, Seonghyeon Ye, Stanley Jungkyu Choi, Honglak Lee, Kyunghoon Bae
EMNLP6
2024 AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
abstract
Recent advances in large language models (LLMs) have empowered AI agents capable of performing various sequential decision-making tasks. However, effectively guiding LLMs to perform well in unfamiliar domains like web navigation, where they lack sufficient knowledge, has proven to be difficult with the demonstration-based in-context learning paradigm. In this paper, we introduce a novel framework, called AutoGuide, which addresses this limitation by automatically generating context-aware guidelines from offline experiences. Importantly, each context-aware guideline is expressed in concise natural language and follows a conditional structure, clearly describing the context where it is applicable. As a result, our guidelines facilitate the provision of relevant knowledge for the agent's current decision-making process, overcoming the limitations of the conventional demonstration-based learning paradigm. Our evaluation demonstrates that AutoGuide significantly outperforms competitive baselines in complex benchmark domains, including real-world web navigation.
Dong-Ki Kim, Jaekyeom Kim, Sungryull Sohn, Lajanugen Logeswaran, Kyunghoon Bae, Honglak Lee
NeurIPS6
2024 ReConPatch : Contrastive Patch Representation Learning for Industrial Anomaly Detection
abstract
Anomaly detection is crucial to the advanced identification of product defects such as incorrect parts, misaligned components, and damages in industrial manufacturing. Due to the rare observations and unknown types of defects, anomaly detection is considered to be challenging in machine learning. To overcome this difficulty, recent approaches utilize the common visual representations pre-trained from natural image datasets and distill the relevant features. However, existing approaches still have the discrepancy between the pre-trained feature and the target data, or require the input augmentation which should be carefully designed, particularly for the industrial dataset. In this paper, we introduce ReConPatch, which constructs discriminative features for anomaly detection by training a linear modulation of patch features extracted from the pre-trained model. ReConPatch employs contrastive representation learning to collect and distribute features in a way that produces a target-oriented and easily separable representation. To address the absence of labeled pairs for the contrastive learning, we utilize two similarity measures between data representations, pairwise and contextual similarities, as pseudo-labels. Our method achieves the state-of-the-art anomaly detection performance (99.72%) for the widely used and challenging MVTec AD dataset. Additionally, we achieved a state-of-the-art anomaly detection performance (95.8%) for the BTAD dataset.
Jeeho Hyun, Giyoung Jeon, Kyunghoon Bae, Byung Jun Kang
WACV5
2022 L-Verse: Bidirectional Generation Between Image and Text
abstract
Far beyond learning long-range interactions of natural language, transformers are becoming the de-facto standard for many vision tasks with their power and scalability. Especially with cross-modal tasks between image and text, vector quantized variational autoencoders (VQ-VAEs) are widely used to make a raw RGB image into a sequence of feature vectors. To better leverage the correlation between image and text, we propose L-Verse, a novel architecture consisting of feature-augmented variational autoencoder (AugVAE) and bidirectional auto-regressive transformer (BiART) for image-to-text and text-to-image generation. Our AugVAE shows the state-of-the-art reconstruction performance on ImageNetlK validation set, along with the robustness to unseen images in the wild. Unlike other models, BiART can distinguish between image (or text) as a conditional reference and a generation target. L-Verse can be directly used for image-to-text or text-to-image generation without any finetuning or extra object detection framework. In quantitative and qualitative experiments, L-Verse shows impressive results against previous methods in both image-to-text and text-to-image generation on MS-COCO Captions. We furthermore assess the scalability of L-Verse architecture on Conceptual Captions and present the initial result of bidirectional vision-language representation learning on general domain.
Gwangmo Song, Sihaeng Lee, Yewon Seo, Soonyoung Lee, Honglak Lee, Kyunghoon Bae
CVPR9
2021 Explaining Convolutional Neural Networks through Attribution-Based Input Sampling and Block-Wise Feature Aggregation
abstract
As an emerging field in Machine Learning, Explainable AI (XAI) has been offering remarkable performance in interpreting the decisions made by Convolutional Neural Networks (CNNs). To achieve visual explanations for CNNs, methods based on class activation mapping and randomized input sampling have gained great popularity. However, the attribution methods based on these techniques provide lower-resolution and blurry explanation maps that limit their explanation power. To circumvent this issue, visualization based on various layers is sought. In this work, we collect visualization maps from multiple layers of the model based on an attribution-based input sampling technique and aggregate them to reach a fine-grained and complete explanation. We also propose a layer selection strategy that applies to the whole family of CNN-based models, based on which our extraction framework is applied to visualize the last layers of each convolutional block of the model. Moreover, we perform an empirical analysis of the efficacy of derived lower-level information to enhance the represented attributions. Comprehensive experiments conducted on shallow and deep models trained on natural and industrial datasets, using both ground-truth and model-truth based evaluation metrics validate our proposed algorithm by meeting or outperforming the state-of-the-art methods in terms of explanation ability and visual quality, demonstrating that our method shows stability regardless of the size of objects or instances to be explained.
Sam Sattarzadeh, Mahesh Sudhakar, Anthony Lem, Shervin Mehryar, Konstantinos N. Plataniotis, Jongseong Jang, Yeonjeong Jeong, Kyunghoon Bae
AAAI10