EDBT 2026 Demo / reviewers in the wild / expert
Jaewoong Cho
dblp:184/3848
· DBLP profile ↗
19ranked-venue papers
2as first author
14since 2021 · last 2025
0009-0007-5864-9514ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure GamesabstractGUI agents powered by LLMs show promise in interacting with diverse digital environments.Among these, video games offer a valuable testbed due to their varied interfaces, with adventure games posing additional challenges through complex, narrative-driven interactions.Existing game benchmarks, however, lack diversity and rarely evaluate agents on completing entire storylines.To address this, we introduce FlashAdventure, a benchmark of 34 Flashbased adventure games designed to test full story arc completion and tackle the observationbehavior gap: the challenge of remembering and acting on earlier gameplay information.We also propose CUA-as-a-Judge, an automated gameplay evaluator, and COAST, an agentic framework leveraging long-term clue memory to better plan and solve sequential tasks.Experiments show current GUI agents struggle with full story arcs, while COAST improves milestone completion by bridging the observationbehavior gap.Nonetheless, a marked discrepancy between humans and best-performing agents warrants continued research efforts to narrow this divide. * Equal contribution. †Work done during an internship at KRAFTON. Flash-Based Adventure GamesInput GUI Agent (Operator) Gameplay Jaewoo Ahn, Junseo Kim, Heeseung Yun, Jaehyeon Son, Dongmin Park, Jaewoong Cho, Gunhee Kim |
EMNLP | 6 |
| 2025 | DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific FactorsabstractLarge-scale latent diffusion models (LDMs) excel in content generation across various modalities, but their reliance on phonemes and durations in text-to-speech (TTS) limits scalability and access from other fields. While recent studies show potential in removing these domain-specific factors, performance remains suboptimal. In this work, we introduce DiTTo-TTS, a Diffusion Transformer (DiT)-based TTS model, to investigate whether LDM-based TTS can achieve state-of-the-art performance without domain-specific factors. Through rigorous analysis and empirical exploration, we find that (1) DiT with minimal modifications outperforms U-Net, (2) variable-length modeling with a speech length predictor significantly improves results over fixed-length approaches, and (3) conditions like semantic alignment in speech latent representations are key to further enhancement. By scaling our training data to 82K hours and the model size to 790M parameters, we achieve superior or comparable zero-shot performance to state-of-the-art TTS models in naturalness, intelligibility, and speaker similarity, all without relying on domain-specific factors. Speech samples are available at https://ditto-tts.github.io. Keon Lee, Dong Won Kim, Jaehyeon Kim, Seungjun Chung, Jaewoong Cho |
ICLR | 5 |
| 2025 | Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM GuidanceabstractState-of-the-art text-to-image (T2I) diffusion models often struggle to generate rare compositions of concepts, e.g., objects with unusual attributes. In this paper, we show that the compositional generation power of diffusion models on such rare concepts can be significantly enhanced by the Large Language Model (LLM) guidance. We start with empirical and theoretical analysis, demonstrating that exposing frequent concepts relevant to the target rare concepts during the diffusion sampling process yields more accurate concept composition. Based on this, we propose a training-free approach, R2F, that plans and executes the overall rare-to-frequent concept guidance throughout the diffusion inference by leveraging the abundant semantic knowledge in LLMs. Our framework is flexible across any pre-trained diffusion models and LLMs, and can be seamlessly integrated with the region-guided diffusion approaches. Extensive experiments on three datasets, including our newly proposed benchmark, RareBench, containing various prompts with rare compositions of concepts, R2F significantly surpasses existing models including SD3.0 and FLUX by up to 28.1%p in T2I alignment. Code is available at https://github.com/krafton-ai/Rare-to-Frequent. Dongmin Park, Sebin Kim, Taehong Moon, Kangwook Lee 0001, Jaewoong Cho |
ICLR | 6 |
| 2025 | Efficient Generative Modeling with Residual Vector Quantization-Based TokensabstractWe introduce ResGen, an efficient Residual Vector Quantization (RVQ)-based generative model for high-fidelity generation with fast sampling. RVQ improves data fidelity by increasing the number of quantization steps, referred to as depth, but deeper quantization typically increases inference steps in generative models. To address this, ResGen directly predicts the vector embedding of collective tokens rather than individual ones, ensuring that inference steps remain independent of RVQ depth. Additionally, we formulate token masking and multi-token prediction within a probabilistic framework using discrete diffusion and variational inference. We validate the efficacy and generalizability of the proposed method on two challenging tasks across different modalities: conditional image generation on ImageNet 256$\times$256 and zero-shot text-to-speech synthesis. Experimental results demonstrate that ResGen outperforms autoregressive counterparts in both tasks, delivering superior performance without compromising sampling speed. Furthermore, as we scale the depth of RVQ, our generative models exhibit enhanced generation fidelity or faster sampling speeds compared to similarly sized baseline models. Jaehyeon Kim, Taehong Moon, Keon Lee, Jaewoong Cho |
ICML | 4 |
| 2025 | Lexico: Extreme KV Cache Compression via Sparse Coding over Universal DictionariesabstractWe introduce Lexico, a novel KV cache compression method that leverages sparse coding with a universal dictionary. Our key finding is that key-value cache in modern LLMs can be accurately approximated using sparse linear combination from a small, input-agnostic dictionary of ~4k atoms, enabling efficient compression across different input prompts, tasks and models. Using orthogonal matching pursuit for sparse approximation, Lexico achieves flexible compression ratios through direct sparsity control. On GSM8K, across multiple model families (Mistral, Llama 3, Qwen2.5), Lexico maintains 90-95% of the original performance while using only 15-25% of the full KV-cache memory, outperforming both quantization and token eviction methods. Notably, Lexico remains effective in low memory regimes where 2-bit quantization fails, achieving up to 1.7x better compression on LongBench and GSM8K while maintaining high accuracy. Junhyuck Kim, Jaewoong Cho, Dimitris S. Papailiopoulos |
ICML | 3 |
| 2025 | Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and EvaluationabstractJaechang Kim, Jinmin Goh, Inseok Hwang, Jaewoong Cho, Jungseul Ok. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jaechang Kim 0001, Jinmin Goh, Inseok Hwang 0001, Jaewoong Cho, Jungseul Ok |
NAACL (Long Papers) | 4 |
| 2025 | Distilling LLM Agent into Small Models with Retrieval and Code ToolsabstractLarge language models (LLMs) excel at complex reasoning tasks but remain computationally expensive, limiting their practical deployment.
To address this, recent works have focused on distilling reasoning capabilities into smaller language models (sLMs) using chain-of-thought (CoT) traces from teacher LLMs.
However, this approach struggles in scenarios requiring rare factual knowledge or precise computation, where sLMs often hallucinate due to limited capability.
In this work, we propose Agent Distillation, a framework for transferring not only reasoning capability but full task-solving behavior from LLM-based agents into sLMs with retrieval and code tools.
We improve agent distillation along two complementary axes: (1) we introduce a prompting method called first-thought prefix to enhance the quality of teacher-generated trajectories;
and (2) we propose a self-consistent action generation for improving test-time robustness of small agents.
We evaluate our method on eight reasoning tasks across factual and mathematical domains, covering both in-domain and out-of-domain generalization.
Our results show that sLMs as small as 0.5B, 1.5B, 3B parameters can achieve performance competitive with next-tier larger 1.5B, 3B, 7B models fine-tuned using CoT distillation, demonstrating the potential of agent distillation for building practical, tool-using small agents. Minki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho, Sung Ju Hwang |
NeurIPS | 4 |
| 2025 | Delving into Large Language Models for Effective Time-Series Anomaly DetectionabstractRecent efforts to apply Large Language Models (LLMs) to time-series anomaly detection (TSAD) have yielded limited success, often performing worse than even simple methods. While prior work has focused solely on downstream performance evaluation, the fundamental question—why do LLMs struggle with TSAD?—has remained largely unexplored. In this paper, we present an in-depth analysis that identifies two core challenges in understanding complex temporal dynamics and accurately localizing anomalies. To address these challenges, we propose a simple yet effective method that combines statistical decomposition with index-aware prompting. Our method outperforms 21 existing prompting strategies on the AnomLLM benchmark, achieving up to a 66.6\% improvement in F1 score. We further compare LLMs with 16 non-LLM baselines on the TSB-AD benchmark, highlighting scenarios where LLMs offer unique advantages via contextual reasoning. Our findings provide empirical insights into how and when LLMs can be effective for TSAD. The code is publicly available at: https://github.com/junwoopark92/LLM-TSAD Junwoo Park, Kyudan Jung, Dohyun Lee 0001, Hyuck Lee, Daehoon Gwak, Chaehun Park, Jaegul Choo, Jaewoong Cho |
NeurIPS | 8 |
| 2024 | CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-SpeechabstractWith the emergence of neural audio codecs, which encode multiple streams of discrete tokens from audio, large language models have recently gained attention as a promising approach for zero-shot Text-to-Speech (TTS) synthesis. Despite the ongoing rush towards scaling paradigms, audio tokenization ironically amplifies the scalability challenge, stemming from its long sequence length and the complexity of modelling the multiple sequences. To mitigate these issues, we present CLaM-TTS that employs a probabilistic residual vector quantization to (1) achieve superior compression in the token length, and (2) allow a language model to generate multiple tokens at once, thereby eliminating the need for cascaded modeling to handle the number of token streams. Our experimental results demonstrate that CLaM-TTS is better than or comparable to state-of-the-art neural codec-based TTS models regarding naturalness, intelligibility, speaker similarity, and inference speed. In addition, we examine the impact of the pretraining extent of the language models and their text tokenization strategies on performances. Jaehyeon Kim, Keon Lee, Seungjun Chung, Jaewoong Cho |
ICLR | 4 |
| 2024 | Image Clustering Conditioned on Text CriteriaabstractClassical clustering methods do not provide users with direct control of the clustering results, and the clustering results may not be consistent with the relevant criterion that a user has in mind. In this work, we present a new methodology for performing image clustering based on user-specified criteria in the form of text by leveraging modern Vision-Language Models and Large Language Models. We call our method Image Clustering Conditioned on Text Criteria (IC$|$TC), and it represents a different paradigm of image clustering. IC$|$TC requires a minimal and practical degree of human intervention and grants the user significant control over the clustering results in return. Our experiments show that IC$|$TC can effectively cluster images with various criteria, such as human action, physical location, or the person's mood, significantly outperforming baselines. Sehyun Kwon, Jaeseung Park, Jaewoong Cho, Ernest K. Ryu, Kangwook Lee 0001 |
ICLR | 4 |
| 2024 | A Simple Early Exiting Framework for Accelerated Sampling in Diffusion ModelsabstractDiffusion models have shown remarkable performance in generation problems over various domains including images, videos, text, and audio. A practical bottleneck of diffusion models is their sampling speed, due to the repeated evaluation of score estimation networks during the inference. In this work, we propose a novel framework capable of adaptively allocating compute required for the score estimation, thereby reducing the overall sampling time of diffusion models. We observe that the amount of computation required for the score estimation may vary along the time step for which the score is estimated. Based on this observation, we propose an early-exiting scheme, where we skip the subset of parameters in the score estimation network during the inference, based on a time-dependent exit schedule. Using the diffusion models for image synthesis, we show that our method could significantly improve the sampling throughput of the diffusion models without compromising image quality. Furthermore, we also demonstrate that our method seamlessly integrates with various types of solvers for faster sampling, capitalizing on their compatibility to enhance overall efficiency. Tae Hong Moon, Moonseok Choi, EungGu Yun 0001, Jongmin Yoon, Gayoung Lee, Jaewoong Cho, Juho Lee 0001 |
ICML | 6 |
| 2024 | Can Mamba Learn How To Learn? A Comparative Study on In-Context Learning TasksabstractState-space models (SSMs), such as Mamba (Gu & Dao, 2023), have been proposed as alternatives to Transformer networks in language modeling, incorporating gating, convolutions, and input-dependent token selection to mitigate the quadratic cost of multi-head attention. Although SSMs exhibit competitive performance, their in-context learning (ICL) capabilities, a remarkable emergent property of modern language models that enables task execution without parameter optimization, remain less explored compared to Transformers. In this study, we evaluate the ICL performance of SSMs, focusing on Mamba, against Transformer models across various tasks. Our results show that SSMs perform comparably to Transformers in standard regression ICL tasks, while outperforming them in tasks like sparse parity learning. However, SSMs fall short in tasks involving non-standard retrieval functionality. To address these limitations, we introduce a hybrid model, MambaFormer, that combines Mamba with attention blocks, surpassing individual models in tasks where they struggle independently. Our findings suggest that hybrid architectures offer promising avenues for enhancing ICL in language models. Jaeseung Park, Zheyang Xiong, Jaewoong Cho, Samet Oymak, Kangwook Lee 0001, Dimitris S. Papailiopoulos |
ICML | 5 |
| 2024 | Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language ModelsabstractAs Large Language Models (LLMs) are increasingly deployed in specialized domains with continuously evolving knowledge, the need for timely and precise knowledge injection has become essential. Fine-tuning with paraphrased data is a common approach to enhance knowledge injection, yet it faces two significant challenges: high computational costs due to repetitive external model usage and limited sample diversity.
To this end, we introduce LaPael, a latent-level paraphrasing method that applies input-dependent noise to early LLM layers.
This approach enables diverse and semantically consistent augmentations directly within the model. Furthermore, it eliminates the recurring costs of paraphrase generation for each knowledge update.
Our extensive experiments on question-answering benchmarks demonstrate that LaPael improves knowledge injection over standard fine-tuning and existing noise-based approaches.
Additionally, combining LaPael with data-level paraphrasing further enhances performance. Minki Kang, Sung Ju Hwang, Gibbeum Lee, Jaewoong Cho |
NeurIPS | 4 |
| 2023 | Censored Sampling of Diffusion Models Using 3 Minutes of Human FeedbackabstractDiffusion models have recently shown remarkable success in high-quality image generation. Sometimes, however, a pre-trained diffusion model exhibits partial misalignment in the sense that the model can generate good images, but it sometimes outputs undesirable images. If so, we simply need to prevent the generation of the bad images, and we call this task censoring. In this work, we present censored generation with a pre-trained diffusion model using a reward model trained on minimal human feedback. We show that censoring can be accomplished with extreme human feedback efficiency and that labels generated with a mere few minutes of human feedback are sufficient. Taeho Yoon, Kibeom Myoung, Keon Lee, Jaewoong Cho, Albert No, Ernest K. Ryu |
NeurIPS | 4 |
| 2020 | A Fair Classifier Using Mutual InformationabstractAs machine learning becomes prevalent in our daily lives involving a widening array of applications such as medicine, finance, job hiring and criminal justice, one morally & legally motivated need for machine learning algorithms is to ensure fairness for disadvantageous against advantageous groups. Fairness in machine learning aims at guaranteeing the irrelevancy of a prediction output to sensitive attributes like race, sex and religion. To this end, we take an information- theoretic approach using mutual information (MI) which can fully capture such independence. Inspired by the fact that MI between prediction and the sensitive attribute being zero is the "sufficient and necessary condition" for independence, we develop an MI-based algorithm that well trades off prediction accuracy for fairness performance often quantified as Disparate Impact (DI) or Equalized Odds (EO). Our experiments both on synthetic and benchmark real datasets demonstrate that our algorithm outperforms prior fair classifiers in tradeoff performance both w.r.t. DI and EO. Jaewoong Cho, Gyeongjo Hwang, Changho Suh |
ISIT | 1 |
| 2020 | A Fair Classifier Using Kernel Density EstimationabstractAs machine learning becomes prevalent in a widening array of sensitive applications such as job hiring and criminal justice, one critical aspect that machine learning classifiers should respect is to ensure fairness: guaranteeing the irrelevancy of a prediction output to sensitive attributes such as gender and race. In this work, we develop a kernel density estimation trick to quantify fairness measures that capture the degree of the irrelevancy. A key feature of our approach is that quantified fairness measures can be expressed as differentiable functions w.r.t. classifier model parameters. This then allows us to enjoy prominent gradient descent to readily solve an interested optimization problem that fully respects fairness constraints. We focus on a binary classification setting and two well-known definitions of group fairness: Demographic Parity (DP) and Equalized Odds (EO). Our experiments both on synthetic and benchmark real datasets demonstrate that our algorithm outperforms prior fair classifiers in accuracy-fairness tradeoff performance both w.r.t. DP and EO. Jaewoong Cho, Gyeongjo Hwang, Changho Suh |
NeurIPS | 1 |
| 2018 | Two-Way Interference Channel Capacity: How to Have the Cake and Eat It Too
Changho Suh, Jaewoong Cho, David Tse |
IEEE Trans. Inf. Theory | 2 |
| 2017 | Two-way interference channel capacity: How to have the cake and eat it tooabstractTwo-way communication is prevalent and its fundamental limits are first studied in the point-to-point setting by Shannon. One natural extension is a two-way interference channel (IC) with four independent messages: two associated with each direction of communication. In this paper, we explore a deterministic two-way IC, which captures the key properties of the wireless Gaussian channel. Our main contribution lies in the complete capacity region characterization of the two-way IC (with respect to the forward and backward sum-rate pair) via a new achievable scheme and a new converse. One surprising consequence of this result is that not only we can get an interaction gain over the one-way non-feedback capacities, we can sometimes get all the way to perfect feedback capacities in both directions simultaneously. In addition, our novel outer bound characterizes channel regimes in which interaction has no bearing on capacity. Changho Suh, Jaewoong Cho, David Tse |
ISIT | 2 |
| 2016 | To feedback or not to feedbackabstractWe explore two-way interference channels (ICs) where there are forward and backward ICs with four independent messages: two associated with the forward IC and the other two with respect to the backward IC. For a linear deterministic model of this channel, we develop inner and outer bounds on the capacity region. As a consequence, we demonstrate that interaction across forward and backward channels enables a more beneficial use of the channels, thereby yielding strict capacity improvements over non-interactive independent transmission. Moreover, our novel outer bound establishes the characterization of channel regimes in which interaction has no bearing on sum capacity. Changho Suh, David Tse, Jaewoong Cho |
ISIT | 3 |