Jaehoon Yoo

dblp:289/0340 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 59% Deep learning architectures and training · 19% Representation and self-supervised learning · 15%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
discrete diffusion model
0.912025
ReDi: Rectified Discrete Flow · NeurIPS 2025
Machine learning › Generative modeling › normalizing flow
discrete flow model
0.912025
ReDi: Rectified Discrete Flow · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
few-step generation
0.912025
ReDi: Rectified Discrete Flow · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › compositional representation
compositional representation learning
0.812024
Learning to Compose: Improving Object Centric Learning by Injecting Compositionality · ICLR 2024
Machine learning › Generative modeling
flow matching
0.812024
Simulation-Free Training of Neural ODEs on Paired Data · NeurIPS 2024
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations
0.812024
Simulation-Free Training of Neural ODEs on Paired Data · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.812024
Learning to Compose: Improving Object Centric Learning by Injecting Compositionality · ICLR 2024
Machine learning › Generative modeling › diffusion model › diffusion model training
simulation-free training
0.812024
Simulation-Free Training of Neural ODEs on Paired Data · NeurIPS 2024
Natural language and speech › Language models and text generation › language modeling › language model architecture
bidirectional transformer
0.712023
Towards End-to-End Generative Modeling of Long Videos with Memory-Efficient Bidirectional Transformers · CVPR 2023
Machine learning › Generative modeling › video generation
long video generation
0.712023
Towards End-to-End Generative Modeling of Long Videos with Memory-Efficient Bidirectional Transformers · CVPR 2023
Machine learning › Deep learning architectures and training
transformer
0.712023
Towards End-to-End Generative Modeling of Long Videos with Memory-Efficient Bidirectional Transformers · CVPR 2023
Machine learning › Generative modeling
video generation
0.712023
Towards End-to-End Generative Modeling of Long Videos with Memory-Efficient Bidirectional Transformers · CVPR 2023
Machine learning › Generative modeling
variational autoencoder
0.512021
SetVAE: Learning Hierarchical Composition for Generative Modeling of Set-Structured Data · CVPR 2021

Methods — techniques the papers use, named apart from their topics

iterative coupling rectification · 0.9conditional total correlation · 0.9slot attention · 0.8neural ODE · 0.8flow matching · 0.8memory-efficient bidirectional transformer · 0.7latent token projection · 0.7cross-attention · 0.7hierarchical latent variable model · 0.5attention · 0.5
YearPublicationVenuePosition
2025 ReDi: Rectified Discrete Flow
abstract
Discrete Flow-based Models (DFMs) are powerful generative models for high-quality discrete data but typically suffer from slow sampling speeds due to their reliance on iterative decoding processes. This reliance on a multi-step process originates from the factorization approximation of DFMs, which is necessary for handling high-dimensional data. In this paper, we analyze the factorization approximation error using Conditional Total Correlation (TC), and reveal its dependence on the coupling. To address the challenge of efficient few-step generation, we propose Rectified Discrete Flow (ReDi), a novel iterative method that reduces the underlying factorization error (measured as Conditional TC) by rectifying the coupling between source and target distributions. We theoretically prove that each ReDi step guarantees a monotonic decreasing Conditional TC, ensuring its convergence. Empirically, ReDi significantly reduces Conditional TC and enables few-step generation. Moreover, we demonstrate that the rectified couplings are well-suited for training efficient one-step models on image generation. ReDi offers a simple and theoretically grounded approach for tackling the few-step challenge, providing a new perspective on efficient discrete data synthesis. Code is available at https://github.com/Ugness/ReDi_discrete.
Jaehoon Yoo, Seunghoon Hong
NeurIPS1
2024 Learning to Compose: Improving Object Centric Learning by Injecting Compositionality
abstract
Learning compositional representation is a key aspect of object-centric learning as it enables flexible systematic generalization and supports complex visual reasoning. However, most of the existing approaches rely on auto-encoding objective, while the compositionality is implicitly imposed by the architectural or algorithmic bias in the encoder. This misalignment between auto-encoding objective and learning compositionality often results in failure of capturing meaningful object representations. In this study, we propose a novel objective that explicitly encourages compositionality of the representations. Built upon the existing object-centric learning framework (e.g., slot attention), our method incorporates additional constraints that an arbitrary mixture of object representations from two images should be valid by maximizing the likelihood of the composite data. We demonstrate that incorporating our objective to the existing framework consistently improves the objective-centric learning and enhances the robustness to the architectural choices.
Whie Jung, Jaehoon Yoo, Sungjin Ahn, Seunghoon Hong
ICLR2
2024 Simulation-Free Training of Neural ODEs on Paired Data
abstract
In this work, we investigate a method for simulation-free training of Neural Ordinary Differential Equations (NODEs) for learning deterministic mappings between paired data. Despite the analogy of NODEs as continuous-depth residual networks, their application in typical supervised learning tasks has not been popular, mainly due to the large number of function evaluations required by ODE solvers and numerical instability in gradient estimation. To alleviate this problem, we employ the flow matching framework for simulation-free training of NODEs, which directly regresses the parameterized dynamics function to a predefined target velocity field. Contrary to generative tasks, however, we show that applying flow matching directly between paired data can often lead to an ill-defined flow that breaks the coupling of the data pairs (e.g., due to crossing trajectories). We propose a simple extension that applies flow matching in the embedding space of data pairs, where the embeddings are learned jointly with the dynamic function to ensure the validity of the flow which is also easier to learn. We demonstrate the effectiveness of our method on both regression and classification tasks, where our method outperforms existing NODEs with a significantly lower number of function evaluations. The code is available at https://github.com/seminkim/simulation-free-node.
Semin Kim 0002, Jaehoon Yoo, Yeonwoo Cha, Saehoon Kim, Seunghoon Hong
NeurIPS2
2023 Towards End-to-End Generative Modeling of Long Videos with Memory-Efficient Bidirectional Transformers
abstract
Autoregressive transformers have shown remarkable success in video generation. However, the transformers are prohibited from directly learning the longterm dependency in videos due to the quadratic complexity of self-attention, and inherently suffering from slow inference time and error propagation due to the autoregressive process. In this paper, we propose Memory-efficient Bidirectional Transformer (MeBT) for end-to-end learning of longterm dependency in videos and fast inference. Based on recent advances in bidirectional transformers, our method learns to decode the entire spatio-temporal volume of a video in parallel from partially observed patches. The proposed transformer achieves a linear time complexity in both encoding and decoding, by projecting observable context tokens into a fixed number of latent tokens and conditioning them to decode the masked tokens through the cross-attention. Empowered by linear complexity and bidirectional modeling, our method demonstrates significant improvement over the autoregressive transformers for generating moderately long videos in both quality and speed. Videos and code are available at https://sites.google.com/view/mebt-cvpr2023.
Jaehoon Yoo, Semin Kim 0002, Doyup Lee, Chiheon Kim, Seunghoon Hong
CVPR1
2021 SetVAE: Learning Hierarchical Composition for Generative Modeling of Set-Structured Data
abstract
Generative modeling of set-structured data, such as point clouds, requires reasoning over local and global structures at various scales. However, adopting multi-scale frameworks for ordinary sequential data to a set-structured data is nontrivial as it should be invariant to the permutation of its elements. In this paper, we propose SetVAE, a hierarchical variational autoencoder for sets. Motivated by recent progress in set encoding, we build SetVAE upon attentive modules that first partition the set and project the partition back to the original cardinality. Exploiting this module, our hierarchical VAE learns latent variables at multiple scales, capturing coarse-to-fine dependency of the set elements while achieving permutation invariance. We evaluate our model on point cloud generation task and achieve competitive performance to the prior arts with substantially smaller model capacity. We qualitatively demonstrate that our model generalizes to unseen set sizes and learns interesting subset relations without supervision. Our implementation is available at https://github.com/jw9730/setvae.
Jaehoon Yoo, Juho Lee 0001, Seunghoon Hong
CVPR2