EDBT 2026 Demo / reviewers in the wild / expert
Daeil Kim
dblp:158/6763
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OASIS: Object-guided Attention for Text-conditional Diffusion Synthesis of Human Interaction SequencesabstractAnalyzing and synthesizing human-object interaction is crucial for advancing intelligent systems that engage with the physical environment. However, simultaneous tracking of human and object data presents inherent challenges, resulting in limitations in dataset scale, diversity, and annotation quality within this domain, thereby hindering the generalization ability of trained models. This study introduces OASIS, a novel framework that extends pretrained text-conditional human motion diffusion models to address the complex task of fullbody 3D hand-object interaction generation. Specifically, we freeze the parameters of the pretrained motion diffusion model, while incorporating additional object-guided attention layers, which we train to adapt the human motion latents to match the input object motion sequence and the text. Our method can be understood as a ControlNet [38] for interaction. Through extensive experimentation, we demonstrate the effectiveness and robustness of our framework in generating realistic handobject interactions from textual descriptions. Our method surpasses the state-of-the-art performance in FID and accuracy interaction fidelity metrics compared to the prior best method IMoS [10], with improvements of 0.08 in FID and $2 \%$ in accuracy for body motion synthesis, and 0.15 in FID and $10 \%$ in accuracy for hand motion synthesis. Chih-Chun Yang 0005, Tianhui Cai, Zoltán Ádám Milacski, Aayush Prakash, Shingo Takagi 0001, Daeil Kim, Fernando De la Torre |
FG | 6 |
| 2025 | Multi-View Image Diffusion via Coordinate Noise and Fourier AttentionabstractRecently, text-to-image generation with diffusion models has made significant advancements in both higher fidelity and generalization capabilities compared to previous baselines. However, generating holistic multi-view consistent images from prompts still remains an important and challenging task. To address this challenge, we propose a diffusion process that attends to time-dependent spatial frequencies of features with a novel attention mechanism as well as novel noise initialization technique and cross-attention loss. This Fourier-based attention block focuses on features from non-overlapping regions of the generated scene in order to better align the global appearance. Our noise initialization technique incorporates shared noise and low spatial frequency information derived from pixel coordinates and depth maps to induce noise correlations across views. The cross-attention loss further aligns features sharing the same prompt across the scene. Our technique improves SOTA on several quantitative metrics with qualitatively better results when compared to other state-of-the-art approaches for multi-view consistency. Justin Theiss, Norman Müller, Daeil Kim, Aayush Prakash |
WACV | 3 |
| 2024 | Generalizable Human Gaussians for Sparse View Synthesis
Youngjoong Kwon, Baole Fang, Yixing Lu, Haoye Dong, Cheng Zhang 0014, Francisco Vicente 0001, Albert Mosella-Montoro, Jianjin Xu, Shingo Takagi 0001, Daeil Kim, Aayush Prakash, Fernando De la Torre |
ECCV (78) | 10 |
| 2024 | Target-Aware Language Modeling via Granular Data SamplingabstractErnie Chang, Pin-Jie Lin, Yang Li, Changsheng Zhao, Daeil Kim, Rastislav Rabatin, Zechun Liu, Yangyang Shi, Vikas Chandra. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ernie Chang, Pin-Jie Lin, Yang Li 0183, Changsheng Zhao 0002, Daeil Kim, Rastislav Rabatin, Zechun Liu, Yangyang Shi, Vikas Chandra |
EMNLP | 5 |
| 2024 | Towards Realistic Generative 3D Face ModelsabstractIn recent years, there has been significant progress in 2D generative face models fueled by applications such as animation, synthetic data generation, and digital avatars. However, due to the absence of 3D information, these 2D models often struggle to accurately disentangle facial attributes like pose, expression, and illumination, limiting their editing capabilities. To address this limitation, this paper proposes a 3D controllable generative face model to produce high-quality albedo and precise 3D shapes by leveraging existing 2D generative models. By combining 2D face generative models with semantic face manipulation, this method enables editing of detailed 3D rendered faces. The proposed framework utilizes an alternating descent optimization approach over shape and albedo. Differentiable rendering is used to train high-quality shapes and albedo without 3D supervision. Moreover, this approach outperforms most state-of-the-art (SOTA) methods in the well-known NoW and REALY benchmarks for 3D face re construction. It also outperforms the SOTA reconstruction models in recovering rendered faces’ identities across novel poses. Additionally, the paper demonstrates direct control of expressions in 3D faces by exploiting latent space leading to text-based editing of 3D faces. Aashish Rai, Hiresh Gupta, Francisco Vicente 0001, Shingo Takagi 0001, Amaury Aubel, Daeil Kim, Aayush Prakash, Fernando De la Torre |
WACV | 7 |
| 2024 | MotionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 PromptingabstractThere are numerous applications for human motion synthesis, including animation, gaming, robotics, or sports science. In recent years, human motion generation from natural language has emerged as a promising alternative to costly and labor-intensive data collection methods relying on motion capture or wearable sensors (e.g., suits). Despite this, generating human motion from textual descriptions remains a challenging and intricate task, primarily due to the scarcity of large-scale supervised datasets capable of capturing the full diversity of human activity.This study proposes a new approach, called MotionGPT, to address the limitations of previous text-based human motion generation methods by utilizing the extensive semantic information available in large language models (LLMs). We first pretrain a doubly text-conditional motion diffusion model on both coarse ("high-level") and detailed ("low-level") ground truth text data. Then during inference, we improve motion diversity and alignment with the training set, by zero-shot prompting GPT-3 for additional "low-level" details. Our method achieves new state-of-the-art quantitative results in terms of Fréchet Inception Distance (FID) and motion diversity metrics, and improves all considered metrics. Furthermore, it has strong qualitative performance, producing natural results. Code is available at https://github.com/humansensinglab/MotionGPT José Ribeiro-Gomes, Tianhui Cai, Zoltán Ádám Milacski, Aayush Prakash, Shingo Takagi 0001, Amaury Aubel, Daeil Kim, Alexandre Bernardino, Fernando De la Torre |
WACV | 8 |
| 2023 | Controllable 3D Generative Adversarial Face Model via Disentangling Shape and Appearanceabstract3D face modeling has been an active area of research in computer vision and computer graphics, fueling applications ranging from facial expression transfer in virtual avatars to synthetic data generation. Existing 3D deep learning generative models (e.g., VAE, GANs) allow generating compact face representations (both shape and texture) that can model non-linearities in the shape and appearance space (e.g., scatter effects, specularities,..). However, they lack the capability to control the generation of subtle expressions. This paper proposes a new 3D face generative model that can decouple identity and expression and provides granular control over expressions. In particular, we propose using a pair of supervised auto-encoder and generative adversarial networks to produce high-quality 3D faces, both in terms of appearance and shape. Experimental results in the generation of 3D faces learned with holistic expression labels, or Action Unit (AU) labels, show how we can decouple identity and expression; gaining fine-control over expressions while preserving identity.1 Fariborz Taherkhani, Aashish Rai, Quankai Gao, Shaunak Srivastava, Xuanbai Chen, Fernando De la Torre, Steven Song, Aayush Prakash, Daeil Kim |
WACV | 9 |
| 2022 | Unpaired Image Translation via Vector Symbolic Architectures
Justin Theiss, Jay Leverett, Daeil Kim, Aayush Prakash |
ECCV (21) | 3 |
| 2021 | RarePlanes: Synthetic Data Takes FlightabstractRarePlanes is a unique open-source machine learning dataset that incorporates both real and synthetically generated satellite imagery. The RarePlanes dataset specifically focuses on the value of synthetic data to aid computer vision algorithms in their ability to automatically detect aircraft and their attributes in satellite imagery. Although other synthetic/real combination datasets exist, RarePlanes is the largest openly-available very-high resolution dataset built to test the value of synthetic data from an overhead perspective. Previous research has shown that synthetic data can reduce the amount of real training data needed and potentially improve performance for many tasks in the computer vision domain. The real portion of the dataset consists of 253 Maxar WorldView-3 satellite scenes spanning 112 locations and 2,142km2with 14,700 hand-annotated aircraft. The accompanying synthetic dataset is generated via AI. Reverie's simulation platform and features 50,000 synthetic satellite images simulating a total area of 9331.2km2with ~ 630,000 aircraft annotations. Both the real and synthetically generated aircraft feature 10 fine grain attributes including: aircraft length, wingspan, wing-shape, wing-position, wingspan class, propulsion, number of engines, number of vertical-stabilizers, presence of canards, and aircraft role. Finally, we conduct extensive experiments to evaluate the real and synthetic datasets and compare performances. By doing so, we show the value of synthetic data for the task of detecting and classifying aircraft from an overhead perspective. Jacob Shermeyer, Thomas Hossler, Adam Van Etten, Daniel Hogan, Ryan Lewis, Daeil Kim |
WACV | 6 |