EDBT 2026 Demo / reviewers in the wild / expert
Boeun Kim
dblp:148/3443
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SimForce: Force and Surface Electromyography from Full Body Video with Graph Neural NetsabstractWe propose a novel framework, named SimForce, for simultaneously estimating skeletal pose, ground reaction force and surface electromyography from an input video. Simforce predicts the more biomechanically accurate 3D human pose and shape of a given subject, along with their proposed muscle activations and resultant ground reaction force which leads to their input motion. Previous research has either focused on estimating these attributes singly and not treated them as a related task by taking into account the inherent shared motion between the three. In contrast, SimForce is designed to take advantage of the shared biological structure of the human body and its intrinsic connections to infer these attributes jointly using past, current, and future frames. SimForce features a newly introduced temporal and attention aware GCN-based architecture. To learn the subtle links between the body parts and how it affects the distribution of weight on the muscles over time, we introduce the Spatially Aware Attention Module. Esha Dasgupta, Boeun Kim, Sang Hoon Yeo, Hyung Jin Chang |
WACV | 2 |
| 2026 | Bidirectional regression for monocular 6DoF head pose estimation and reference system alignment
Sungho Chun, Boeun Kim, Hyung Jin Chang, Ju Yong Chang |
Pattern Recognit. | 2 |
| 2025 | 3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze Estimationabstract3D and 2D gaze estimation share the fundamental objective of capturing eye movements but are traditionally treated as two distinct research domains. In this paper, we introduce a novel cross-task few-shot 2D gaze estimation approach, aiming to adapt a pre-trained 3D gaze estimation network for 2D gaze prediction on unseen devices using only a few training images. This task is highly challenging due to the domain gap between 3D and 2D gaze, unknown screen poses, and limited training data. To address these challenges, we propose a novel framework that bridges the gap between 3D and 2D gaze. Our framework contains a physics-based differentiable projection module with learnable parameters to model screen poses and project 3D gaze into 2D gaze. The framework is fully differentiable and can integrate into existing 3D gaze networks without modifying their original architecture. Additionally, we introduce a dynamic pseudo-labelling strategy for flipped images, which is particularly challenging for 2D labels due to unknown screen poses. To overcome this, we reverse the projection process by converting 2D labels to 3D space, where flipping is performed. Notably, this 3D space is not aligned with the camera coordinate system, so we learn a dynamic transformation matrix to compensate for this misalignment. We evaluate our method on MPIIGaze, EVE, and GazeCapture datasets, collected respectively on laptops, desktop computers, and mobile devices. The superior performance highlights the effectiveness of our approach, and demonstrates its strong potential for real-world applications. Yihua Cheng, Hengfei Wang, Zhongqun Zhang, Boeun Kim, Feng Lu 0005, Hyung Jin Chang |
CVPR | 5 |
| 2025 | PersonaBooth: Personalized Text-to-Motion GenerationabstractThis paper introduces Motion Personalization, a new task that generates personalized motions aligned with text descriptions using several basic motions containing Persona. To support this novel task, we introduce a new large-scale motion dataset called PerMo (PersonaMotion), which captures the unique personas of multiple actors. We also propose a multi-modal finetuning method of a pretrained motion diffusion model called PersonaBooth. PersonaBooth addresses two main challenges: i) A significant distribution gap between the persona-focused PerMo dataset and the pretraining datasets, which lack persona-specific data, and ii) the difficulty of capturing a consistent persona from the motions vary in content (action type). To tackle the dataset distribution gap, we introduce a persona token to accept new persona features and perform multi-modal adaptation for both text and visuals during finetuning. To capture a consistent persona, we incorporate a contrastive learning technique to enhance intra-cohesion among samples with the same persona. Furthermore, we introduce a context-aware fusion mechanism to maximize the integration of persona cues from multiple input motions. PersonaBooth outperforms state-of-the-art motion style transfer methods, establishing a new benchmark for motion personalization. Boeun Kim, Hea In Jeong, JungHoon Sung, Yihua Cheng, Jeongmin Lee 0007, Ju Yong Chang, Sang-Il Choi, Younggeun Choi 0001, Saim Shin, Hyung Jin Chang |
CVPR | 1 |
| 2025 | High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationabstractModeling high-resolution spatiotemporal representations, including both global dynamic contexts (e.g., holistic human motion tendencies) and local motion details (e.g., high-frequency changes of keypoints), is essential for video-based human pose estimation (VHPE). Current state-of-the-art methods typically unify spatiotemporal learning within a single type of modeling structure (convolution or attention-based blocks), which inherently have difficulties in balancing global and local dynamic modeling and may bias the network to one of them, leading to suboptimal performance. Moreover, existing VHPE models suffer from quadratic complexity when capturing global dependencies, limiting their applicability especially for high-resolution sequences. Recently, the state space models (known as Mamba) have demonstrated significant potential in modeling long-range contexts with linear complexity; however, they are restricted to 1D sequential data. In this paper, we present a novel framework that extends Mamba from two aspects to separately learn global and local high-resolution spatiotemporal representations for VHPE. Specifically, we first propose a Global Spatiotemporal Mamba, which performs 6D selective space-time scan and spatial- and temporal-modulated scan merging to efficiently extract global representations from high-resolution sequences. We further introduce a windowed space-time scan-based Local Refinement Mamba to enhance the high-frequency details of localized keypoint motions. Extensive experiments on four benchmark datasets demonstrate that the proposed model outperforms state-of-the-art VHPE approaches while achieving better computational trade-offs. Runyang Feng, Hyung Jin Chang, Tze Ho Elden Tse, Boeun Kim, Yi Chang 0001, Yixing Gao 0001 |
ICCV | 4 |
| 2025 | Roll Your Eyes: Gaze Redirection via Explicit 3D Eyeball RotationabstractWe propose a novel 3D gaze redirection framework that leverages an explicit 3D eyeball structure. Existing gaze redirection methods are typically based on neural radiance fields, which employ implicit neural representations via volume rendering. Unlike these NeRF-based approaches, where the rotation and translation of 3D representations are not explicitly modeled, we introduce a dedicated 3D eyeball structure to represent the eyeballs with 3D Gaussian Splatting (3DGS). Our method generates photorealistic images that faithfully reproduce the desired gaze direction by explicitly rotating and translating the 3D eyeball structure. In addition, we propose an adaptive deformation module that enables the replication of subtle muscle movements around the eyes. Through experiments conducted on the ETH-XGaze dataset, we demonstrate that our framework is capable of generating diverse novel gaze images, achieving superior image quality and gaze estimation accuracy compared to previous state-of-the-art methods. YoungChan Choi, HengFei Wang, YiHua Cheng, Boeun Kim, Hyung Jin Chang, Younggeun Choi 0001, Sang-Il Choi |
ACM Multimedia | 4 |
| 2025 | A unified framework for unsupervised action learning via global-to-local motion transformer
Boeun Kim, Hyung Jin Chang, Tae-Hyun Oh |
Pattern Recognit. | 1 |
| 2024 | MoST: Motion Style Transformer Between Diverse Action ContentsabstractWhile existing motion style transfer methods are effective between two motions with identical content, their performance significantly diminishes when transferring style between motions with different contents. This challenge lies in the lack of clear separation between content and style of a motion. To tackle this challenge, we propose a novel motion style transformer that effectively disentangles style from content and generates a plausible motion with transferred style from a source motion. Our distinctive approach to achieving the goal of disentanglement is twofold: (1) a new architecture for motion style transformer with 'part-attentive style modulator across body parts' and ‘Siamese encoders that encode style and content features separately’; (2) style disentanglement loss. Our method outperforms existing methods and demonstrates exceptionally high quality, particularly in motion pairs with different contents, without the need for heuristic post-processing. Codes are available at https://github.com/Boeun-Kim/MoST. Boeun Kim, Hyung Jin Chang, Jin Young Choi 0002 |
CVPR | 1 |
| 2023 | Omission-Free Inpainting: A Three-Stage Approach to Ensure Object GenerationabstractThis paper proposes a novel inpainting framework, omission-free inpainting, which ensures generating the desired object in the masked region. Despite recent advancements in text-driven and class-conditional inpainting models, they often fail to restore the missing object. To address this issue, the proposed framework includes a separate object generation stage, resulting in omission-free inpainting. The framework consists of three stages: background generation, object generation and refinement. The background generation stage restores a harmonious background with the surrounding pixels, while the object generation stage creates the desired object using a blending mask that allows the object to be influenced by the background’s color and brightness. Finally, the refinement stage blends the object and background to produce a visually realistic image. We compare the results qualitatively with the state-of-the-art methods, and our method outperforms the existing methods in CLIP score. Hea In Jeong, Boeun Kim, Chungil Kim, Saim Shin |
ICIP | 2 |
| 2022 | Global-Local Motion Transformer for Unsupervised Skeleton-Based Action Learning
Boeun Kim, Hyung Jin Chang, Jin Young Choi 0002 |
ECCV (4) | 1 |
| 2022 | Learning spectral transform for 3D human motion prediction
Boeun Kim, Jin Young Choi 0002 |
Comput. Vis. Image Underst. | 1 |
| 2012 | FaceReview: Supporting Interactive Exploration of Linked Heterogeneous Datasets for Unilateral Cleft Lip and Palate
Jinwook Seo, Boeun Kim, Bongshin Lee, Bo Hyoung Kim, Bohyung Han, Nina Anderson, Richard Bruun, Stephen Shusterman |
AMIA | 3 |