Sebin Lee

dblp:279/6692 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HiGlassRM: Learning to Remove High-prescription Glasses via Synthetic Dataset Generation
abstract
Existing eyeglass removal methods can handle frames and shadows but fail to correct lens-induced geometric distortions, as public datasets lack the necessary supervision. To address this, we introduce the HiGlass Dataset, the first large-scale synthetic dataset providing explicit flow-based supervision for refractive warping. We also propose Hi-GlassRM, a novel pipeline whose core is a network that explicitly estimates a displacement flowmap to de-warp distorted facial geometry. Experiments on both synthetic and real images show that this flowmap-centric approach, trained on our data, significantly improves identity preservation and perceptual quality over existing methods. Our work demonstrates that explicitly modeling and correcting geometric distortion via flowmap estimation, enabled by targeted supervision, is key to faithful eyeglass removal. Project page: https://higlassrm.github.io/
Sebin Lee
WACV1
2025 VTuber's Atelier: The Design Space, Challenges, and Opportunities for VTubing
abstract
VTubing, the practice of live streaming using virtual avatars, has gained worldwide popularity among streamers seeking to maintain anonymity. While previous research has primarily focused on the social and cultural aspects of VTubing, there is a noticeable lack of studies examining the practical challenges VTubers face in creating and operating their avatars. To address this gap, we surveyed VTubers' equipment and expanded the live-streaming design space by introducing six new dimensions related to avatar creation and control. Additionally, we conducted interviews with 16 professional VTubers to comprehensively explore their practices, strategies, and challenges throughout the VTubing process. Our findings reveal that VTubers face significant burdens compared to real-person streamers due to fragmented tools and the multi-tasking nature of VTubing, leading to unique workarounds. Finally, we summarize these challenges and propose design opportunities to improve the effectiveness and efficiency of VTubing.
Daye Kim, Sebin Lee, Yoonseo Jun, Yujin Shin, Jungjin Lee
CHI2
2025 Bring the VibeOn: Designing a Multimodal Interface for Shared Emotional Experiences in Live-streamed Concerts
Gyeongjin Kim, Sebin Lee, Daye Kim, Jungjin Lee
ACM Multimedia2
2024 SemCity: Semantic Scene Generation with Triplane Diffusion
abstract
We present “SemCity,” a 3D diffusion model for semantic scene generation in real-world outdoor environments. Most 3D diffusion models focus on generating a single object, synthetic indoor scenes, or synthetic outdoor scenes, while the generation of real-world outdoor scenes is rarely addressed. In this paper, we concentrate on generating a real-outdoor scene through learning a diffusion model on a realworld outdoor dataset. In contrast to synthetic data, real-outdoor datasets often contain more empty spaces due to sensor limitations, causing challenges in learning realoutdoor distributions. To address this issue, we exploit a triplane representation as a proxy form of scene distributions to be learned by our diffusion model. Furthermore, we propose a triplane manipulation that integrates seamlessly with our triplane diffusion model. The manipulation improves our diffusion model's applicability in a variety of downstream tasks related to outdoor scene generation such as scene inpainting, scene outpainting, and semantic scene completion refinements. In experimental results, we demonstrate that our triplane diffusion model shows meaningful generation results compared with existing work in a real-outdoor dataset, SemanticKITTI. We also show our triplane manipulation facilitates seamlessly adding, removing, or modifying objects within a scene. Further, it also enables the expansion of scenes toward a city-level scale. Finally, we evaluate our method on semantic scene completion refinements where our diffusion model enhances predictions of semantic scene completion networks by learning scene distribution. Our code is available at https://github.com/zoom.in-lee/SemCity.
Jumin Lee, Sebin Lee, Changho Jo, Woobin Im, Juhyeong Seon, Sung-Eui Yoon
CVPR2
2024 Regularizing Dynamic Radiance Fields with Kinematic Fields
Woobin Im, Geonho Cha, Sebin Lee, Jumin Lee, Juhyeong Seon, Dongyoon Wee, Sung-Eui Yoon
ECCV (39)3
2024 Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
abstract
Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audiovisual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense prediction problems, prior works have investigated the introduction of SAM into AVS with audio as a new modality of the prompt. Nevertheless, constrained by SAM’s single-frame segmentation scheme, the temporal context across multiple frames of audio-visual data remains insufficiently utilized. To this end, we study the extension of SAM’s capabilities to the sequence of audio-visual scenes by analyzing contextual cross-modal relationships across the frames. To achieve this, we propose a Spatio-Temporal, Bidirectional Audio-Visual Attention (ST-BAVA) module integrated into the middle of SAM’s image encoder and mask decoder. It adaptively updates the audio-visual features to convey the spatio-temporal correspondence between the video frames and audio streams. Extensive experiments demonstrate that our proposed model outperforms the state-of-the-art methods on AVS benchmarks, especially with an 8.3% mIoU gain on a challenging multi-sources subset.
Juhyeong Seon, Woobin Im, Sebin Lee, Jumin Lee, Sung-Eui Yoon
ICIP3
2024 LiDAR-camera Online Calibration by Representing Local Feature and Global Spatial Context
abstract
LiDAR-camera calibration plays a crucial role in autonomous driving. However, operation-induced factors such as physical vibrations and temperature variations degrade the pre-deployment calibration accuracy, leading to the environmental perception performance deterioration. Recent recalibration methods have achieved online calibration without a target board by leveraging the relative attributes of LiDAR and camera. Nevertheless, we proposes a novel framework for LiDAR-camera online calibration which employs a Transformer network to learn crucial interactions between cameras and LiDAR sensors. Additionally, our novel framework design enables the effective calibration by utilizing correspondence point information between the two sensors. This allows the utilization of global spatial context and achieves high performance by integrating information across modalities. Experimental results indicate that our method demonstrates superior performance compared to state-of-the-art benchmarks.
SeongJoo Moon, Sebin Lee, Sung-Eui Yoon
IROS2
2023 The Effects of Viewing Formats and Song Genres on Audience Experiences in Virtual Avatar Concerts
abstract
With the recent advancements in multimedia and computer graphics technology, virtual avatar concerts have become increasingly popular among both well-known celebrities and subculture-based virtual YouTubers, attracting large audiences. Although this new form of performance art can be delivered through various platforms that use different viewing formats, research on the components of enjoying virtual avatar concerts and their effects on the audience experience is lacking. This study aims to explore the effects of different viewing formats and song genres on sickness, presence, and immersion and identify preferences for enjoying virtual avatar concerts. We first analyzed data from 64 virtual avatar concert cases to identify the most commonly used viewing media, methods, and song genres. Based on our analysis, we designed and conducted a user experiment with four virtual avatar concert scenarios. The experiment compared two viewing methods (audience's point of view and director's edited view) and two song genres (dance and ballad) for 2D display and a head-mounted display. The results show that the audience's point of view and a dance song can be advantageous in increasing the immersive experience irrespective of the viewing medium. Furthermore, the participants show different expectations regarding performances according to song genres. We discuss these findings and propose guidelines to help design and conduct virtual avatar concerts effectively.
Sebin Lee, Daye Kim, Jungjin Lee
ACM Multimedia1
2023 Multi-resolution distillation for self-supervised monocular depth estimation
Sebin Lee, Woobin Im, Sung-Eui Yoon
Pattern Recognit. Lett.1
2022 Semi-supervised Learning of Optical Flow by Flow Supervisor
Woobin Im, Sebin Lee, Sung-Eui Yoon
ECCV (35)2
2021 Dynamic Humanoid Locomotion Over Rough Terrain With Streamlined Perception-Control Pipeline
abstract
Vision aided dynamic exploration on bipedal robots poses an integrated challenge for perception and control. Rapid walking motions as well as the vibrations caused by the landing-foot contact-force introduce critical uncertainty in the visual-inertial system, which can cause the robot to misplace its feet placing on complex terrains and even fall over. In this paper, we present a streamlined integration of an efficient geometric footstep planner and the corresponding walking controller for a humanoid robot to dynamically walk across rough terrain at speeds up to 0.3 m/s. To handle perception uncertainty that arises during dynamic locomotion, we present a geometric safety scoring method in our footstep planner to optimally select feasible path candidates. In addition, the real-time performance of the perception pipeline allows for reactive locomotion such as generating a new corresponding swing leg trajectory in mid-gait if a sudden change in the terrain is detected. The proposed perception-control pipeline is evaluated and demonstrated with real experiments using a full-scale humanoid to traverse across various terrains.
Moonyoung Lee, Youngsun Kwon, Sebin Lee, Jonghun Choe, Junyong Park 0002, Hyobin Jeong, Yujin Heo, Min-Su Kim 0005, Sungho Jo, Sung-Eui Yoon, Jun-Ho Oh
IROS3