EDBT 2026 Demo / reviewers in the wild / expert
Seongmin Lee 0002
dblp:186/1126-2
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-1564-5077ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Domain Crossover Non-Rigid Registration for 3D Human MeshesabstractNon-rigid registration is essential for reconstructing dynamic and incomplete 3D human meshes, yet traditional methods often fail to achieve robust alignment in the sequence of high-motion deformations and missing geometry. We propose a domain crossover non-rigid registration (DCNRR) framework that addresses these challenges by effectively transferring informative features from 2D image space into the 3D mesh domain of three key stages: multi-view projection, hierarchical non-rigid registration, and topology-consistent completion. In the first stage, multi-view projections are used to extract 2D joint locations and deep features, which guide deformation in the 3D space. In the second stage, hierarchical joint priors and deep features collaboratively guide mesh alignment, enabling more accurate deformation in distal regions and complex poses. In the final stage, we apply a diffusion-based completion process in UV coordinates to reconstruct incomplete surface normals and refine missing mesh areas with topological consistency. Our approach achieves highly detailed and perceptually accurate mesh deformation. To validate our approach, we evaluate performance on a newly constructed dynamic human motion (DHM) dataset, as well as public datasets. Our method demonstrates state-of-the-art results in both geometric accuracy and stability, showing particular robustness in dynamic and incomplete mesh sequences. Kyungjune Lee, Seongjean Kim, Hoseok Tong, Hyucksang Lee, Seongmin Lee 0002, Weisi Lin, Ping An 0001, Sanghoon Lee 0001 |
ACM Multimedia | 5 |
| 2025 | Visibility-Aware Multi-View Stereo by Surface Normal Weighting for Occlusion RobustnessabstractRecent learning-based multi-view stereo (MVS) still exhibits insufficient accuracy in large occlusion cases, such as environments with significant inter-camera distance or when capturing objects with complex shapes. This is because incorrect image features extracted from occluded areas serve as significant noise in the cost volume construction. To address this, we propose a visibility-aware MVS using surface normal weighting (SnowMVSNet) based on explicit 3D geometry. It selectively suppresses mismatched features in the cost volume construction by computing inter-view visibility. Additionally, we present a geometry-guided cost volume regularization that enhances true depth among depth hypotheses using a surface normal prior. We also propose intra-view visibility that distinguishes geometrically more visible pixels within a reference view. Using intra-view visibility, we introduce the visibility-weighted training and depth estimation methods. These methods enable the network to achieve accurate 3D point cloud reconstruction by focusing on visible regions. Based on simple inter-view and intra-view visibility computations, SnowMVSNet accomplishes substantial performance improvements relative to computational complexity, particularly in terms of occlusion robustness. To evaluate occlusion robustness, we constructed a multi-view human (MVHuman) dataset containing general human body shapes prone to self-occlusion. Extensive experiments demonstrated that SnowMVSNet significantly outperformed state-of-the-art methods in both low- and high-occlusion scenarios. Hyucksang Lee, Seongmin Lee 0002, Sanghoon Lee 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | DMESH: A Structure-Preserving Diffusion Model for 3-D Mesh DenoisingabstractDenoising diffusion models have shown a powerful capacity for generating high-quality image samples by progressively removing noise. Inspired by this, we present a diffusion-based mesh denoiser that progressively removes noise from mesh. In general, the iterative algorithm of diffusion models attempts to manipulate the overall structure and fine details of target meshes simultaneously. For this reason, it is difficult to apply the diffusion process to a mesh denoising task that removes artifacts while maintaining a structure. To address this, we formulate a structure-preserving diffusion process. Instead of diffusing the mesh vertices to be distributed as zero-centered isotopic Gaussian distribution, we diffuse each vertex into a specific noise distribution, in which the entire structure can be preserved. In addition, we propose a topology-agnostic mesh diffusion model by projecting the vertex into multiple 2-D viewpoints to efficiently learn the diffusion using a deep network. This enables the proposed method to learn the diffusion of arbitrary meshes that have an irregular topology. Finally, the denoised mesh can be obtained via refinement based on 2-D projections obtained from reverse diffusion. Through extensive experiments, we demonstrate that our method outperforms the state-of-the-art mesh denoising methods in both quantitative and qualitative evaluations. Seongmin Lee 0002, Suwoong Heo, Sanghoon Lee 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | 3D Facial Shape Similarity with Deep Perceptual RepresentationsabstractComparing different 3D shapes is challenging due to their irregularities. Motivated by the human visual system mechanism, where the entire 3D geometry is clearly perceived as a series of multiple projections, we propose a novel facial shape similarity measurement using multiview deep perceptual representations. We introduce a multiview disentangling scheme that accurately represents a facial mesh in multiple coordinates and the training strategy with view specificity and regional consistency to reliably train the network with multiple projections. View specificity pertains to the human visual perception to better recognize facial similarity. Regional consistency mitigates regional redundancy among views. Hence, robust perceptual features with respect to views are embedded and accurate similarity can be measured. Consequently, the view-specific integration scheme incorporates the similarities of all views, allowing for highly consistent measurement. The experiments demonstrate that the proposed similarity outperforms state-of-the-arts and significantly improves the details in terms of geometry and human perception. Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Speech-Driven Emotional 3d Talking Face Animation Using Emotional EmbeddingsabstractExisting emotional talking 3D facial animation primarily focus on animating emotional faces using a specific emotion condition. However, in real-world situations, no one consistently speaks with just one emotion. Thus, previous emotion-based approaches have very limited applicability in real-world applications. To address this issue, we propose SDETalk, a novel learning framework that animates the emotional talking faces by leveraging the emotional source from a speech. Unlike previous studies, which use static one-hot emotion conditions, the proposed network regresses complex emotional states from speech. It enables the network to animate natural facial animation from an emotional speech without using a specific emotional condition. Furthermore, we design the proposed method to produce head motions because head motion is an important factor to enhance the naturalness of talking face animation. By doing this, our approach simultaneously achieves accurate lip motion, natural expressions, and rhythmical head motions from emotional speech. Through extensive experiments in both qualitative and quantitative manners, it is demonstrated that our method outperforms other state-of-the-art methods by animating realistic and expressive 3D faces. Seongmin Lee 0002, Jeonghaeng Lee, Hyewon Song, Sanghoon Lee 0001 |
ICASSP | 1 |
| 2024 | InViTe: Individual Virtual Transfer for Personalized 3D Face Generation System
Mingyu Jang, Kyungjune Lee, Seongmin Lee 0002, Hoseok Tong, Juwan Chung, Yusung Ro, Sanghoon Lee 0001 |
IJCAI | 3 |
| 2024 | AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning
Chanhyuk Park, Jungbin Cho, Junwan Kim, Seongmin Lee 0002, Jungsu Kim, Sanghoon Lee 0001 |
IJCAI | 4 |
| 2024 | 3D-PSSIM: Projective Structural Similarity for 3D Mesh Quality Assessment Robust to Topological IrregularitiesabstractDespite acceleration in the use of 3D meshes, it is difficult to find effective mesh quality assessment algorithms that can produce predictions highly correlated with human subjective opinions. Defining mesh quality features is challenging due to the irregular topology of meshes, which are defined on vertices and triangles. To address this, we propose a novel 3D projective structural similarity index ( 3D- PSSIM) for meshes that is robust to differences in mesh topology. We address topological differences between meshes by introducing multi-view and multi-layer projections that can densely represent the mesh textures and geometrical shapes irrespective of mesh topology. It also addresses occlusion problems that occur during projection. We propose visual sensitivity weights that capture the perceptual sensitivity to the degree of mesh surface curvature. 3D- PSSIM computes perceptual quality predictions by aggregating quality-aware features that are computed in multiple projective spaces onto the mesh domain, rather than on 2D spaces. This allows 3D- PSSIM to determine which parts of a mesh surface are distorted by geometric or color impairments. Experimental results show that 3D- PSSIM can predict mesh quality with high correlation against human subjective judgments, across the presence of noise, even when there are large topological differences, outperforming existing mesh quality assessment models. Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001, Weisi Lin, Alan C. Bovik |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Fusing Explicit and Implicit Flow for Optical Flow EstimationabstractEstimating optical flow for large movement remains a challenging issue due to inconsistency in features between frames. To resolve this challenge, we propose a novel sequence-based deep learning network that jointly trains explicit flow and implicit flow to accurately estimate optical flow. To do this, we implemented three submodules: explicit flow embedder, implicit flow embedder, and flow fusion network. Explicit flow embedder learns the pair-wise correlation between visible pixels based on the spatial attention made per image. Implicit flow embedder learns implicit flow based on the temporal context of motion from all frames in the sequence. To effectively learn the implicit flow, we give a longer sequence of frames as input. Flow fusion network fuses features from explicit and implicit embedder to output the final optical flow. Through extensive experiments, our model demonstrates its robustness against the large motion while providing accurate flow estimation for pixels without pairs in the next frame. Hyunse Yoon, Seongmin Lee 0002, Sanghoon Lee 0001 |
ICIP | 2 |
| 2023 | Video-Based Stabilized 3D Face Alignment Using Temporal Multi-DiscriminationabstractExisting 3D face alignment primarily aim to achieve accurate face alignment result for a static facial image. While these methods have strong alignment performance under large poses, occlusion, and extreme lighting conditions, they often result in trembling artifacts in video-based sequential 3D face alignment. Reducing temporal misalignment remains a challenging task because a single misaligned frame can propagate errors to other frames along the temporal axis. To address this issue, we propose a novel temporal discriminating scheme that learns the distribution gap between the face alignment results and ground truth face animation. By leveraging the discrimination results as a guide, the proposed method can effectively align the 3D faces to the input video by reducing temporal trembling artifacts. To effectively learn the distribution gap, we introduce a multi-discriminating scheme that separately discriminates facial animation based on identity and expression changes. It enables the proposed method to produce a stabilized alignment result, especially in dynamic and fast movement. Through extensive experiments in both qualitative and quantitative evaluations, it is confirmed that our method outperforms state-of-the-art 3D face alignment methods by animating stabilized results in the video. Seongmin Lee 0002, Hyunse Yoon, Jiwoo Kang 0001, Jungsu Kim, Jiwan Son, Jungwoo Huh, Sanghoon Lee 0001 |
MMSP | 1 |
| 2022 | Gradient Flow Evolution for 3D Fusion From a Single Depth SensorabstractWe present a novel real-time framework for non-rigid 3D reconstruction that is robust to noise, camera poses, and large deformation from a single depth camera. KinectFusion has achieved high-quality 3D object reconstructions in real-time by implicitly representing an object’s surface with a signed distance field (SDF) representation from a single depth camera. Many studies for incremental reconstruction have been presented since then, with the surface estimation improving over time. Previous works primarily focused on improving conventional SDF matching and deformation schemes. In contrast to these works, the proposed framework tackles the problem of temporal inconsistency caused by SDF approximation and fusion to manipulate SDFs and reconstruct a target more accurately over time. In our reconstruction pipeline, we introduce a refinement evolution method, where an erroneous SDF from a depth sensor is recovered more accurately in a few iterations by propagating erroneous SDF values from the surface. Reliable gradients of refined SDFs enable more accurate non-rigid tracking of a target object. Furthermore, we propose a level-set evolution for SDF fusion, enabling SDFs to be manipulated stably in the reconstruction pipeline over time. The proposed methods are fully parallelizable and can be executed in real-time. Qualitative and quantitative evaluations show that incorporating the refinement and fusion methods into the reconstruction pipeline improves 3D reconstruction accuracy and temporal reliability by avoiding cumulative errors over time. Evaluation results show that our pipeline results in more accurate reconstruction that is robust to noise and large motions, as well as outperforms previous state-of-the-art reconstruction methods. Jiwoo Kang 0001, Seongmin Lee 0002, Mingyu Jang, Sanghoon Lee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Competitive Learning of Facial Fitting and Synthesis Using UV EnergyabstractThe three-dimensional morphable model (3DMM) is the most widely used representative model for obtaining a three-dimensional (3-D) face from a target on an image. Although 3DMMs have demonstrated the powerful capability to represent various facial shapes on natural images, they are limited to capturing texture variations of in-the-wild human faces. Based on the fact that fitting a 3-D facial model to an image determines the corresponding UV map, we propose a novel method for facial fitting and synthesis by competitively training two deep learning networks for facial alignment and UV texture completion. When the completion network is trained using well-aligned UV maps, it can model facial textures precisely and, consequently, fill the missing regions more completely. Accordingly, we use a UV completion network, denoted as a UV energy-based generative adversarial network (UV EB-GAN), to discriminate whether a UV map from the alignment network is well aligned by defining the generative loss of the completion network as the energy. Competitive learning facilitates training the completion network without ground-truth facial UV maps and training the alignment network without hard constraints and regularization terms. The proposed network can be trained in an end-to-end manner. The facial texture, albedo, lighting parameters, and 3-D facial shape can be obtained through this network. The results of the experiments on 2-D alignment, 3-D reconstruction, texture synthesis, and illumination estimation verified that the proposed method achieves remarkable improvements over the state-of-the-art methods. Jiwoo Kang 0001, Seongmin Lee 0002, Sanghoon Lee 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | WarpingFusion: Accurate Multi-View TSDF Fusion with Local Perspective WarpabstractIn this paper, we propose the novel 3D reconstruction framework, where the surface of a target object is reconstructed accurately and robustly from multi-view depth maps. A depth map of a moving object tends to have the spatially-varying perspective warps due to motion blur and rolling shutter artifacts. Incorporating those misaligned points from the views into the world coordinate leads to significant artifacts in the reconstructed shape. We address the mismatches by the patch-based depth-to-surface alignment using implicit surface-based distance measurement. The patch-based minimization finds spatial warps on the depth map fast and accurately with the global transformation preserved. The proposed framework efficiently optimizes the local alignments against depth occlusions and local variants thanks to the point to surface distance based on an implicit representation. The proposed method shows significant improvements over the other reconstruction methods, demonstrating efficiency and benefits of our method in the multi-view reconstruction. Jiwoo Kang 0001, Seongmin Lee 0002, Mingyu Jang, Hyunse Yoon, Sanghoon Lee 0001 |
ICIP | 2 |
| 2021 | Deep Chessboard Corner Detection Using Multi-task LearningabstractCamera calibration is an indispensable step in the fields of robotics and computer vision, which includes augmented reality, 3D reconstruction, and camera motion estimation. Before camera calibration, detecting matching correspondence is necessary to understand the structure of the world from multiple images. For an accurate result, a calibration object, such as a chessboard, is used. Existing handcrafted feature methods precisely detect chessboard corners but are weak against blurs, noises, and severe lens distortion. Conversely, neural network-based methods can detect corners regardless of noises in the image. Both methods do not utilize the information of camera priors, which are lens distortion and intrinsic parameters, affecting the location of chessboard corners. Learning of lens distortion and intrinsic parameters enables the proposed network to understand the alignment of corners more precisely. Therefore, in this paper, we propose a novel multi-task learning framework to detect chessboard corners and simultaneously estimate lens distortion and intrinsic parameters. In order to train these three tasks, synthetic images of the chessboard are generated with ground-truth labels corresponding to each task. Hence, by learning the camera priors, the proposed network can more precisely locate the corners than other state-of-the-art corner detection methods while robust to noises, blurs, and distortion. Hyunse Yoon, Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001 |
MMSP | 2 |
| 2019 | A Deep Cybersickness Predictor Based on Brain Signal Analysis for Virtual Reality ContentsabstractWhat if we could interpret the cognitive state of a user while experiencing a virtual reality (VR) and estimate the cognitive state from a visual stimulus? In this paper, we address the above question by developing an electroencephalography (EEG) driven VR cybersickness prediction model. The EEG data has been widely utilized to learn the cognitive representation of brain activity. In the first stage, to fully exploit the advantages of the EEG data, it is transformed into the multi-channel spectrogram which enables to account for the correlation of spectral and temporal coefficient. Then, a convolutional neural network (CNN) is applied to encode the cognitive representation of the EEG spectrogram. In the second stage, we train a cybersickness prediction model on the VR video sequence by designing a Recurrent Neural Network (RNN). Here, the encoded cognitive representation is transferred to the model to train the visual and cognitive features for cybersickness prediction. Through the proposed framework, it is possible to predict the cybersickness level that reflects brain activity automatically. We use 8-channels EEG data to record brain activity while more than 200 subjects experience 44 different VR contents. After rigorous training, we demonstrate that the proposed framework reliably estimates cognitive states without the EEG data. Furthermore, it achieves state-of-the-art performance comparing to existing VR cybersickness prediction models. Jinwoo Kim 0005, Woojae Kim, Heeseok Oh, Seongmin Lee 0002, Sanghoon Lee 0001 |
ICCV | 4 |