EDBT 2026 Demo / reviewers in the wild / expert
I-Chen Lin
dblp:11/3246
· DBLP profile ↗
24ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0001-9924-4723ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exemplar-based image colorization with awareness of object co-saliencyabstractAbstract While exemplar-based colorization colorizes a target image according to the chromatic information of a given reference image, we observed that existing methods were easily disturbed by colors of multiple objects in a reference image. To tackle the issue, we propose enhancing the exemplar-based colorization process according to the correlation among salient objects within the target and reference images. The proposed framework predicts the regions of common objects that occur in both the target and reference images. It first focuses on colorizing the co-salient objects and the remaining region, respectively. Based on the above preliminary features, our region-aware module with attention mechanism, which can moderately tolerate imperfect predicted regions, then progressively colorizes the whole image. Experiments show that the proposed framework can substantially alleviate the color disturbance from unrelated objects and generate more precise colors, especially for salient common objects. An extended metric is presented to complement the limitation of existing metrics for exemplar-based colorization. Moreover, while we extend our framework to colorize only specific salient regions of an image, it can become an intelligent editing tool for exemplar-based colorization. Yi-Hua Chiu, Keng-Hao Chang, I-Chen Lin |
Multim. Tools Appl. | 3 |
| 2025 | Human-MoE: Multimodal Full-Body Human Image Synthesis with Component-driven Mixture of ExpertsabstractConditional full-body human synthesis is to generate and edit realistic images based on given conditions. Previous methods lay a solid foundation but may have limitations in adjusting human poses and appearance. They usually adopt a monolithic design, and the details of generated images are prone to be indistinct or distorted due to the high variation of human appearances. To tackle the above-mentioned challenges, we propose Human-MoE for multi-modal full-body human synthesis. Users can control the image generation through three types of input representations: parsing maps for geometry, text descriptions for appearance attributes, and pose maps to distinguish postures. Our framework specifically designs a mixture-of-experts module to capture and synthesize details in specific regions with high fidelity. These synthesized details are then applied to refine the appearance. Our method achieves top scores in experiments by multiple metrics, especially FID and SSIM, demonstrating its advance in visual quality and controllability. Yu-Jiu Huang, I-Chen Lin |
ICME | 2 |
| 2024 | Estimation of Hand-Interacting Object Poses with Boundary Guidance
Sin-Yu Fu, I-Chen Lin |
ICPR (17) | 2 |
| 2024 | Hairstyle-and-identity-aware facial image style transfer with region-guiding masks
Hsin-Ying Wang, Chiu-Wei Chien, Ming-Han Tsai, I-Chen Lin |
Multim. Tools Appl. | 4 |
| 2024 | Characteristic-Preserving Latent Space for Unpaired Cross-Domain Translation of 3D Point CloudsabstractThis article aims at unpaired shape-to-shape transformation for 3D point clouds, for instance, turning a chair to its table counterpart. Recent work for 3D shape transfer or deformation highly relies on paired inputs or specific correspondences. However, it is usually not feasible to assign precise correspondences or prepare paired data from two domains. A few methods start to study unpaired learning, but the characteristics of a source model may not be preserved after transformation. To overcome the difficulty of unpaired learning for transformation, we propose alternately training the autoencoder and translators to construct shape-aware latent space. This latent space based on novel loss functions enables our translators to transform 3D point clouds across domains and maintain the consistency of shape characteristics. We also crafted a test dataset to objectively evaluate the performance of point-cloud translation. The experiments demonstrate that our framework can construct high-quality models and retain more shape characteristics during cross-domain translation compared to the state-of-the-art methods. Moreover, we also present shape editing applications with our proposed latent space, including shape-style mixing and shape-type shifting, which do not require retraining a model. Jia-Wen Zheng, Jhen-Yung Hsu, Chih-Chia Li, I-Chen Lin |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Efficient Video Matting on Human Video Clips for Real-Time ApplicationabstractThis paper presents an efficient and effective matting framework for human video clips. To alleviate the inefficiency problem in existing models, we propose using a refiner dedicated to error-prone regions, and reduce the computation at higher resolutions, so the proposed framework can achieve real-time performance for 1080p 60fps videos. Also, with the recurrent architecture, our model is aware of temporal information and produces temporally more consistent matting results compared to models processing each frame individually. Moreover, it contains a module for capturing semantic information. That makes our model easy to use without troublesome setup, such as annotating trimaps or other additional inputs. Experiments show that our proposed method outperforms previous matting methods, and reaches the state of the art on the VideoMatte240K dataset. Chao-Liang Yu, I-Chen Lin |
ICME | 2 |
| 2023 | Generalizable person re-identification with part-based multi-scale network
Jia-Jen Wu, Keng-Hao Chang, I-Chen Lin |
Multim. Tools Appl. | 3 |
| 2022 | Confidence-Based 6D Object Pose EstimationabstractThe aim of this paper is to estimate the six-degree-of-freedom (6DOF) poses of objects from a single RGB image in which the target objects are partially occluded. Most recent studies have formulated methods for predicting the projected two-dimensional (2D) locations of three-dimensional keypoints through a deep neural network and then used a PnP algorithm to compute the 6DOF poses. Several researchers have pointed out the uncertainty of the predicted locations and modelled it according to predefined rules or functions, but the performance of such approaches may still be degraded if occlusion is present. To address this problem, we formulated 2D keypoint locations as probabilistic distributions in our novel loss function and developed a confidence-based pose estimation network. This network not only predicts the 2D keypoint locations from each visible patch of a target object but also provides the corresponding confidence values in an unsupervised fashion. Through the proper fusion of the most reliable local predictions, the proposed method can improve the accuracy of pose estimation when target objects are partially occluded. Experiments demonstrated that our method outperforms state-of-the-art methods on a main occlusion data set used for estimating 6D object poses. Moreover, this framework is efficient and feasible for realtime multimedia applications. Chun-Yi Hung, I-Chen Lin |
IEEE Trans. Multim. | 3 |
| 2018 | Enhancing the Realism of Sketch and Painted Portraits With Adaptable PatchesabstractAbstract Realizing unrealistic faces is a complicated task that requires a rich imagination and comprehension of facial structures. When face matching, warping or stitching techniques are applied, existing methods are generally incapable of capturing detailed personal characteristics, are disturbed by block boundary artefacts, or require painting‐photo pairs for training. This paper presents a data‐driven framework to enhance the realism of sketch and portrait paintings based only on photo samples. It retrieves the optimal patches of adaptable shapes and numbers according to the content of the input portrait and collected photos. These patches are then seamlessly stitched by chromatic gain and offset compensation and multi‐level blending. Experiments and user evaluations show that the proposed method is able to generate realistic and novel results for a moderately sized photo collection. Yin-Hsuan Lee, Yu-Kai Chang, Yu-Lun Chang, I-Chen Lin, Yu-Shuen Wang, Wen-Chieh Lin |
Comput. Graph. Forum | 4 |
| 2016 | Augmented reality instruction for object assembly based on markerless trackingabstractConventional object assembly instructions are usually written or illustrated in a paper manual. Users have to associate these static instructions with real objects in 3D space. In this paper, a novel augmented reality system is presented for a user to interact with objects and instructions. While most related methods pasted obvious markers onto objects for tracking and constrained their orientations or shapes, we adopt a markerless strategy for more intuitive interaction. Based on live information from an off-the-shelf RGB-D camera, the proposed tracking procedure identifies components in a scene, tracks their 3D positions and orientations, and evaluates whether there are combinations of components. According to the detected events and poses, our indication procedure then dynamically displays indication lines, circular arrows and other hints to guide a user to manipulate the components into correct poses. The experiment shows that the proposed system can robustly track the components and respond intuitive instructions at an interactive rate. Most of users in evaluation are interested and willing to use this novel technique for object assembly. Li-Chen Wu, I-Chen Lin, Ming-Han Tsai |
I3D | 2 |
| 2015 | Real-time upper body pose estimation from depth imagesabstractEstimating upper body poses from a sequence of depth images is a challenging problem. Lately, the state-of-art work adopted a randomized forest method to label human parts in real time. However, it requires enormous training data to obtain favorable results. In this paper, we propose using a novel two-stage method to estimate the probability maps of upper body parts of users. These maps are then used to evaluate the region fitness of body parts for pose recovery. Experiments show that the proposed method can obtain satisfactory outcome in real time and it requires a moderate size of training data. Ming-Han Tsai, Kuan-Hua Chen, I-Chen Lin |
ICIP | 3 |
| 2015 | Interactive Visual Analysis for Vehicle Detector DataabstractAbstract Visualization of vehicle detection (VD) data is essential because the data play an important role in traffic control and policy development. Most previous works focus on visualizing trajectories obtained from global positioning system (GPS), which are detailed but less representative. In contrast, VD data report the traffic statistic at each sensing site during a time span, including speed, flow, and occupancy of each lane, which contain comprehensive traffic information for analysis. In this work, we visualize three‐year VD data of freeways in Taiwan. The visualization depicts the traffic situation at a site over time using a color‐coded chart that extends from left to right over time. The charts are vertically stacked and horizontally aligned according to VD's located mileage and data time, respectively, to provide global insight. Our system allows semantic zoom, which changes the chart appearance in a continuous manner, to enable macro‐ and micro‐ scopic visualizations. Analysts can explore events that span an area with different sizes and that persist a time span with various lengths. To ensure the feasibility of our visualization, before the system design, we conducted a study with experts who work in the national freeway bureau and the institute of transportation of Taiwan. We also showed our results to the experts after the prototype system was built. The feedback shows that our VD data visualization is helpful to traffic control and policy development. Yu-Shuen Wang, Wen-Chieh Lin, Wei-Xiang Huang, I-Chen Lin |
Comput. Graph. Forum | 5 |
| 2015 | SI-Cut: Structural Inconsistency Analysis for Image Foreground ExtractionabstractThis paper presents a novel approach for extracting foreground objects from an image. Existing methods involve separating the foreground and background mainly according to their color distributions and neighbor similarities. This paper proposes using a more discriminative strategy, structural inconsistency analysis, in which the localities of color and texture are considered. Given an indicated rectangle, the proposed system iteratively maximizes the consensus regions between the original image and predicted structures from the known background. The object contour can then be extracted according to inconsistency in the predicted background and foreground structures. The proposed method includes an efficient image completion technique for structural prediction. The results of experiments showed that the extraction accuracy of the proposed method is higher than that of related methods for structural scenes, and is also comparable to that of related methods for less structural situations. I-Chen Lin, Yu-Chien Lan, Po-Wen Cheng |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | Human face aging with guided prediction and detail synthesis
Ming-Han Tsai, Yen-Kai Liao, I-Chen Lin |
Multim. Tools Appl. | 3 |
| 2013 | Markerless 3D Hand Posture Estimation from Monocular Video by Two-Level SearchingabstractIn this paper, a marker less 3D hand tracking system for monocular RGB video is presented. We propose a novel two-level approach to efficiently grasp the personal characteristics and high varieties of hand postures. Our system first searches the approximate nearest neighbors in a small personalized real-hand image set, and retrieves more details from a large synthetic 3D hand posture database. Temporal consistency property is also utilized for disambiguating and noise reduction. Our prototype system can approximate hand poses including rigid and non-rigid out-of-image-plane rotation, slow and fast gesture changing during rotation. It can also recover from a short-term missing hand situation in an interactive rate. Iek-Kuong Pun, I-Chen Lin, Tsung-Hsien Tang |
CAD/Graphics | 2 |
| 2013 | Skeleton-driven surface deformation through lattices for real-time character animation
Cheng-Hao Chen, Ming-Han Tsai, I-Chen Lin, Pin-Hua Lu |
Vis. Comput. | 3 |
| 2011 | Lattice-Based Skinning and Deformation for Real-Time Skeleton-Driven AnimationabstractIn this paper, we present an efficient framework to deform polygonal models for skeleton-driven animation. Standard solutions of skeleton-driven animation, such as linear blend skinning, require intensive artist intervention and focus on primary deformations. The proposed approach can generate both low- and high-frequency surface motions such as muscle deformation and vibrations with little user intervention. Given a surface mesh, we construct a lattice of cubic cells embracing the mesh and we apply lattice-based smooth skinning to drive the surface primary deformation with volume preservation. Lattice shape matching with dynamic particles, in the meantime, is utilized for secondary deformations. Due to the highly parallel lattice structure, the proposed method is liable to GPU computation. Our results show that it is adequate to vividly real-time animation. Cheng-Hao Chen, I-Chen Lin, Ming-Han Tsai, Pin-Hua Lu |
CAD/Graphics | 2 |
| 2011 | Adaptive Motion Data Representation with Repeated Motion AnalysisabstractIn this paper, we present a representation method for motion capture data by exploiting the nearly repeated characteristics and spatiotemporal coherence in human motion. We extract similar motion clips of variable lengths or speeds across the database. Since the coding costs between these matched clips are small, we propose the repeated motion analysis to extract the referred and repeated clip pairs with maximum compression gains. For further utilization of motion coherence, we approximate the subspace-projected clip motions or residuals by interpolated functions with range-aware adaptive quantization. Our experiments demonstrate that the proposed feature-aware method is of high computational efficiency. Furthermore, it also provides substantial compression gains with comparable reconstruction and perceptual errors. I-Chen Lin, Jen-Yu Peng, Chao-Chih Lin, Ming-Han Tsai |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2010 | Image-based detail reconstruction of non-Lambertian surfacesabstractAbstract This paper presents a novel optimization framework for estimating the static or dynamic surfaces with details. The proposed method uses dense depths from a structured‐light system or sparse ones from motion capture as the initial positions, and exploits non‐Lambertian reflectance models to approximate surface reflectance. Multi‐stage shape‐from‐shading (SFS) is then applied to optimize both shape geometry and reflectance properties. Because this method uses non‐Lambertian properties, it can compensate for triangulation reconstruction errors caused by view‐dependent reflections. This approach can also estimate detailed undulations on textureless regions, and employs spatial‐temporal constraints for reliably tracking time‐varying surfaces. Experiment results demonstrate that accurate and detailed 3D surfaces can be reconstructed from images acquired by off‐the‐shelf devices. Copyright © 2010 John Wiley & Sons, Ltd. I-Chen Lin, Wen-Hsing Chang, Yung-Sheng Lo, Jen-Yu Peng, Chan-Yu Lin |
Comput. Animat. Virtual Worlds | 1 |
| 2007 | Interactive and flexible motion transitionabstractAbstract In this paper, we present an example‐based motion synthesis technique. Users can interactively control the virtual character to perform desired actions in any order. The desired action can be not only recorded or pre‐computed motion, but also parametric synthesized one to attain the precise control of avatars. Moreover, a user can change their commands any time to switch to another action according to the instant response of opponents in fighting. The quality transition motions between consecutive actions are rapidly synthesized through traversing a simple graph structure which represents the transition relationships between different poses. The graph is constructed according to clustering on frames in a corpus of motion capture data. With the pre‐computation of path finding, our approach can also be applied to real‐time applications. Besides, this pre‐computed graph structure can be used to transit those motions not included in the database. Furthermore, our approach is automatic without any human intervention. The final results demonstrate the potential of our algorithm. Copyright © 2007 John Wiley & Sons, Ltd. Jen-Yu Peng, I-Chen Lin, Jui-Hsiang Chao, Yan-Ju Chen, Gwo-Hao Juang |
Comput. Animat. Virtual Worlds | 2 |
| 2005 | Mirror MoCap: Automatic and efficient capture of dense 3D facial motion parameters from video
I-Chen Lin, Ouhyoung Ming |
Vis. Comput. | 1 |
| 2004 | Surface Detail Capturing for Realistic Facial Animation
Pei-Hsuan Tu, I-Chen Lin, Jeng-Sheng Yeh, Rung-Huei Liang, Ouhyoung Ming |
J. Comput. Sci. Technol. | 2 |
| 2001 | Realistic 3D facial animation parameters from mirror-reflected multi-view videoabstractA robust, accurate and inexpensive approach to estimate 3D facial motion from multi-view video is proposed, where two mirrors located near one's cheeks can reflect the side views of markers on one face. Nice properties of mirrored images are utilized to simplify the proposed tracking algorithm significantly, while a Kalman filter is employed to reduce the noise and to predict the occluded marker positions. More than 50 markers on one face are continuously tracked at 30 frames per second. The estimated 3D facial motion data has been practically applied to our facial animation system. In addition, the dataset of facial motion can also be applied to the analysis of co-articulation effects, facial expressions, and audio-visual hybrid recognition system. I-Chen Lin, Jeng-Sheng Yeh, Ouhyoung Ming |
CA | 1 |
| 1999 | A Speech Driven Talking Head System Based on a Single Face ImageabstractIn this paper, a lifelike talking head system is proposed. The talking head, which is driven by speaker independent speech recognition, requires only one single face image to synthesize lifelike facial expression. The proposed system uses speech recognition engines to get utterances and corresponding time stamps in the speech data. Associated facial expressions can be fetched from an expression pool and the synthetic facial expression can then be synchronized with speech. When applied to Internet, our web-enabled talking head system can be a vivid merchandise narrator, and only requires 50 K bytes/minute with an additional face image (about 40 Kbytes in CIF format, 24 bit-color, JPEG compression). The system can synthesize facial animation more than 30 frames/sec on a Pentium II 266 MHz PC. I-Chen Lin, Cheng-Sheng Hung, Tzong-Jer Yang, Ouhyoung Ming |
PG | 1 |