Jean-Charles Bazin

dblp:83/6046 · DBLP profile ↗
← Back
64ranked-venue papers
17as first author
13since 2021 · last 2026
0000-0001-7660-4802ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 36 · 15 first-author · 7 since 2021Systems, architecture and hardware · 12 · 6 first-authorHuman-computer interaction and ubiquitous computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 HoloQA: Full Reference Video Quality Assessor of Rendered Human Avatars in Virtual Reality
abstract
We present HoloQA, a new state-of-the-art Full Reference Video Quality Assessment (VQA) model that was designed using principles of visual neuroscience, information theory, and self-supervised deep learning to accurately predict the quality of rendered digital human avatars in Virtual Reality (VR) and Augmented Reality (AR) systems. The growing adoption of VR/AR applications that aim to transmit digital human avatars over bandwidth-limited video networks has driven the need for VQA algorithms that better account for the kinds of distortions that reduce the quality of rendered and viewed avatars. As we will show, standard VQA models often fail to capture distortions unique to the rendering, transmission, and compression of videos containing human avatars. Towards solving this difficult problem, we adopt a multi-level Mixture-of-Experts approach. This involves computing distortion-aware perceptual features and high-level content-aware deep features that capture semantic attributes of human body avatars. The high-level features are computed using a self-supervised, pre-trained deep learning network. We show that HoloQA is able to achieve state-of-the-art performance on the recently introduced LIVE-Meta Rendered Human Avatar VQA database, demonstrating its efficacy in predicting the quality of rendered human avatars in VR. Furthermore, we demonstrate the competitive performance of HoloQA on other digital human avatar databases and on another synthetically generated video quality use case: cloud gaming. The code associated with this work will be made available on https://github.com/avinabsaha/HologramQAGitHub.
Avinab Saha, Yu-Chih Chen, Christian Häne, Jean-Charles Bazin, Ioannis Katsavounidis, Alexandre Chapiro, Alan C. Bovik
IEEE Trans. Image Process.4
2024 FAMOUS: High-Fidelity Monocular 3D Human Digitization Using View Synthesis
Vishnu Mani Hema, Shubhra Aich, Christian Häne, Jean-Charles Bazin, Fernando De la Torre
ECCV (82)4
2024 Doubly Hierarchical Geometric Representations for Strand-based Human Hairstyle Generation
abstract
We introduce a doubly hierarchical generative representation for strand-based 3D hairstyle geometry that progresses from coarse, low-pass filtered guide hair to densely populated hair strands rich in high-frequency details. We employ the Discrete Cosine Transform (DCT) to separate low-frequency structural curves from high-frequency curliness and noise, avoiding the Gibbs' oscillation issues associated with the standard Fourier transform in open curves. Unlike the guide hair sampled from the scalp UV map grids which may lose capturing details of the hairstyle in existing methods, our method samples optimal sparse guide strands by utilising $k$-medoids clustering centres from low-pass filtered dense strands, which more accurately retain the hairstyle's inherent characteristics. The proposed variational autoencoder-based generation network, with an architecture inspired by geometric deep learning and implicit neural representations, facilitates flexible, off-the-grid guide strand modelling and enables the completion of dense strands in any quantity and density, drawing on principles from implicit neural representations. Empirical evaluations confirm the capacity of the model to generate convincing guide hair and dense strands, complete with nuanced high-frequency details.
Yunlu Chen, Francisco Vicente 0001, Christian Häne, Giljoo Nam, Jean-Charles Bazin, Fernando De la Torre
NeurIPS5
2024 FaceMap: Distortion-Driven Perceptual Facial Saliency Maps
Zhongshi Jiang, Kishore Venkateshan, Giljoo Nam, Meixu Chen, Romain Bachy, Jean-Charles Bazin, Alexandre Chapiro
SIGGRAPH Asia6
2024 Personalized Face Inpainting with Diffusion Models by Parallel Visual Attention
abstract
Face inpainting is important in various applications, such as photo restoration, image editing, and virtual reality. Despite the significant advances in face generative models, ensuring that a person’s unique facial identity is maintained during the inpainting process is still an elusive goal. Current state-of-the-art techniques, exemplified by MyStyle, necessitate resource-intensive fine-tuning and a substantial number of images for each new identity. Furthermore, existing methods often fall short in accommodating user-specified semantic attributes, such as beard or expression.To improve inpainting results, and reduce the computational complexity during inference, this paper proposes the use of Parallel Visual Attention (PVA) in conjunction with diffusion models. Specifically, we insert parallel attention matrices to each cross-attention module in the denoising network, which attends to features extracted from reference images by an identity encoder. We train the added attention modules and identity encoder on CelebAHQ-IDI, a dataset proposed for identity-preserving face inpainting. Experiments demonstrate that PVA attains unparalleled identity resemblance in both face inpainting and face inpainting with language guidance tasks, in comparison to various benchmarks, including MyStyle, Paint by Example, and Custom Diffusion. Our findings reveal that PVA ensures good identity preservation while offering effective language-controllability. Additionally, in contrast to Custom Diffusion, PVA requires just 40 fine-tuning steps for each new identity, which translates to a significant speed increase of over 20 times.
Jianjin Xu, Saman Motamed, Praneetha Vaddamanu, Chen Henry Wu, Christian Häne, Jean-Charles Bazin, Fernando De la Torre
WACV6
2024 Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual Reality
abstract
We study the visual quality judgments of human subjects on digital human avatars (sometimes referred to as "holograms" in the parlance of virtual reality [VR] and augmented reality [AR] systems) that have been subjected to distortions. We also study the ability of video quality models to predict human judgments. As streaming human avatar videos in VR or AR become increasingly common, the need for more advanced human avatar video compression protocols will be required to address the tradeoffs between faithfully transmitting high-quality visual representations while adjusting to changeable bandwidth scenarios. During transmission over the internet, the perceived quality of compressed human avatar videos can be severely impaired by visual artifacts. To optimize trade-offs between perceptual quality and data volume in practical workflows, video quality assessment (VQA) models are essential tools. However, there are very few VQA algorithms developed specifically to analyze human body avatar videos, due, at least in part, to the dearth of appropriate and comprehensive datasets of adequate size. Towards filling this gap, we introduce the LIVE-Meta Rendered Human Avatar VQA Database, which contains 720 human avatar videos processed using 20 different combinations of encoding parameters, labeled by corresponding human perceptual quality judgments that were collected in six degrees of freedom VR headsets. To demonstrate the usefulness of this new and unique video resource, we use it to study and compare the performances of a variety of state-of-the-art Full Reference and No Reference video quality prediction models, including a new model called HoloQA. As a service to the research community, we publicly releases the metadata of the new database at https://live.ece.utexas.edu/research/LIVE-Meta-rendered-human-avatar/index.html.
Yu-Chih Chen, Avinab Saha, Alexandre Chapiro, Christian Häne, Jean-Charles Bazin, Stefano Zanetti, Ioannis Katsavounidis, Alan C. Bovik
IEEE Trans. Image Process.5
2023 MMG-Ego4D: Multi-Modal Generalization in Egocentric Action Recognition
abstract
In this paper, we study a novel problem in egocentric action recognition, which we term as “Multimodal Generalization“ (MMG). MMG aims to study how systems can generalize when data from certain modalities is limited or even completely missing. We thoroughly investigate MMG in the context of standard supervised action recognition and the more challenging few-shot setting for learning new action categories. MMG consists of two novel scenarios, designed to support security, and efficiency considerations in real-world applications: (1) missing modality generalization where some modalities that were present during the train time are missing during the inference time, and (2) cross-modal zero-shot generalization, where the modalities present during the inference time and the training time are disjoint. To enable this investigation, we construct a new dataset MMG-Ego4D containing data points with video, audio, and inertial motion sensor (IMU) modalities. Our dataset is derived from Ego4D [27] dataset, but processed and thoroughly re-annotated by human experts to facilitate research in the MMG problem. We evaluate a diverse array of models on MMG-Ego4D and propose new methods with improved generalization ability. In particular, we introduce a new fusion module with modality dropout training, contrastive-based alignment training, and a novel cross-modal prototypical loss for better few-shot performance. We hope this study will serve as a benchmark and guide future research in multimodal generalization problems. The benchmark and code are available at https://github.com/facebookresearch/MMG_Ego4D
Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang
CVPR4
2023 PATMAT: Person Aware Tuning of Mask-Aware Transformer for Face inpainting
abstract
Generative models such as StyleGAN2 and Stable Diffusion have achieved state-of-the-art performance in computer vision tasks such as image synthesis, inpainting, and de-noising. However, current generative models for face inpainting often fail to preserve fine facial details and the identity of the person, despite creating aesthetically convincing image structures and textures. In this work, we propose Person Aware Tuning (PAT) of Mask-Aware Transformer (MAT) for face inpainting, which addresses this issue. Our proposed method, PATMAT1, effectively preserves identity by incorporating reference images of a subject and fine-tuning a MAT architecture trained on faces. By using ~40 reference images, PATMAT creates anchor points in MAT’s style module, and tunes the model using the fixed anchors to adapt the model to a new face identity. Moreover, PATMAT’s use of multiple images per anchor during training allows the model to use fewer reference images than competing methods. We demonstrate that PATMAT outperforms state-of-the-art models in terms of image quality, the preservation of person-specific details, and the identity of the subject. Our results suggest that PATMAT can be a promising approach for improving the quality of personalized face inpainting2.
Saman Motamed, Jianjin Xu, Chen Henry Wu, Christian Häne, Jean-Charles Bazin, Fernando De la Torre
ICCV5
2023 A Perceptual Measure for Deep Single Image Camera and Lens Calibration
abstract
Image editing and compositing have become ubiquitous in entertainment, from digital art to AR and VR experiences. To produce beautiful composites, the camera needs to be geometrically calibrated, which can be tedious and requires a physical calibration target. In place of the traditional multi-image calibration process, we propose to infer the camera calibration parameters such as pitch, roll, field of view, and lens distortion directly from a single image using a deep convolutional neural network. We train this network using automatically generated samples from a large-scale panorama dataset, yielding competitive accuracy in terms of standard$\ell ^{2}$error. However, we argue that minimizing such standard error metrics might not be optimal for many applications. In this work, we investigate human sensitivity to inaccuracies in geometric camera calibration. To this end, we conduct a large-scale human perception study where we ask participants to judge the realism of 3D objects composited with correct and biased camera calibration parameters. Based on this study, we develop a new perceptual measure for camera calibration and demonstrate that our deep calibration network outperforms previous single-image based calibration methods both on standard metrics as well as on this novel perceptual measure. Finally, we demonstrate the use of our calibration network for several applications, including virtual object insertion, image retrieval, and compositing.
Yannick Hold-Geoffroy, Dominique Piché-Meunier, Kalyan Sunkavalli, Jean-Charles Bazin, François Rameau, Jean-François Lalonde
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Hong Kong World: Leveraging Structural Regularity for Line-Based SLAM
abstract
Manhattan and Atlanta worlds hold for the structured scenes with only vertical and horizontal dominant directions (DDs). To describe the scenes with additional sloping DDs, a mixture of independent Manhattan worlds seems plausible, but may lead to unaligned and unrelated DDs. By contrast, we propose a novel structural model called Hong Kong world. It is more general than Manhattan and Atlanta worlds since it can represent the environments with slopes, e.g., a city with hilly terrain, a house with sloping roof, and a loft apartment with staircase. Moreover, it is more compact and accurate than a mixture of independent Manhattan worlds by enforcing the orthogonality constraints between not only vertical and horizontal DDs, but also horizontal and sloping DDs. We further leverage the structural regularity of Hong Kong world for the line-based SLAM. Our SLAM method is reliable thanks to three technical novelties. First, we estimate DDs/vanishing points in Hong Kong world in a semi-searching way. We use a new consensus voting strategy for search, instead of traditional branch and bound. This method is the first one that can simultaneously determine the number of DDs, and achieve quasi-global optimality in terms of the number of inliers. Second, we compute the camera pose by exploiting the spatial relations between DDs in Hong Kong world. This method generates concise polynomials, and thus is more accurate and efficient than existing approaches designed for unstructured scenes. Third, we refine the estimated DDs in Hong Kong world by a novel filter-based method. Then we use these refined DDs to optimize the camera poses and 3D lines, leading to higher accuracy and robustness than existing optimization algorithms. In addition, we establish the first dataset of sequential images in Hong Kong world. Experiments showed that our approach outperforms state-of-the-art methods in terms of accuracy and/or efficiency.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Pyojin Kim, Kyungdon Joo, Zhenjun Zhao, Yun-Hui Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Quasi-Globally Optimal and Near/True Real-Time Vanishing Point Estimation in Manhattan World
abstract
Image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim to cluster them by three unknown-but-sought VPs. The VP estimation can be reformulated as computing the rotation between the Manhattan frame and camera frame. To estimate three degrees of freedom (DOF) of this rotation, state-of-the-art methods are based on either data sampling or parameter search. However, they fail to guarantee high accuracy and efficiency simultaneously. In contrast, we propose a set of approaches that hybridize these two strategies. We first constrain two or one DOF of the rotation by two or one sampled image line. Then we search for the remaining one or two DOF based on branch and bound. Our sampling accelerates our search by reducing the search space and simplifying the bound computation. Our search achieves quasi-global optimality. Specifically, it guarantees to retrieve the maximum number of inliers on the condition that two or one DOF is constrained. Our hybridization of two-line sampling and one-DOF search can estimate VPs in real time. Our hybridization of one-line sampling and two-DOF search can estimate VPs in near real time. Experiments on both synthetic and real-world datasets demonstrated that our approaches outperform state-of-the-art methods in terms of accuracy and/or efficiency.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Yun-Hui Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Motionsnap: A Motion Sensor-Based Approach for Automatic Capture and Editing of Photos and Videos on Smartphones
abstract
Taking photos and videos with smartphones has become part of our daily life. However, it is still challenging to capture a brief action at the right time (e.g. jump photos), and video editing (e.g. local slow motion) remains a manual, time-consuming task. To address this problem, we present a motion sensor-based approach that leverages the advanced technical features of modern smartphones to facilitate the capture and editing tasks. Concretely, we simultaneously record the motion sensor data from a smartphone carried by the object of interest, as well as a video from a remote smartphone cam-era. Taking advantage of the motion sensor data, our approach can automatically "snap" the video editing effect to the input video, i.e. apply the effect (e.g. local slow motion) at the right time. The proposed approach is shown to be effective in various applications. Moreover, we implemented it as an app for Android smartphones, running in a fully automatic manner.
Adil Karjauv, Sanzhar Bakhtiyarov, Chaoning Zhang, Jean-Charles Bazin, In-So Kweon
ICME4
2021 ResNet or DenseNet? Introducing Dense Shortcuts to ResNet
abstract
ResNet or DenseNet? Nowadays, most deep learning based approaches are implemented with seminal backbone networks, among them the two arguably most famous ones are ResNet and DenseNet. Despite their competitive performance and overwhelming popularity, inherent drawbacks exist for both of them. For ResNet, the identity shortcut that stabilizes training might limit its representation capacity, and DenseNet mitigates it with multi-layer feature concatenation. However, the dense concatenation causes a new problem of requiring high GPU memory and more training time. Partially due to this, it is not a trivial choice between ResNet and DenseNet. This paper provides a unified perspective of dense summation to analyze them, which facilitates a better understanding of their core difference. We further propose dense weighted normalized shortcuts as a solution to the dilemma between them. Our proposed dense shortcut inherits the design philosophy of simple design in ResNet and DenseNet. On several benchmark datasets, the experimental results show that the proposed DSNet achieves significantly better results than ResNet, and achieves comparable performance as DenseNet but requiring fewer computation resources.
Chaoning Zhang, Philipp Benz, Dawit Mureja Argaw, Seokju Lee, Junsik Kim 0001, François Rameau, Jean-Charles Bazin, In-So Kweon
WACV7
2020 Linear RGB-D SLAM for Atlanta World
abstract
We present a new linear method for RGB-D based simultaneous localization and mapping (SLAM). Compared to existing techniques relying on the Manhattan world assumption defined by three orthogonal directions, our approach is designed for the more general scenario of the Atlanta world. It consists of a vertical direction and a set of horizontal directions orthogonal to the vertical direction and thus can represent a wider range of scenes. Our approach leverages the structural regularity of the Atlanta world to decouple the non-linearity of camera pose estimations. This allows us separately to estimate the camera rotation and then the translation, which bypasses the inherent non-linearity of traditional SLAM techniques. To this end, we introduce a novel tracking-by-detection scheme to estimate the underlying scene structure by Atlanta representation. Thereby, we propose an Atlanta frame-aware linear SLAM framework which jointly estimates the camera motion and a planar map supporting the Atlanta structure through a linear Kalman filter. Evaluations on both synthetic and real datasets demonstrate that our approach provides favorable performance compared to existing state-of-the-art methods while extending their working range to the Atlanta world.
Kyungdon Joo, Tae-Hyun Oh, François Rameau, Jean-Charles Bazin, In-So Kweon
ICRA4
2020 Robust and Efficient Estimation of Absolute Camera Pose for Monocular Visual Odometry
abstract
Given a set of 3D-to-2D point correspondences corrupted by outliers, we aim to robustly estimate the absolute camera pose. Existing methods robust to outliers either fail to guarantee high robustness and efficiency simultaneously, or require an appropriate initial pose and thus lack generality. In contrast, we propose a novel approach based on the robust "L2-minimizing estimate" (L2E) loss. We first define a novel cost function by integrating the projection constraint into the L2E loss. Then to efficiently obtain the global minimum of this function, we propose a hybrid strategy of a local optimizer and branch-and-bound. For branch-and-bound, we derive effective function bounds. Our approach can handle high outlier ratios, leading to high robustness. It can run reliably regardless of whether the initial pose is appropriate, providing high generality. Moreover, given a decent initial pose, it is suitable for real-time applications. Experiments on synthetic and real-world datasets showed that our approach outperforms state-of-the-art methods in terms of robustness and/or efficiency.
Haoang Li, Wen Chen 0021, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001
ICRA4
2020 DeepPTZ: Deep Self-Calibration for PTZ Cameras
abstract
Rotating and zooming cameras, also called PTZ (Pan-Tilt-Zoom) cameras, are widely used in modern surveillance systems. While their zooming ability allows acquiring detailed images of the scene, it also makes their calibration more challenging since any zooming action results in a modification of their intrinsic parameters. Therefore, such camera calibration has to be computed online; this process is called self-calibration. In this paper, given an image pair captured by a PTZ camera, we propose a deep learning based approach to automatically estimate the focal length and distortion parameters of both images as well as the rotation angles between them. The proposed approach relies on a dual-Siamese structure, imposing bidirectional constraints. The proposed network is trained on a large-scale dataset automatically generated from a set of panoramas. Empirically, we demonstrate that our proposed approach achieves competitive performance with respect to both deep learning based and traditional state-of-the art methods. Our code and model will be publicly available at https://github.com/ChaoningZhang/DeepPTZ.
Chaoning Zhang, François Rameau, Junsik Kim 0001, Dawit Mureja Argaw, Jean-Charles Bazin, In-So Kweon
WACV5
2020 Globally Optimal Inlier Set Maximization for Atlanta World Understanding
abstract
In this work, we describe man-made structures via an appropriate structure assumption, called the Atlanta world assumption, which contains a vertical direction (typically the gravity direction) and a set of horizontal directions orthogonal to the vertical direction. Contrary to the commonly used Manhattan world assumption, the horizontal directions in Atlanta world are not necessarily orthogonal to each other. While Atlanta world can encompass a wider range of scenes, this makes the search space much larger and the problem more challenging. Our input data is a set of surface normals, for example, acquired from RGB-D cameras or 3D laser scanners, as well as lines from calibrated images. Given this input data, we propose the first globally optimal method of inlier set maximization for Atlanta direction estimation. We define a novel search space for Atlanta world, as well as its parametrization, and solve this challenging problem using a branch-and-bound (BnB) framework. To alleviate the computational bottleneck in BnB, i.e., the bound computation, we present two bound computation strategies: rectangular bound and slice bound in an efficient measurement domain, i.e., the extended Gaussian image (EGI). In addition, we propose an efficient two-stage method which automatically estimates the number of horizontal directions of a scene. Experimental results with synthetic and real-world datasets have successfully confirmed the validity of our approach.
Kyungdon Joo, Tae-Hyun Oh, In-So Kweon, Jean-Charles Bazin
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Robust Estimation of Absolute Camera Pose via Intersection Constraint and Flow Consensus
abstract
Estimating the absolute camera pose requires 3D-to-2D correspondences of points and/or lines. However, in practice, these correspondences are inevitably corrupted by outliers, which affects the pose estimation. Existing outlier removal strategies for robust pose estimation have some limitations. They are only applicable to points, rely on prior pose information, or fail to handle high outlier ratios. By contrast, we propose a general and accurate outlier removal strategy. It can be integrated with various existing pose estimation methods originally vulnerable to outliers, and is applicable to points, lines, and the combination of both. Moreover, it does not rely on any prior pose information. Our strategy has a nested structure composed of the outer and inner modules. First, our outer module leverages our intersection constraint, i.e., the projection rays or planes defined by inliers intersect at the camera center. Our outer module alternately computes the inlier probabilities of correspondences and estimates the camera pose. It can run reliably and efficiently under high outlier ratios. Second, our inner module exploits our flow consensus. The 2D displacement vectors or 3D directed arcs generated by inliers exhibit a common directional regularity, i.e., follow a dominant trend of flow. Our inner module refines the inlier probabilities obtained at each iteration of our outer module. This refinement improves the accuracy and facilitates the convergence of our outer module. Experiments on both synthetic data and real-world images have shown that our method outperforms state-of-the-art approaches in terms of accuracy and robustness.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Yun-Hui Liu 0001
IEEE Trans. Image Process.3
2020 Capturing Subjective First-Person View Shots with Drones for Automated Cinematography
abstract
We propose an approach to capture subjective first-person view (FPV) videos by drones for automated cinematography. FPV shots are intentionally not smooth to increase the level of immersion for the audience, and are usually captured by a walking camera operator holding traditional camera equipment. Our goal is to automatically control a drone in such a way that it imitates the motion dynamics of a walking camera operator, and, in turn, capture FPV videos. For this, given a user-defined camera path, orientation, and velocity, we first present a method to automatically generate the operator’s motion pattern and the associated motion of the camera, considering the damping mechanism of the camera equipment. Second, we propose a general computational approach that generates the drone commands to imitate the desired motion pattern. We express this task as a constrained optimization problem, where we aim to fulfill high-level user-defined goals, while imitating the dynamics of the walking camera operator and taking the drone’s physical constraints into account. Our approach is fully automatic, runs in real time, and is interactive, which provides artistic freedom in designing shots. It does not require a motion capture system, and works both indoors and outdoors. The validity of our approach has been confirmed via quantitative and qualitative evaluations.
Amirsaman Ashtari, Stefan Stevsic, Tobias Naegeli, Jean-Charles Bazin, Otmar Hilliges
ACM Trans. Graph.4
2019 Revisiting Residual Networks with Nonlinear Shortcuts
Chaoning Zhang, François Rameau, Seokju Lee, Junsik Kim 0001, Philipp Benz, Dawit Mureja Argaw, Jean-Charles Bazin, In-So Kweon
BMVC7
2019 Quasi-Globally Optimal and Efficient Vanishing Point Estimation in Manhattan World
abstract
The image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim at clustering them by three unknown-but-sought VPs. The VP estimation can be reformulated as computing the rotation between the Manhattan frame and the camera frame. To compute this rotation, state-of-the-art methods are based on either data sampling or parameter search, and they fail to guarantee the accuracy and efficiency simultaneously. In contrast, we propose to hybridize these two strategies. We first compute two degrees of freedom (DOF) of the above rotation by two sampled image lines, and then search for the optimal third DOF based on the branch-and-bound. Our sampling accelerates our search by reducing the search space and simplifying the bound computation. Our search is not sensitive to noise and achieves quasi-global optimality in terms of maximizing the number of inliers. Experiments on synthetic and real-world images showed that our method outperforms state-of-the-art approaches in terms of accuracy and/or efficiency.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Wen Chen 0021, Zhe Liu 0022, Yun-Hui Liu 0001
ICCV3
2019 Leveraging Structural Regularity of Atlanta World for Monocular SLAM
abstract
A wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocular SLAM. First, we robustly cluster image lines. Based on these clusters, we compute the local Atlanta frames in the camera frame by solving polynomial equations. Our method provides the global optimum and satisfies inherent geometric constraints. Second, we define the posterior probabilities to refine the initial clusters and Atlanta frames alternately by the maximum a posteriori estimation. Third, based on multiple local Atlanta frames, we compute the global Atlanta frames in the world frame using Kalman filtering. We optimize rotations by the global alignment and then refine translations and 3D line-based map under the directional constraints. Experiments on both synthesized and real data have demonstrated that our approach outperforms state-of-the-art methods.
Haoang Li, Yazhou Xing, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001
ICRA4
2019 Line-based Absolute and Relative Camera Pose Estimation in Structured Environments
abstract
3D lines in structured environments encode particular regularity like parallelism and orthogonality. We leverage this structural regularity to estimate the absolute and relative camera poses. We decouple the rotation and translation, and propose a novel rotation estimation method. We decompose the absolute and relative rotations and reformulate the problem as computing the rotation from the Manhattan frame to the camera frame. To compute this rotation, we propose an accurate and efficient two-step method. We first estimate its two degrees of freedom (DOF) by two image lines, and then estimate its third DOF by another image line. For these lines, we assume their associated 3D lines are mutually orthogonal, or two 3D lines are parallel to each other and orthogonal to the third. Thanks to our two-step DOF estimation, our absolute and relative pose estimation methods are accurate and efficient. Moreover, our relative pose estimation method relies on weaker assumptions or less correspondences than existing approaches. We also propose a novel strategy to reject outliers and identify dominant directions of the scene. We integrate it into our pose estimation methods, and show that it is more robust than RANSAC. Experiments on synthetic and real-world datasets demonstrated that our methods outperform state-of-the-art approaches.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Wen Chen 0021, Kai Chen 0028, Yun-Hui Liu 0001
IROS3
2019 Deep360Up: A Deep Learning-Based Approach for Automatic VR Image Upright Adjustment
abstract
Spherical VR cameras can capture high-quality immersive VR images with a 360° field of view. However, in practice, when the camera orientation is not straight, the acquired VR image appears tilted when displayed on a VR headset, which diminishes the quality of the VR experience. To overcome this problem, we present a deep learning-based approach that can automatically estimate the orientation of a VR image and return its upright version. In contrast to existing methods, our approach does not require the presence of lines or horizon in the image, and thus can be applied on a wide range of scenes. Extensive experiments and comparisons with state-of-the-art methods have successfully confirmed the validity of our approach.
Raehyuk Jung, Aiden Seung Joon Lee, Amirsaman Ashtari, Jean-Charles Bazin
VR4
2019 A computational approach for spider web-inspired fabrication of string art
abstract
Abstract Creating objects with threads is commonly referred to as string art. It is typically a manual, tedious work reserved for skilled artists. In this paper, we investigate how to automatically fabricate string art pieces from one single continuous thread in such a way that it looks like an input image. The proposed system consists of a thread connection optimization algorithm and a custom‐made fabrication machine. It allows casual users to create their own personalized string art pieces in a fully automatic manner. Quantitative and qualitative evaluations demonstrated our system can create visually appealing results.
Seungwoo Je, Yekaterina Abileva, Andrea Bianchi, Jean-Charles Bazin
Comput. Animat. Virtual Worlds4
2018 Globally Optimal Inlier Set Maximization for Atlanta Frame Estimation
abstract
In this work, we describe man-made structures via an appropriate structure assumption, called Atlanta world, which contains a vertical direction (typically the gravity direction) and a set of horizontal directions orthogonal to the vertical direction. Contrary to the commonly used Manhattan world assumption, the horizontal directions in Atlanta world are not necessarily orthogonal to each other. While Atlanta world permits to encompass a wider range of scenes, this makes the solution space larger and the problem more challenging. Given a set of inputs, such as lines in a calibrated image or surface normals, we propose the first globally optimal method of inlier set maximization for Atlanta direction estimation. We define a novel search space for Atlanta world, as well as its parameterization, and solve this challenging problem by a branch-and-bound framework. Experimental results with synthetic and real-world datasets have successfully confirmed the validity of our approach.
Kyungdon Joo, Tae-Hyun Oh, In-So Kweon, Jean-Charles Bazin
CVPR4
2018 A Monocular SLAM System Leveraging Structural Regularity in Manhattan World
abstract
The structural features in Manhattan world encode useful geometric information of parallelism, orthogonality and/or coplanarity in the scene. By fully exploiting these structural features, we propose a novel monocular SLAM system which provides accurate estimation of camera poses and 3D map. The foremost contribution of the proposed system is a structural feature-based optimization module which contains three novel optimization strategies. First, a rotation optimization strategy using the parallelism and orthogonality of 3D lines is presented. We propose a global binding method to compute an accurate estimation of the absolute rotation of the camera. Then we propose an approach for calculating the relative rotation to further refine the absolute rotation. Second, a translation optimization strategy leveraging coplanarity is proposed. Coplanar features are effectively identified, and we leverage them by a unified model handling both points and lines to calculate the relative translation, and then the optimal absolute translation. Third, a 3D line optimization strategy utilizing parallelism, orthogonality and coplanarity simultaneously is proposed to obtain an accurate 3D map consisting of structural line segments with low computational complexity. Experiments in man-made environments have demonstrated that the proposed system outperforms existing state-of-the-art monocular SLAM systems in terms of accuracy and robustness.
Haoang Li, Jian Yao 0002, Jean-Charles Bazin, Xiaohu Lu, Yazhou Xing, Kang Liu 0003
ICRA3
2018 Robust Camera Pose Estimation via Consensus on Ray Bundle and Vector Field
abstract
Estimating the camera pose requires point correspondences. However, in practice, correspondences are inevitably corrupted by outliers, which affects the pose estimation. We propose a general and accurate outlier removal strategy for robust camera pose estimation. The proposed strategy can detect outliers by leveraging the fact that only inliers comply with two effective consensuses, i.e., 3D ray bundle consensus and 2D vector field consensus. Our strategy has a nested structure. First, the outer module utilizes the 3D ray bundle consensus. We define the likelihood based on the probabilistic mixture model and maximize it by the expectation-maximization (EM) algorithm. The inlier probability of each correspondence and the camera pose are determined alternately. Second, the inner module exploits the 2D vector field consensus to refine the probabilities obtained by the outer module. The refinement based on the Bayesian rule facilitates the convergence of the outer module and improves the accuracy of the entire framework. Our strategy can be integrated into various existing camera pose estimation methods which are originally vulnerable to outliers. Experiments on both synthesized data and real images have shown that our approach outperforms state-of-the-art outlier rejection methods in terms of accuracy and robustness.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Jian Yao 0002
IROS3
2018 Overlay Design Methodology for virtual environment design within digital games
Ikhwan Kim, SukJoo Hong, Ji-Hyun Lee, Jean-Charles Bazin
Adv. Eng. Informatics4
2018 An Omnistereoscopic Video Pipeline for Capture and Display of Real-World VR
abstract
In this article, we describe a complete pipeline for the capture and display of real-world Virtual Reality video content, based on the concept of omnistereoscopic panoramas. We address important practical and theoretical issues that have remained undiscussed in previous works. On the capture side, we show how high-quality omnistereo video can be generated from a sparse set of cameras (16 in our prototype array) instead of the hundreds of input views previously required. Despite the sparse number of input views, our approach allows for high quality, real-time virtual head motion, thereby providing an important additional cue for immersive depth perception compared to static stereoscopic video. We also provide an in-depth analysis of the required camera array geometry in order to meet specific stereoscopic output constraints, which is fundamental for achieving a plausible and fully controlled VR viewing experience. Finally, we describe additional insights on how to integrate omnistereo video panoramas with rendered CG content. We provide qualitative comparisons to alternative solutions, including depth-based view synthesis and the Facebook Surround 360 system. In summary, this article provides a first complete guide and analysis for reimplementing a system for capturing and displaying real-world VR, which we demonstrate on several real-world examples captured with our prototype.
Christopher Schroers, Jean-Charles Bazin, Alexander Sorkine-Hornung
ACM Trans. Graph.2
2017 Rapid one-shot acquisition of dynamic VR avatars
abstract
We present a system for rapid acquisition of bespoke, animatable, full-body avatars including face texture and shape. A blendshape rig with a skeleton is used as a template for customization. Identity blendshapes are used to customize the body and face shape at the fitting stage, while animation blendshapes allow the face to be animated. The subject assumes a T-pose and a single snapshot is captured using a stereo RGB plus depth sensor rig. Our system automatically aligns a photo texture and fits the 3D shape of the face. The body shape is stylized according to body dimensions estimated from segmented depth. The face identity blendweights are optimised according to image-based facial landmarks, while a custom texture map for the face is generated by warping the input images to a reference texture according to the facial landmarks. The total capture and processing time is under 10 seconds and the output is a light-weight, game-engine-ready avatar which is recognizable as the subject. We demonstrate our system in a VR environment in which each user sees the other users' animated avatars through a VR headset with real-time audio-based facial animation and live body motion tracking, affording an enhanced level of presence and social engagement compared to generic avatars.
Charles Malleson, Maggie Kosek, Martin Klaudiny, Ivan Huerta Casado, Jean-Charles Bazin, Alexander Sorkine-Hornung, Mark Mine, Kenny Mitchell
VR5
2017 Demonstration: Rapid one-shot acquisition of dynamic VR avatars
abstract
In this demonstration, we showcase a system for rapid acquisition of bespoke avatars for each participant (subject) in a social VR environment is presented. For each subject, the system automatically customizes a parametric avatar model to match the captured subject by adjusting its overall height, body and face shape parameters and generating a custom face texture.
Charles Malleson, Maggie Kosek, Martin Klaudiny, Ivan Huerta Casado, Jean-Charles Bazin, Alexander Sorkine-Hornung, Mark Mine, Kenny Mitchell
VR5
2016 An Immersive Bidirectional System for Life-size 3D Communication
abstract
Telecommunication and video conferencing are an integral part of modern society with implications in many aspects of everyday life. However, compared to a meeting in person, the sense of presence is still limited in electronic communication. In this paper, we present a novel system for life-size 3D telecommunication. It is designed to create an immersive user experience by seamlessly embedding a remote conversation partner into the local environment. To achieve this, users are captured in 3D by hybrid (color+depth) sensors and displayed on a life-size transparent 3D display. We have built two instances of this system in Zurich and Singapore. They form a complete and fully functional prototype enabling bidirectional communication in real-time over a long distance. We further demonstrate alternative hardware setups, which make our system flexible and adaptable to different usage scenarios.
Claudia Plüss, Nicola Ranieri, Jean-Charles Bazin, Pierre-Yves Laffont, Tiberiu Popa, Markus Gross 0001
CASA3
2016 ActionSnapping: Motion-Based Video Synchronization
Jean-Charles Bazin, Alexander Sorkine-Hornung
ECCV (5)1
2016 Motion based remote camera control with mobile devices
abstract
With current digital cameras and smartphones, taking photos and videos has never been easier. However, it is still difficult to take a photo of a brief action at the right time. In addition, editing captured videos, such as modifying the playback speed of some parts of a video, remains a time consuming task.
Sabir Akhadov, Marcel Lancelle, Jean-Charles Bazin, Markus Gross 0001
MobileHCI3
2016 Physically Based Video Editing
abstract
Abstract Convincing manipulation of objects in live action videos is a difficult and often tedious task. Skilled video editors achieve this with the help of modern professional tools, but complex motions might still lack physical realism since existing tools do not consider the laws of physics. On the other hand, physically based simulation promises a high degree of realism, but typically creates a virtual 3D scene animation rather than returning an edited version of an input live action video. We propose a framework that combines video editing and physics‐based simulation. Our tool assists unskilled users in editing an input image or video while respecting the laws of physics and also leveraging the image content. We first fit a physically based simulation that approximates the object's motion in the input video. We then allow the user to edit the physical parameters of the object, generating a new physical behavior for it. The core of our work is the formulation of an image‐aware constraint within physics simulations. This constraint manifests as external control forces to guide the object in a way that encourages proper texturing at every frame, yet producing physically plausible motions. We demonstrate the generality of our method on a variety of physical interactions: rigid motion, multi‐body collisions, clothes and elastic bodies.
Jean-Charles Bazin, Claudia Plüss, Alec Jacobson, Markus Gross 0001
Comput. Graph. Forum1
2016 Partial Sum Minimization of Singular Values in Robust PCA: Algorithm and Applications
abstract
Robust Principal Component Analysis (RPCA) via rank minimization is a powerful tool for recovering underlying low-rank structure of clean data corrupted with sparse noise/outliers. In many low-level vision problems, not only it is known that the underlying structure of clean data is low-rank, but the exact rank of clean data is also known. Yet, when applying conventional rank minimization for those problems, the objective function is formulated in a way that does not fully utilize a priori target rank information about the problems. This observation motivates us to investigate whether there is a better alternative solution when using rank minimization. In this paper, instead of minimizing the nuclear norm, we propose to minimize the partial sum of singular values, which implicitly encourages the target rank constraint. Our experimental analyses show that, when the number of samples is deficient, our approach leads to a higher success rate than conventional rank minimization, while the solutions obtained by the two approaches are almost identical when the number of samples is more than sufficient. We apply our approach to various low-level vision problems, e.g., high dynamic range imaging, motion edge detection, photometric stereo, image alignment and recovery, and show that our results outperform those obtained by the conventional nuclear norm rank minimization method.
Tae-Hyun Oh, Yu-Wing Tai, Jean-Charles Bazin, Hyeongwoo Kim, In-So Kweon
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Intrinsic Decomposition of Image Sequences from Local Temporal Variations
abstract
We present a method for intrinsic image decomposition, which aims to decompose images into reflectance and shading layers. Our input is a sequence of images with varying illumination acquired by a static camera, e.g. an indoor scene with a moving light source or an outdoor timelapse. We leverage the local color variations observed over time to infer constraints on the reflectance and solve the ill-posed image decomposition problem. In particular, we derive an adaptive local energy from the observations of each local neighborhood over time, and integrate distant pairwise constraints to enforce coherent decomposition across all surfaces with consistent shading changes. Our method is solely based on multiple observations of a Lambertian scene under varying illumination and does not require user interaction, scene geometry, or an explicit lighting model. We compare our results with several intrinsic decomposition methods on a number of synthetic and captured datasets.
Pierre-Yves Laffont, Jean-Charles Bazin
ICCV2
2015 FaceDirector: Continuous Control of Facial Performance in Video
abstract
We present a method to continuously blend between multiple facial performances of an actor, which can contain different facial expressions or emotional states. As an example, given sad and angry video takes of a scene, our method empowers the movie director to specify arbitrary weighted combinations and smooth transitions between the two takes in post-production. Our contributions include (1) a robust nonlinear audio-visual synchronization technique that exploits complementary properties of audio and visual cues to automatically determine robust, dense spatiotemporal correspondences between takes, and (2) a seamless facial blending approach that provides the director full control to interpolate timing, facial expression, and local appearance, in order to generate novel performances after filming. In contrast to most previous works, our approach operates entirely in image space, avoiding the need of 3D facial reconstruction. We demonstrate that our method can synthesize visually believable performances with applications in emotion transition, performance correction, and timing control.
Charles Malleson, Jean-Charles Bazin, Oliver Wang, Derek Bradley, Thabo Beeler, Adrian Hilton 0001, Alexander Sorkine-Hornung
ICCV2
2015 Sampling based scene-space video processing
abstract
Many compelling video processing effects can be achieved if per-pixel depth information and 3D camera calibrations are known. However, the success of such methods is highly dependent on the accuracy of this "scene-space" information. We present a novel, sampling-based framework for processing video that enables high-quality scene-space video effects in the presence of inevitable errors in depth and camera pose estimation. Instead of trying to improve the explicit 3D scene representation, the key idea of our method is to exploit the high redundancy of approximate scene information that arises due to most scene points being visible multiple times across many frames of video. Based on this observation, we propose a novel pixel gathering and filtering approach. The gathering step is general and collects pixel samples in scene-space, while the filtering step is application-specific and computes a desired output video from the gathered sample sets. Our approach is easily parallelizable and has been implemented on GPU, allowing us to take full advantage of large volumes of video data and facilitating practical runtimes on HD video using a standard desktop computer. Our generic scene-space formulation is able to comprehensively describe a multitude of video processing applications such as denoising, deblurring, super resolution, object removal, computational shutter functions, and other scene-space camera effects. We present results for various casually captured, hand-held, moving, compressed, monocular videos depicting challenging scenes recorded in uncontrolled environments.
Felix Klose, Oliver Wang, Jean-Charles Bazin, Marcus A. Magnor, Alexander Sorkine-Hornung
ACM Trans. Graph.3
2014 Globally Optimal Inlier Set Maximization with Unknown Rotation and Focal Length
Jean-Charles Bazin, Yongduek Seo, Richard I. Hartley, Marc Pollefeys
ECCV (2)1
2014 Automatic jumping photos on smartphones
abstract
Jumping photos are very popular, particularly in the contexts of holidays and entertainment. However, triggering the camera at the right time to take a visually appealing jumping photo is quite difficult in practice, especially for casual photographers or self-portraits. We propose a fully automatic method that solves this practical problem. By analyzing the ongoing jump motion online at a fast rate, our method predicts the time at which the jumping person will reach the highest point and takes trigger delays into account to compute when the camera has to be triggered. Since smartphones are more and more ubiquitous, we focus on these devices which leads to some challenges such as limited computational power and data transfer rates. We developed an Android app for smart-phones and used it to conduct experiments with various jump styles confirming the validity of our approach.
Cecilia Garcia, Jean-Charles Bazin, Marcel Lancelle, Markus Gross 0001
ICIP2
2014 Registration of multiple RGBD cameras via local rigid transformations
abstract
RGBD cameras, such as the Kinect, have recently revolutionized the field of real-time geometry and appearance acquisition. While impressive 3D reconstruction results have been obtained, combining data acquired by multiple RGBD cameras constitutes a technical challenge. Several methods have been proposed to estimate the internal parameters of each RGBD camera (such as depth mapping function and focal length). Despite that the textured geometry obtained by each RGBD camera individually is visually attractive, even state-of-the-art methods have difficulties in correctly combining the textured geometries obtained by several RGBD cameras via a rigid transformation. Based on this observation, our approach registers the RGBD cameras by a smooth field of rigid transformations, instead of a single rigid transformation. Experimental results on challenging data demonstrate the validity of the proposed approach.
Teng Deng, Jean-Charles Bazin, Claudia Plüss, Jianfei Cai 0001, Tiberiu Popa, Markus Gross 0001
ICME2
2014 Gaze correction witha single webcam
abstract
Eye contact is a critical aspect of human communication. However, when talking over a video conferencing system, such as Skype, it is not possible for users to have eye contact when looking at the conversation partner's face displayed on the screen. This is due to the location disparity between the video conferencing window and the camera. This issue has been tackled by expensive high-end systems or hybrid depth+color cameras, but such equipment is still largely unavailable at the consumer level and on platforms such as laptops or tablets. In contrast, we propose a gaze correction method that needs just a single webcam. We apply recent shape deformation techniques to generate a 3D face model that matches the user's face. We then render a gaze-corrected version of this face model and seamlessly insert it into the original image. Experiments on real data and various platforms confirm the validity of the approach and demonstrate that the visual quality of our results is at least equivalent to those obtained by state-of-the-art methods requiring additional equipment.
Dominik Giger, Jean-Charles Bazin, Claudia Plüss, Tiberiu Popa, Markus Gross 0001
ICME2
2014 Spatio-temporal geometry fusion for multiple hybrid cameras using moving least squares surfaces
abstract
Abstract Multi‐view reconstruction aims at computing the geometry of a scene observed by a set of cameras. Accurate 3D reconstruction of dynamic scenes is a key component for a large variety of applications, ranging from special effects to telepresence and medical imaging. In this paper we propose a method based on Moving Least Squares surfaces which robustly and efficiently reconstructs dynamic scenes captured by a calibrated set of hybrid color+depth cameras. Our reconstruction provides spatio‐temporal consistency and seamlessly fuses color and geometric information. We illustrate our approach on a variety of real sequences and demonstrate that it favorably compares to state‐of‐the‐art methods.
Claudia Plüss, Jean-Charles Bazin, A. Cengiz Öztireli, Teng Deng, Tiberiu Popa, Markus Gross 0001
Comput. Graph. Forum2
2013 Partial Sum Minimization of Singular Values in RPCA for Low-Level Vision
abstract
Robust Principal Component Analysis (RPCA) via rank minimization is a powerful tool for recovering underlying low-rank structure of clean data corrupted with sparse noise/outliers. In many low-level vision problems, not only it is known that the underlying structure of clean data is low-rank, but the exact rank of clean data is also known. Yet, when applying conventional rank minimization for those problems, the objective function is formulated in a way that does not fully utilize a priori target rank information about the problems. This observation motivates us to investigate whether there is a better alternative solution when using rank minimization. In this paper, instead of minimizing the nuclear norm, we propose to minimize the partial sum of singular values. The proposed objective function implicitly encourages the target rank constraint in rank minimization. Our experimental analyses show that our approach performs better than conventional rank minimization when the number of samples is deficient, while the solutions obtained by the two approaches are almost identical when the number of samples is more than sufficient. We apply our approach to various low-level vision problems, e.g. high dynamic range imaging, photometric stereo and image alignment, and show that our results outperform those obtained by the conventional nuclear norm rank minimization method.
Tae-Hyun Oh, Hyeongwoo Kim, Yu-Wing Tai, Jean-Charles Bazin, In-So Kweon
ICCV4
2013 Scalable Music: Automatic Music Retargeting and Synthesis
abstract
Abstract In this paper we propose a method for dynamic rescaling of music, inspired by recent works on image retargeting, video reshuffling and character animation in the computer graphics community. Given the desired target length of a piece of music and optional additional constraints such as position and importance of certain parts, we build on concepts from seam carving, video textures and motion graphs and extend them to allow for a global optimization of jumps in an audio signal. Based on an automatic feature extraction and spectral clustering for segmentation, we employ length‐constrained least‐costly path search via dynamic programming to synthesize a novel piece of music that best fulfills all desired constraints, with imperceptible transitions between reshuffled parts. We show various applications of music retargeting such as part removal, decreasing or increasing music duration, and in particular consistent joint video and audio editing.
Simon Wenner, Jean-Charles Bazin, Alexander Sorkine-Hornung, Changil Kim 0001, Markus Gross 0001
Comput. Graph. Forum2
2013 A Branch-and-Bound Approach to Correspondence and Grouping Problems
abstract
Data correspondence/grouping under an unknown parametric model is a fundamental topic in computer vision. Finding feature correspondences between two images is probably the most popular application of this research field, and is the main motivation of our work. It is a key ingredient for a wide range of vision tasks, including three-dimensional reconstruction and object recognition. Existing feature correspondence methods are based on either local appearance similarity or global geometric consistency or a combination of both in some heuristic manner. None of these methods is fully satisfactory, especially in the presence of repetitive image textures or mismatches. In this paper, we present a new algorithm that combines the benefits of both appearance-based and geometry-based methods and mathematically guarantees a global optimization. Our algorithm accepts the two sets of features extracted from two images as input, and outputs the feature correspondences with the largest number of inliers, which verify both the appearance similarity and geometric constraints. Specifically, we formulate the problem as a mixed integer program and solve it efficiently by a series of linear programs via a branch-and-bound procedure. We subsequently generalize our framework in the context of data correspondence/grouping under an unknown parametric model and show it can be applied to certain classes of computer vision problems. Our algorithm has been validated successfully on synthesized data and challenging real images.
Jean-Charles Bazin, Hongdong Li, In-So Kweon, Cédric Demonceaux, Pascal Vasseur, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Painting by feature: texture boundaries for example-based image creation
abstract
In this paper we propose a reinterpretation of the brush and the fill tools for digital image painting. The core idea is to provide an intuitive approach that allows users to paint in the visual style of arbitrary example images. Rather than a static library of colors, brushes, or fill patterns, we offer users entire images as their palette, from which they can select arbitrary contours or textures as their brush or fill tool in their own creations. Compared to previous example-based techniques related to the painting-by-numbers paradigm we propose a new strategy where users can generate salient texture boundaries by our randomized graph-traversal algorithm and apply a content-aware fill to transfer textures into the delimited regions. This workflow allows users of our system to intuitively create visually appealing images that better preserve the visual richness and fluidity of arbitrary example images. We demonstrate the potential of our approach in various applications including interactive image creation, editing and vector image stylization.
Michal Lukác, Jakub Fiser, Jean-Charles Bazin, Ondrej Jamriska, Alexander Sorkine-Hornung, Daniel Sýkora
ACM Trans. Graph.3
2012 Globally Optimal Consensus Set Maximization through Rotation Search
Jean-Charles Bazin, Yongduek Seo, Marc Pollefeys
ACCV (2)1
2012 Globally optimal line clustering and vanishing point estimation in Manhattan world
abstract
The projections of world parallel lines in an image intersect at a single point called the vanishing point (VP). VPs are a key ingredient for various vision tasks including rotation estimation and 3D reconstruction. Urban environments generally exhibit some dominant orthogonal VPs. Given a set of lines extracted from a calibrated image, this paper aims to (1) determine the line clustering, i.e. find which line belongs to which VP, and (2) estimate the associated orthogonal VPs. None of the existing methods is fully satisfactory because of the inherent difficulties of the problem, such as the local minima and the chicken-and-egg aspect. In this paper, we present a new algorithm that solves the problem in a mathematically guaranteed globally optimal manner and can inherently enforce the VP orthogonality. Specifically, we formulate the task as a consensus set maximization problem over the rotation search space, and further solve it efficiently by a branch-and-bound procedure based on the Interval Analysis theory. Our algorithm has been validated successfully on sets of challenging real images as well as synthetic data sets.
Jean-Charles Bazin, Yongduek Seo, Cédric Demonceaux, Pascal Vasseur, Katsushi Ikeuchi, In-So Kweon, Marc Pollefeys
CVPR1
2012 3-line RANSAC for orthogonal vanishing point detection
abstract
A wide range of robotic systems needs to estimate their rotation for diverse tasks like automatic control and stabilization, among many others. In regards of the limitations of traditional navigation equipments (like GPS and inertial sensors), this paper follows a vision approach based on the observation of vanishing points (VPs). Urban environments (outdoor as well as indoor) generally contain orthogonal VPs which constitutes an important constraint to fulfill in order to correctly acquire the structure of the scenes. In contrast to existing VP-based techniques, our method inherently enforces the orthogonality of the VPs by directly incorporating the orthogonality constraint into the model estimation step of the RANSAC procedure, which allows real-time applications. The model is estimated from only 3 lines, which corresponds to the theoretical minimal sampling for rotation estimation and constitutes our 3-line RANSAC. We also propose a 1-line RANSAC when the horizon plane is known. Our algorithm has been validated successfully on challenging real datasets.
Jean-Charles Bazin, Marc Pollefeys
IROS1
2012 Gaze correction for home video conferencing
abstract
Effective communication using current video conferencing systems is severely hindered by the lack of eye contact caused by the disparity between the locations of the subject and the camera. While this problem has been partially solved for high-end expensive video conferencing systems, it has not been convincingly solved for consumer-level setups. We present a gaze correction approach based on a single Kinect sensor that preserves both the integrity and expressiveness of the face as well as the fidelity of the scene as a whole, producing nearly artifact-free imagery. Our method is suitable for mainstream home video conferencing: it uses inexpensive consumer hardware, achieves real-time performance and requires just a simple and short setup. Our approach is based on the observation that for our application it is sufficient to synthesize only the corrected face. Thus we render a gaze-corrected 3D model of the scene and, with the aid of a face tracker, transfer the gaze-corrected facial portion in a seamless manner onto the original image.
Claudia Plüss, Tiberiu Popa, Jean-Charles Bazin, Craig Gotsman, Markus Gross 0001
ACM Trans. Graph.3
2011 2D/3D virtual face modeling
abstract
We propose a novel and simple framework that solves two popular problems in digital photography: 2D face synthesis and 3D face modeling. 2D face synthesis aims at creating a new face, usually by mixing two or more portraits. We extend this notion to the combination of human and statue faces. The goal of 3D face modeling is to reconstruct a face in three dimensions from one or several images. These two tasks are often treated as separate problems although they both consider face modeling. In this paper, we propose a unified and general framework for both 2D and 3D cases that runs in a fully automatic manner. Our work also creates stereoscopic views for entertainment 3D display. Experimental results and subjective tests have confirmed the validity of our approach.
Soonkee Chung, Jean-Charles Bazin, In-So Kweon
ICIP2
2010 An original approach for automatic plane extraction by omnidirectional vision
abstract
Whereas some methods for plane extraction have been proposed, this problem still remains an open issue due to the complexity of the task. This paper especially focuses on the extraction of points lying on a plane (such as the ground and buildings walls) in sequences acquired by a central omnidirectional camera. Our approach is based on the epipolar constraint for planar scenes (i.e. homography) on a pair of omnidirectional images to detect some interest points belonging to a plane. Our main contribution is the introduction of a new method, called “2-point algorithm for homography”, that imposes some constraints on the homography using vanishing point (VP) information. Compared to the widely used DLT (4-point) algorithm, experiments on real data demonstrated that the proposed “2-point algorithm for homography” is more robust to noise and false matching, even when the plane to extract is not dominant in the image. Finally, we show that our system provides key clues for ground segmentation by GrabCut.
Jean-Charles Bazin, Pierre-Yves Laffont, In-So Kweon, Cédric Demonceaux, Pascal Vasseur
IROS1
2010 Motion estimation by decoupling rotation and translation in catadioptric vision
Jean-Charles Bazin, Cédric Demonceaux, Pascal Vasseur, In-So Kweon
Comput. Vis. Image Underst.1
2009 Particle Filter Approach Adapted to Catadioptric Images for Target Tracking Application
abstract
International audience
Jean-Charles Bazin, Kuk-Jin Yoon, In-So Kweon, Cédric Demonceaux, Pascal Vasseur
BMVC1
2009 Automatic closed eye correction
abstract
On a large group picture, having all people open their eyes can turn out to be a difficult task for photographers. Therefore, in this paper, we describe an original method to automatically correct closed eyes on everyday pictures. For this aim, we explore the combination possibilities of (1) active shape model (ASM) to detect facial features, such as eyes, nose and head shape, and (2) Poisson editing to clone open eyes seamlessly. To improve the performance of seamless cloning, we suggest a pre-processing method that adjusts skin luminosity between two pictures. A nearest neighbor-based search to find the best suited pair of eyes among a set of donor candidates is also presented. We applied the proposed algorithm on several pictures and obtained very natural results, which demonstrates the validity of our approach.
Jean-Charles Bazin, Dang-Quang Pham, In-So Kweon, Kuk-Jin Yoon
ICIP1
2009 Dynamic programming and skyline extraction in catadioptric infrared images
abstract
Unmanned Aerial Vehicles (UAV) are the subject of an increasing interest in many applications and a key requirement for autonomous navigation is the attitude/position stabilization of the vehicle. Some previous works have suggested using catadioptric vision, instead of traditional perspective cameras, in order to gather much more information from the environment and therefore improve the robustness of the UAV attitude/position estimation. This paper belongs to a series of recent publications of our research group concerning catadioptric vision for UAVs. Currently, we focus on the extraction of skyline in catadioptric images since it provides important information about the attitude/position of the UAV. For example, the DEM-based methods can match the extracted skyline with a Digital Elevation Map (DEM) by process of registration, which permits to estimate the attitude and the position of the camera. Like any standard cameras, catadioptric systems cannot work in low luminosity situations because they are based on visible light. To overcome this important limitation, in this paper, we propose using a catadioptric infrared camera and extending one of our methods of skyline detection towards catadioptric infrared images. The task of extracting the best skyline in images is usually converted in an energy minimization problem that can be solved by dynamic programming. The major contribution of this paper is the extension of dynamic programming for catadioptric images using an adapted neighborhood and an appropriate scanning direction. Finally, we present some experimental results to demonstrate the validity of our approach.
Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur
ICRA1
2008 Improvement of feature matching in catadioptric images using gyroscope data
abstract
Most of vision-based algorithms for motion and localization estimation requires matching some interest points in a pair of images. After building feature correspondence, it is possible to estimate camera motion/localization using epipolar geometry. However feature matching is still a challenging problem because of time constraint or image variability for example. In several robotic applications, the camera rotation may be known thanks to a gyroscope or another orientation sensor. Therefore, in this paper, we aim to answer the following question: can the knowledge of rotation from a gyroscope be used to improve feature matching. To analyze this new approach of camera and gyroscope data fusion, we proceed in two steps. First, we rotationally align the images using rotation information of the gyroscope. And second, we compare the quality of feature matching in the original and rotationally aligned images. Experimental results on a real catadioptric sequence show that gyroscope data permits to sensibly improve the number of inliers according to epipolar geometry.
Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur
ICPR1
2008 UAV Attitude estimation by vanishing points in catadioptric images
abstract
Unmanned aerial vehicles (UAV) are the subject of an increasing interest in many applications and a key requirement is the stabilization of the vehicle. Some previous works have suggested using catadioptric vision, instead of traditional perspective cameras, in order to gather much more information from the environment and therefore improve the robustness of the UAV attitude estimation. This paper belongs to a series of recent publications of our research group concerning catadioptric vision for UAVs. Currently, we focus on the estimation of the complete attitude of a UAV flying in urban environment. In order to avoid the limitations of horizon-based approaches, the difficulties of traditional epipolar methods (such as rotation-translation ambiguity, lack of features, retrieving motion parameters from matrix decomposition, etc..) and improve UAV dynamic control, we suggest computing infinite homography. We show how catadioptric vision plays a key role to: first, extract a large number of lines, second robustly estimate the associated vanishing points and third, track them even during long video sequences. Therefore it is not only possible to estimate the relative rotation between consecutive frames but also compute the absolute rotation between two distant frames without error accumulation. Finally, we present some experimental results with ground truth data to demonstrate the accuracy and the robustness of our method.
Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur
ICRA1
2008 A robust top-down approach for rotation estimation and vanishing points extraction by catadioptric vision in urban environment
abstract
A key requirement for unmanned aerial vehicles (UAV) applications is the attitude stabilization of the aircraft, which requires the knowledge of its orientation. It is now well established that traditional navigation equipments, like GPS or INS, suffer from several disadvantages. That is why some works have suggested a vision-based approach of the problem. Especially, catadioptric vision is more and more used since it permits to gather much more information from the environment, compared to traditional perspective cameras, and therefore the robustness of the UAV attitude estimation is improved. Rotation estimation from conventional and catadioptric images has been extensively studied. Whereas interesting results can be obtained, the existing methods have non-negligible limitations such as difficult features matching (e.g. repeated texture, blurring or illumination changing) or a high computational cost (e.g. vanishing point extraction or analyze in frequency domain). In order to overcome these limitations, this paper presents a top-down approach for estimating the rotation and extracting the vanishing points in catadioptric images. This new framework is accurate and can run in real-time. To obtain the ground truth data, we also calibrate our catadioptric camera with a gyroscope. Finally, experimental results on a real video sequence are presented and compared to the ground truth data obtained by the gyroscope.
Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur
IROS1
2008 Automatic calibration of catadioptric cameras in urban environment
abstract
Camera calibration is an important step for vision-based stabilization of unmanned aerial vehicles (UAV). The goal of this paper is to develop a method for automatic calibration of a catadioptric camera so that it can be easily run before mounting the camera on the UAV or even during the flight to deal with vibrations or shocks. Whereas existing works can provide interesting results, they suffer from several practical limitations (manual line extraction, inaccurate conic fitting, calibration pattern, camera motion, execution time, etc...) and therefore cannot be applied in our application. The proposed algorithm aims to determine the most probable calibration that verifies some geometric constraints induced by catadioptric projection. In order to efficiently maximize this probability, we use a particle filtering approach. Experimental results have demonstrated the effectiveness of the proposed method.
Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur
IROS1
2007 Rectangle Extraction in Catadioptric Images
abstract
Nowadays, robotic systems are more and more equipped with catadioptric cameras. However several problems associated to catadioptric vision have been studied only slightly. Especially algorithms for detecting rectangles in catadioptric images have not yet been developed whereas it is required in diverse applications such as building extraction in aerial images. We show that working in the equivalent sphere provides an appropriate framework to detect lines, parallelism, orthogonality and therefore rectangles. Finally, we present experimental results on synthesized and real data.
Jean-Charles Bazin, In-So Kweon, Cédric Demonceaux, Pascal Vasseur
ICCV1