EDBT 2026 Demo / reviewers in the wild / expert
Diego Thomas
dblp:84/9447 · also Diego Gabriel Francis Thomas
· DBLP profile ↗
39ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-8525-7133ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 8 since 2021Systems, architecture and hardware · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCCVT: Differentiable Clipped Centroidal Voronoi TessellationabstractWhile Marching Cubes (MC) and Marching Tetrahedra (MTet) are widely adopted in 3D reconstruction pipelines due to their simplicity and efficiency, their differentiable variants remain suboptimal for mesh extraction. This often limits the quality of 3D meshes reconstructed from point clouds or images in learning-based frameworks. In contrast, clipped CVTs offer stronger theoretical guarantees and yield higher-quality meshes. However, the lack of a differentiable formulation has prevented their integration into modern machine learning pipelines. To bridge this gap, we propose DCCVT, a differentiable algorithm that extracts high-quality 3D meshes from noisy signed distance fields (SDFs) using clipped CVTs. We derive a fully differentiable formulation for computing clipped CVTs and demonstrate its integration with deep learning-based SDF estimation to reconstruct accurate 3D meshes from input point clouds. Our experiments with synthetic data demonstrate the superior ability of DCCVT against state-of-theart methods in mesh quality and reconstruction fidelity. https://wylliamcantincharawi.dev/DCCVT.github.io/ Wylliam Cantin Charawi, Adrien Gruson, Jane Wu, Christian Desrosiers, Diego Thomas |
3DV | 5 |
| 2025 | ProbeSDF: Light Field Probes For Neural Surface ReconstructionabstractSDF-based differential rendering frameworks have achieved state-of-the-art multiview 3D shape reconstruction. In this work, we re-examine this family of approaches by minimally reformulating its core appearance model in a way that simultaneously yields faster computation and increased performance. To this goal, we exhibit a physically-inspired minimal radiance parametrization decoupling angular and spatial contributions, by encoding them with a small number of features stored in two respective volumetric grids of different resolutions. Requiring as little as four parameters per voxel, and a tiny MLP call inside a single fully fused kernel, our approach allows to enhance performance with both surface and image (PSNR) metrics, while providing a significant training speedup and real-time rendering. We show this performance to be consistently achieved on real data over two widely different and popular application fields, generic object and human subject shape reconstruction, using four representative and challenging datasets.1 Briac Toussaint, Diego Thomas, Jean-Sébastien Franco |
CVPR | 2 |
| 2025 | Neural SDF for Shadow-Aware Unsupervised Structured LightabstractAmong various active 3D measurement techniques, Structured Light (SL) is one of the most popular methods for its robustness and high accuracy. The ordinary SL system consists of a camera and a projector, and by projecting a pre-defined pattern, we can obtain pixel-to-pixel correspondences between the camera and the projector for triangulation. However, if we lack knowledge of the projected pattern for some reason, e.g., the projected pattern is not as expected due to lens distortion, inaccurate calibration, undesired optical phenomena like inter-reflection, and so on, the accuracy of conventional SL is severely degraded. As a remedy, we propose unsupervised structured light (USSL), which does not explicitly use prior knowledge of the pattern. Inspired by the fact that humans can recognize the scene structure illuminated by an unknown light source (e.g. rotating mirror ball), and some prior works have succeeded in novel-view-synthesis under unknown illumination conditions, we implement USSL on Neural Signed Distance Fields (Neural SDF) pipeline with implicit reflection module powered by a neural network. Additionally, since every SL method causes occlusion (shadow) by pattern projection, we must consider it for accurate shape reconstruction. To this end, we integrate shadow volume rendering into the proposed pipeline. Experiments with synthetic and real datasets are conducted to confirm the feasibility of the proposed method. Kazuto Ichimaru, Diego Thomas, Takafumi Iwaguchi, Hiroshi Kawasaki |
WACV | 2 |
| 2025 | VortSDF: 3D Modeling with Centroidal Voronoi Tessellation on Signed Distance FieldabstractVolumetric shape representations have become ubiquitous in multi-view reconstruction tasks. They often build on regular voxel grids as discrete representations of 3D shape functions, such as SDF or radiance fields, either as the full shape model or as sampled instantiations of continuous representations, as with neural networks. Despite their proven efficiency, voxel representations come with the precision versus complexity trade-off. This inherent limitation can significantly impact performance when moving away from simple and uncluttered scenes. In this paper we investigate an alternative discretization strategy with the Centroidal Voronoi Tessellation (CVT). CVTs allow to better partition the observation space with respect to shape occupancy and to focus the discretization around shape surfaces. To leverage this discretization strategy for multi-view reconstruction, we introduce a volumetric optimization framework that combines explicit SDF fields with a shallow color network, in order to estimate 3D shape properties over tetrahedral grids. Experimental results with Chamfer statistics validate this approach with unprecedented reconstruction quality on various scenarios such as objects, open scenes or human. Diego Thomas, Briac Toussaint, Jean-Sébastien Franco, Edmond Boyer |
WACV | 1 |
| 2025 | Sparse-View 3D Reconstruction of Clothed Humans via Normal MapsabstractWe present a novel deep learning-based approach to the 3D reconstruction of clothed humans using weak supervision via 2D normal maps. Given a single RGB image or multiview images, our network is optimized to infer a person-specific signed distance function (SDF) discretized on a tetrahedral mesh surrounding the body in a rest pose. Subsequently, estimated human pose and camera parameters are used to generate a normal map from the SDF. A key aspect of our approach is the direct use of the Marching Tetrahedra algorithm in end-to-end optimization, and in order to do so we derive analytical gradients to facilitate straightforward differentiation (and thus backpropagation). Additionally, predicted normal maps allow us to leverage pretrained image-to-normal networks in order to minimize a surface error instead of a photometric error. We demonstrate the efficacy of our approach on both labeled and in-the-wild data in the context of existing clothed human reconstruction methods. Jane Wu, Diego Thomas, Ronald Fedkiw |
WACV | 2 |
| 2024 | ActiveNeuS: Neural Signed Distance Fields for Active Stereoabstract3D-shape reconstruction in extreme environments, such as low illumination or scattering condition, has been an open problem and intensively researched. Active stereo is one of potential solution for such environments for its robustness and high accuracy. However, active stereo systems usually consist of specialized system configurations with complicated algorithms, which narrow their application. In this paper, we propose Neural Signed Distance Field for active stereo systems to enable implicit correspondence search and triangulation in generalized Structured Light. With our technique, textureless or equivalent surfaces by low light condition are successfully reconstructed even with a small number of captured images. Experiments were conducted to confirm that the proposed method could achieve state-of-the-art reconstruction quality under such severe condition. We also demonstrated that the proposed method worked in an underwater scenario. Kazuto Ichimaru, Takaki Ikeda, Diego Thomas, Takafumi Iwaguchi, Hiroshi Kawasaki |
3DV | 3 |
| 2024 | Neural Active Structure-from-Motion in Dark and Textureless Environment
Kazuto Ichimaru, Diego Thomas, Takafumi Iwaguchi, Hiroshi Kawasaki |
ACCV (10) | 2 |
| 2024 | A Practical Calibration Method for Cameras and Multiple Line-Lasers in Light Sectioning Systems for Underwater EnvironmentsabstractIn recent years, the increasing demand for underwater 3D measurement for various applications has brought about challenges such as low accuracy in 3D shape acquisition and difficulties in localizing sensor positions. This paper introduces a robust calibration method for underwater 3D sensors, comprising line lasers and cameras, utilizing a physically accurate model. Specifically, our proposed line laser calibration method estimates laser plane parameters using two types of planar constraints, avoiding the need for the costly process of backward-projection of refraction for optimization. For camera calibration, we advocate a two-step approach incorporating a simple yet effective deep-learning-based marker detection algorithm to estimate parameters of refraction, representing a physically correct lens model. Through experiments, we validate the superior performance of our methods over previous approximation-based approaches, as demonstrated in simulations and actual experiments conducted in a swimming pool. Takaki Ikeda, Takafumi Iwaguchi, Diego Thomas, Hiroshi Kawasaki |
ICIP | 3 |
| 2024 | Two-stage pose optimization algorithm using color information for underwater SLAM with light-sectioning-based 3D scanning methodabstractThe demand for 3D shape measurement of underwater scene is increasing in various applications. Especially, simultaneous localization and mapping (SLAM) technique utilizing remotely operated vehicle (ROV) attached with 3D sensors has been intensively researched. This paper focuses on solving pose optimization problem for underwater robots with camera/multiple-line-lasers setup, especially for the scene with some textures (color information). To this end, a two-stage pose optimization technique is proposed. In the first stage, due to the sparse nature of the reconstructed shape in the light-sectioning method consisting of several 3D curves, we bundle 10 to 20 consecutive frames to form a block shape, refining significant errors in the initial sensor poses using a novel bundle adjustment algorithm. In the second stage, remaining pose errors are corrected by a block-based matching algorithm utilizing iterative closest point (ICP) algorithm with color information. Through experiments in underwater environment with a real system, it was validated that the proposed method demonstrates superior performance compared to past underwater SLAM techniques. Takaki Ikeda, Takafumi Iwaguchi, Diego Thomas, Hiroshi Kawasaki |
IROS | 3 |
| 2024 | Millimetric Human Surface Capture in MinutesabstractInternational audience Briac Toussaint, Laurence Boissieux, Diego Thomas, Edmond Boyer, Jean-Sébastien Franco |
SIGGRAPH Asia | 3 |
| 2024 | Fast direct multi-person radiance fields from sparse input with dense pose priorsabstractVolumetric radiance fields have been popular in reconstructing small-scale 3D scenes from multi-view images. With additional constraints such as person correspondences, reconstructing a large 3D scene with multiple persons becomes possible. However, existing methods fail for sparse input views or when person correspondences are unavailable. In such cases, the conventional depth image supervision may be insufficient because it only captures the relative position of each person with respect to the camera center. In this paper, we investigate an alternative approach by supervising the optimization framework with a dense pose prior that represents correspondences between the SMPL model and the input images. The core ideas of our approach consist in exploiting dense pose priors estimated from the input images to perform person segmentation and incorporating such priors into the learning of the radiance field. Our proposed dense pose supervision is view-independent, significantly speeding up computational time and improving 3D reconstruction accuracy , with less floaters and noise. We confirm the advantages of our proposed method with extensive evaluation in a subset of the publicly available CMU Panoptic dataset. When training with only five input views, our proposed method achieves an average improvement of 6.1% in PSNR , 3.5% in SSIM , 17.2% in LPIPS vgg , 19.3% in LPIPS alex , and 39.4% in training time. Joao Paulo Silva do Monte Lima, Hideaki Uchiyama, Diego Thomas, Veronica Teichrieb |
Comput. Graph. | 3 |
| 2023 | A Two-Step Approach for Interactive Animatable Avatars
Takumi Kitamura, Naoya Iwamoto, Hiroshi Kawasaki, Diego Thomas |
CGI | 4 |
| 2022 | Unsupervised Multi-view Multi-person 3D Pose Estimation Using Reprojection Error
Diógenes Wallis de França Silva, Joao Paulo Silva do Monte Lima, David Macedo, Cleber Zanchettin, Diego Thomas, Hideaki Uchiyama, Veronica Teichrieb |
ICANN (3) | 5 |
| 2022 | Self-calibration of multiple-line-lasers based on coplanarity and Epipolar constraints for wide area shape scan using moving cameraabstractHigh-precision three-dimensional scanning systems have been intensively researched and developed. Recently, for acquisition of large scale scene with high density, simultaneous localisation and mapping (SLAM) technique is preferred because of its simplicity; a single sensor that is moved around freely during 3D scanning. However, to integrate multiple scans, captured data as well as position of each sensor must be highly accurate, making these systems difficult to use in environments not accessible by humans, such as underwater, internal body, or outer space. In this paper, we propose a new, flexible system with multiple line lasers that reconstructs dense and accurate 3D scenes. The advantages of our proposed system are (1) no need of synchronization nor precalibration between lasers and a camera, and (2) the system can reconstruct 3D scenes in extreme conditions, such as underwater. We propose a new self-calibration method leveraging coplanarity and Epipolar constraints is proposed. We also propose a new bundle adjustment (BA) technique that is tailored to the system for a dense integration of multiple line laser scans. Experimental evaluation in both air and underwater environments confirms the advantages of the proposed method. Genki Nagamatsu, Takaki Ikeda, Takafumi Iwaguchi, Diego Thomas, Jun Takamatsu, Hiroshi Kawasaki |
ICPR | 4 |
| 2022 | Deep Gesture Generation for Social Robots Using Type-Specific LibrariesabstractBody language such as conversational gesture is a powerful way to ease communication. Conversational gestures do not only make a speech more lively but also contain semantic meaning that helps to stress important information in the discussion. In the field of robotics, giving conversational agents (humanoid robots or virtual avatars) the ability to properly use gestures is critical, yet remain a task of extraordinary difficulty. This is because given only a text as input, there are many possibilities and ambiguities to generate an appropriate gesture. Different to previous works we propose a new method that explicitly takes into account the gesture types to reduce these ambiguities and generate human-like conversational gestures. Key to our proposed system is a new gesture database built on the TED dataset that allows us to map a word to one of three types of gestures: “Imagistic” gestures, which express the content of the speech, “Beat” gestures, which emphasize words, and “No gestures.” We propose a system that first maps the words in the input text to their corresponding gesture type, generate type-specific gestures and combine the generated gestures into one final smooth gesture. In our comparative experiments, the effectiveness of the proposed method was confirmed in user studies for both avatar and humanoid robot. Hitoshi Teshima, Naoki Wake, Diego Thomas, Yuta Nakashima, Hiroshi Kawasaki, Katsushi Ikeuchi |
IROS | 3 |
| 2022 | 3D pedestrian localization using multiple cameras: a generalizable approach
Joao Paulo Silva do Monte Lima, Rafael Roberto, Lucas Silva Figueiredo, Francisco Simões, Diego Thomas, Hideaki Uchiyama, Veronica Teichrieb |
Mach. Vis. Appl. | 5 |
| 2021 | PoseRN: A 2D Pose Refinement Network For Bias-Free Multi-View 3D Human Pose EstimationabstractWe propose a new 2D pose refinement network that learns to predict the human bias in the estimated 2D pose. There are biases in 2D pose estimations that are due to differences between annotations of 2D joint locations based on annotators’ perception and those defined by motion capture (MoCap) systems. These biases are crafted into publicly available 2D pose datasets and cannot be removed with existing error reduction approaches. Our proposed pose refinement network allows us to efficiently remove the human bias in the estimated 2D poses and achieve highly accurate multi-view 3D human pose estimation. Akihiko Sayo, Diego Thomas, Hiroshi Kawasaki, Yuta Nakashima, Katsushi Ikeuchi |
ICIP | 2 |
| 2021 | Self-calibrated dense 3D sensor using multiple cross line-lasers based on light sectioning method and visual odometryabstractAmong various 3D capturing systems, since the system with line lasers based on the light sectioning method is simple and accurate, it has widely attracted many developers and used for many purposes. In addition, there is no need to synchronize the camera and the laser and also the configuration of the camera and the lasers is flexible, and thus, the system can be used for extreme conditions, such as underwater. There are two open problems for the system. The first problem is a low density of the 3D shape obtained from a single image, i.e., just several curves. The second problem is the accuracy of line detection in the wild. In this paper, we propose a self-calibration method using visual odometry (VO) to bundle a large number of frames to increase the density to solve the first problem. We also propose a robust line detection algorithm using CNN to solve the second problem. Comparative experiments prove the effectiveness of our proposed method. In addition, the system was tested in the extreme condition for demonstration. Genki Nagamatsu, Jun Takamatsu, Takafumi Iwaguchi, Diego Thomas, Hiroshi Kawasaki |
IROS | 4 |
| 2020 | TetraTSDF: 3D Human Reconstruction From a Single Image With a Tetrahedral Outer ShellabstractRecovering the 3D shape of a person from its 2D appearance is ill-posed due to ambiguities. Nevertheless, with the help of convolutional neural networks (CNN) and prior knowledge on the 3D human body, it is possible to overcome such ambiguities to recover detailed 3D shapes of human bodies from single images. Current solutions, however, fail to reconstruct all the details of a person wearing loose clothes. This is because of either (a) huge memory requirement that cannot be maintained even on modern GPUs or (b) the compact 3D representation that cannot encode all the details. In this paper, we propose the tetrahedral outer shell volumetric truncated signed distance function (TetraTSDF) model for the human body, and its corresponding part connection network (PCN) for 3D human body shape regression. Our proposed model is compact, dense, accurate, and yet well suited for CNN-based regression task. Our proposed PCN allows us to learn the distribution of the TSDF in the tetrahedral volume from a single image in an end-to-end manner. Results show that our proposed method allows to reconstruct detailed shapes of humans wearing loose clothes from single RGB images. Hayato Onizuka, Zehra Hayirci, Diego Thomas, Akihiro Sugimoto, Hideaki Uchiyama, Rin-Ichiro Taniguchi |
CVPR | 3 |
| 2020 | Unsupervised 3D Human Pose Estimation in Multi-view-multi-pose Videoabstract3D human pose estimation from a single 2D video is an extremely difficult task because computing 3D geometry from 2D images is an ill-posed problem. Recent popular solutions adopt fully-supervised learning strategy, which requires to train a deep network on a large-scale ground truth dataset of 3D poses and 2D images. However, such a large-scale dataset with natural images does not exist, which limits the usability of existing methods. While building a complete 3D dataset is tedious and expensive, abundant 2D in-the-wild data is already publicly available. As a consequence, there is a growing interest in the computer vision community to design efficient techniques that use the unsupervised learning strategy, which does not require any ground truth 3D data. Such methods can be trained with only natural 2D images of humans. In this paper we propose an unsupervised method for estimating 3D human pose in videos. The standard approach for unsupervised learning is to use the Generative Adversarial Network (GAN) framework. To improve the performance of 3D human pose estimation in videos, we propose a new GAN network that enforces body consistency over frames in a video. We evaluate the efficiency of our proposed method on a public 3D human body dataset. Diego Thomas, Hiroshi Kawasaki |
ICPR | 2 |
| 2020 | On-the-fly Extrinsic Calibration of Non-Overlapping in-Vehicle Cameras based on Visual SLAM under 90-degree Backing-up ParkingabstractCalibration of relative poses between cameras is a challenging problem, known as extrinsic calibration, for nonoverlapping cameras that do not share the field of view. We propose a method for calibrating non-overlapping in-vehicle cameras placed at front, back, left and right positions by using visual SLAM(vSLAM). Our proposal is to calibrate the cameras during the motion of 90-degree backing-up parking on the fly, without using any dedicated calibration equipment. With this motion, the adjacent cameras are able to have the close field of view at different moments. The relative poses can be computed if the maps computed with vSLAM on each camera are merged by using the common structures. Therefore, we propose an efficient calibration framework with this feature. The proposed method is divided into three steps: map reconstruction with vSLAM on each camera, map merging for all the cameras, and extrinsic calibration. Especially, we propose to separately utilize the frames for vSLAM and the ones for the calibration so that the accuracy of vSLAM can be maximized for the calibration. In the evaluation, the calibration was performed in a practical environment to investigate the performance in comparison with the ground truth acquired by using a calibration equipment. Kazuki Nishiguchi, Hideaki Uchiyama, Kazutaka Hayakawa, Jun Adachi, Diego Thomas, Atsushi Shimada 0001, Rin-Ichiro Taniguchi |
IV | 5 |
| 2019 | Mobile Photometric Stereo with Keypoint-Based SLAM for Dense 3D ReconstructionabstractThe standard photometric stereo is a technique to densely reconstruct objects' surfaces using light variation under the assumption of a static camera with a moving light source. In this work, we use photometric stereo to reconstruct dense 3D scenes while moving the camera and the light altogether. In such non-static case, camera poses as well as correspondences between pixels of each frame to apply photometric stereo are required. ORB-SLAM is a technique that can be used to estimate camera poses. To retrieve correspondences, our idea is to start from a sparse 3D mesh obtained with ORB SLAM and then densify the mesh by a plane sweep method using a multi-view photometric consistency. By combining ORB-SLAM and photometric stereo, it is possible to reconstruct dense 3D scenes with a off-the-shelf smartphone and its embedded torchlight. Note that SLAM systems usually struggle with textureless object, which is effectively compensated by the photometric stereo in our method. Experiments are conducted to show that our proposed method gives better results than SLAM alone or COLMAP, especially for partially textureless surfaces. Remy Maxence, Hideaki Uchiyama, Hiroshi Kawasaki, Diego Thomas, Vincent Nozick, Hideo Saito 0001 |
3DV | 4 |
| 2019 | Revisiting Depth Image Fusion with Variational Message Passing
Diego Thomas, Ekaterina Sirazitdinova, Akihiro Sugimoto, Rin-Ichiro Taniguchi |
3DV | 1 |
| 2019 | Human Shape Reconstruction with Loose Clothes from Partially Observed Data by Pose Specific Deformation
Akihiko Sayo, Hayato Onizuka, Diego Thomas, Yuta Nakashima, Hiroshi Kawasaki, Katsushi Ikeuchi |
PSIVT | 3 |
| 2019 | 3D Positioning System Based on One-handed Thumb Interactions for 3D Annotation PlacementabstractThis paper presents a 3D positioning system based on one-handed thumb interactions for simple 3D annotation placement with a smart-phone. To place an annotation at a target point in the real environment, the 3D coordinate of the point is computed by interactively selecting the corresponding points in multiple views by users while performing SLAM. Generally, it is difficult for users to precisely select an intended pixel on the touchscreen. Therefore, we propose to compute the 3D coordinate from multiple observations with a robust estimator to have the tolerance to the inaccurate user inputs. In addition, we developed three pixel selection methods based on one-handed thumb interactions. A pixel is selected at the thumb position at a live view in FingAR, the position of a reticle marker at a live view in SnipAR, or that of a movable reticle marker at a freezed view in FreezAR. In the preliminary evaluation, we investigated the 3D positioning accuracy of each method. So Tashiro, Hideaki Uchiyama, Diego Thomas, Rin-Ichiro Taniguchi |
VR | 3 |
| 2018 | SegmentedFusion: 3D Human Body Reconstruction Using Stitched Bounding BoxesabstractThis paper presents SegmentedFusion, a method possessing the capability of reconstructing non-rigid 3D models of a human body by using a single depth camera with skeleton information. Our method estimates a dense volumetric 6D motion field that warps the integrated model into the live frame by segmenting a human body into different parts and building a canonical space for each part. The key feature of this work is that a deformed and connected canonical volume for each part is created, and it is used to integrate data. The dense volumetric warp field of one volume is represented efficiently by blending a few rigid transformations. Overall, SegmentedFusion is able to scan a non-rigidly deformed human surface as well as to estimate the dense motion field by using a consumer-grade depth camera. The experimental results demonstrate that SegmentedFusion is robust against fast inter-frame motion and topological changes. Since our method does not require prior assumption, SegmentedFusion can be applied to a wide range of human motions. Shih-Hsuan Yao, Diego Thomas, Akihiro Sugimoto, Shang-Hong Lai, Rin-Ichiro Taniguchi |
3DV | 2 |
| 2018 | Live Structural Modeling Using RGB-D SLAMabstractThis paper presents a method for localizing primitive shapes in a dense point cloud computed by the RGB-D SLAM system. To stably generate a shape map containing only primitive shapes, the primitive shape is incrementally modeled by fusing the shapes estimated at previous frames in the SLAM, so that an accurate shape can be finally generated. Specifically, the history of the fusing process is used to avoid the influence of error accumulation in the SLAM. The point cloud of the shape is then updated by fusing the points in all the previous frames into a single point cloud. In the experimental results, we show that metric primitive modeling in texture-less and unprepared environments can be achieved online. Nicolas Olivier, Hideaki Uchiyama, Masashi Mishima, Diego Thomas, Rin-Ichiro Taniguchi, Rafael Alves Roberto, Joao Paulo Silva do Monte Lima, Veronica Teichrieb |
ICRA | 4 |
| 2018 | FusionMLS: Highly dynamic 3D reconstruction with consumer-grade RGB-D camerasabstractMulti-view dynamic three-dimensional reconstruction has typically required the use of custom shutter-synchronized camera rigs in order to capture scenes containing rapid movements or complex topology changes. In this paper, we demonstrate that multiple unsynchronized low-cost RGB-D cameras can be used for the same purpose. To alleviate issues caused by unsynchronized shutters, we propose a novel depth frame interpolation technique that allows synchronized data capture from highly dynamic 3D scenes. To manage the resulting huge number of input depth images, we also introduce an efficient moving least squares-based volumetric reconstruction method that generates triangle meshes of the scene. Our approach does not store the reconstruction volume in memory, making it memory-efficient and scalable to large scenes. Our implementation is completely GPU based and works in real time. The results shown herein, obtained with real data, demonstrate the effectiveness of our proposed method and its advantages compared to state-of-the-art approaches. Siim Meerits, Diego Thomas, Vincent Nozick, Hideo Saito 0001 |
Comput. Vis. Media | 2 |
| 2017 | Fast 3D point cloud segmentation using supervoxels with geometry and color for 3D scene understandingabstractSegmentation of 3D colored point clouds is a research field with renewed interest thanks to recent availability of inexpensive consumer RGB-D cameras and its importance as an unavoidable low-level step in many robotic applications. However, 3D data's nature makes the task challenging and, thus, many different techniques are being proposed, all of which require expensive computational costs. This paper presents a novel fast method for 3D colored point cloud segmentation. It starts with supervoxel partitioning of the cloud, i.e., an oversegmentation of the points in the cloud. Then it leverages on a novel metric exploiting both geometry and color to iteratively merge the supervoxels to obtain a 3D segmentation where the hierarchical structure of partitions is maintained. The algorithm also presents computational complexity linear to the size of the input. Experimental results over two publicly available datasets demonstrate that our proposed method outperforms state-of-the-art techniques. Francesco Verdoja, Diego Thomas, Akihiro Sugimoto |
ICME | 2 |
| 2017 | Synthesis of Environment Maps for Mixed RealityabstractWhen rendering virtual objects in a mixed reality application, it is helpful to have access to an environment map that captures the appearance of the scene from the perspective of the virtual object. It is straightforward to render virtual objects into such maps, but capturing and correctly rendering the real components of the scene into the map is much more challenging. This information is often recovered from physical light probes, such as reflective spheres or fisheye cameras, placed at the location of the virtual object in the scene. For many application areas, however, real light probes would be intrusive or impractical. Ideally, all of the information necessary to produce detailed environment maps could be captured using a single device. We introduce a method using an RGBD camera and a small fisheye camera, contained in a single unit, to create environment maps at any location in an indoor scene. The method combines the output from both cameras to correct for their limited field of view and the displacement from the virtual object, producing complete environment maps suitable for rendering the virtual content in real time. Our method improves on previous probeless approaches by its ability to recover high-frequency environment maps. We demonstrate how this can be used to render virtual objects which shadow, reflect and refract their environment convincingly. David R. Walton, Diego Thomas, Anthony Steed, Akihiro Sugimoto |
ISMAR | 2 |
| 2017 | Modeling large-scale indoor scenes with rigid fragments using RGB-D cameras
Diego Thomas, Akihiro Sugimoto |
Comput. Vis. Image Underst. | 1 |
| 2017 | Parametric Surface Representation with Bump Image for Dense 3D Modeling Using an RBG-D Camera
Diego Thomas, Akihiro Sugimoto |
Int. J. Comput. Vis. | 1 |
| 2016 | Augmented Blendshapes for Real-Time Simultaneous 3D Head Modeling and Facial Motion CaptureabstractWe propose a method to build in real-time animated 3D head models using a consumer-grade RGB-D camera. Our framework is the first one to provide simultaneously comprehensive facial motion tracking and a detailed 3D model of the user's head. Anyone's head can be instantly reconstructed and his facial motion captured without requiring any training or pre-scanning. The user starts facing the camera with a neutral expression in the first frame, but is free to move, talk and change his face expression as he wills otherwise. The facial motion is tracked using a blendshape representation while the fine geometric details are captured using a Bump image mapped over the template mesh. We propose an efficient algorithm to grow and refine the 3D model of the head on-the-fly and in real-time. We demonstrate robust and high-fidelity simultaneous facial motion tracking and 3D head modeling results on a wide range of subjects with various head poses and facial expressions. Our proposed method offers interesting possibilities for animation production and 3D video telecommunications. Diego Thomas, Rin-Ichiro Taniguchi |
CVPR | 1 |
| 2016 | Multi-view facial landmark detector learned by the Structured Output SVM
Michal Uricár, Vojtech Franc, Diego Thomas, Akihiro Sugimoto, Václav Hlavác |
Image Vis. Comput. | 3 |
| 2013 | A Flexible Scene Representation for 3D Reconstruction Using an RGB-D CameraabstractUpdating a global 3D model with live RGB-D measurements has proven to be successful for 3D reconstruction of indoor scenes. Recently, a Truncated Signed Distance Function (TSDF) volumetric model and a fusion algorithm have been introduced (KinectFusion), showing significant advantages such as computational speed and accuracy of the reconstructed scene. This algorithm, however, is expensive in memory when constructing and updating the global model. As a consequence, the method is not well scalable to large scenes. We propose a new flexible 3D scene representation using a set of planes that is cheap in memory use and, nevertheless, achieves accurate reconstruction of indoor scenes from RGB-D image sequences. Projecting the scene onto different planes reduces significantly the size of the scene representation and thus it allows us to generate a global textured 3D model with lower memory requirement while keeping accuracy and easiness to update with live RGB-D measurements. Experimental results demonstrate that our proposed flexible 3D scene representation achieves accurate reconstruction, while keeping the scalability for large indoor scenes. Diego Thomas, Akihiro Sugimoto |
ICCV | 1 |
| 2013 | Learning to discover objects in RGB-D images using correlation clusteringabstractWe introduce a method to discover objects from RGB-D image collections which does not require a user to specify the number of objects expected to be found. We propose a probabilistic formulation to find pairwise similarity between image segments, using a classifier trained on labelled pairs from the recently released RGB-D Object Dataset. We then use a correlation clustering solver to both find the optimal clustering of all the segments in the collection and to recover the number of clusters. Unlike traditional supervised learning methods, our training data need not be of the same class or category as the objects we expect to discover. We show that this parameter-free supervised clustering method has superior performance to traditional clustering methods. Michael Firman, Diego Thomas, Simon J. Julier, Akihiro Sugimoto |
IROS | 2 |
| 2013 | Range Image Registration Using a Photometric Metric under Unknown LightingabstractBased on the spherical harmonics representation of image formation, we derive a new photometric metric for evaluating the correctness of a given rigid transformation aligning two overlapping range images captured under unknown, distant, and general illumination. We estimate the surrounding illumination and albedo values of points of the two range images from the point correspondences induced by the input transformation. We then synthesize the color of both range images using albedo values transferred using the point correspondences to compute the photometric reprojection error. This way allows us to accurately register two range images by finding the transformation that minimizes the photometric reprojection error. We also propose a practical method using the proposed photometric metric to register pairs of range images devoid of salient geometric features, captured under unknown lighting. Our method uses a hypothesize-and-test strategy to search for the transformation that minimizes our photometric metric. Transformation candidates are efficiently generated by employing the spherical representation of each range image. Experimental results using both synthetic and real data demonstrate the usefulness of the proposed metric. Diego Thomas, Akihiro Sugimoto |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Illumination-free photometric metric for range image registrationabstractThis paper presents an illumination-free photometric metric for evaluating the goodness of a rigid transformation aligning two overlapping range images, under the assumption of Lambertian surface. Our metric is based on photometric re-projection error but not on feature detection and matching. We synthesize the color of one image using albedo of the other image to compute the photometric re-projection error. The unknown illumination and albedo are estimated from the correspondences induced by the input transformation using the spherical harmonics representation of image formation. This way allows us to derive an illumination-free photometric metric for range image alignment. We use a hypothesize-and-test method to search for the transformation that minimizes our illumination-free photometric function. Transformation candidates are efficiently generated by employing the spherical representation of each image. Experimental results using synthetic and real data show the usefulness of the proposed metric. Diego Thomas, Akihiro Sugimoto |
WACV | 1 |
| 2011 | Robustly registering range images using local distribution of albedo
Diego Thomas, Akihiro Sugimoto |
Comput. Vis. Image Underst. | 1 |