EDBT 2026 Demo / reviewers in the wild / expert
Shohei Nobuhara
dblp:22/4064
· DBLP profile ↗
50ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0002-3204-8696ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 37 · 23 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REACH: Hand Pose Estimation from Room Corners
Shu Nakamura, Ryo Kawahara, Genki Kinoshita, Ryosuke Hirai, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino |
FG | 6 |
| 2025 | Physical Plausibility-aware Trajectory Prediction via Locomotion EmbodimentabstractHumans can predict future human trajectories even from momentary observations by using human pose-related cues. However, previous Human Trajectory Prediction (HTP) methods leverage the pose cues implicitly, resulting in implausible predictions. To address this, we propose Locomotion Embodiment, a framework that explicitly evaluates the physical plausibility of the predicted trajectory by locomotion generation under the laws of physics. While the plausibility of locomotion is learned with an indifferentiable physics simulator, it is replaced by our differentiable Locomotion Value function to train an HTP network in a data-driven manner. In particular, our proposed Embodied Locomotion loss is beneficial for efficiently training a stochastic HTP network using multiple heads. Furthermore, the Locomotion Value filter is proposed to filter out implausible trajectories at inference. Experiments demonstrate that our method enhances even the state-of-the-art HTP methods across diverse datasets and problem settings. Our code is available at: https://github.com/ImIntheMiddle/EmLoco. Hiromu Taketsugu, Takeru Oba, Takahiro Maeda 0001, Shohei Nobuhara, Norimichi Ukita |
CVPR | 4 |
| 2024 | DeepShaRM: Multi-View Shape and Reflectance Map Recovery Under Unknown LightingabstractGeometry reconstruction of textureless, non-Lambertian objects under unknown natural illumination (i.e., in the wild) remains challenging as correspondences cannot be established and the reflectance cannot be expressed in simple analytical forms. We derive a novel multi-view method, DeepShaRM, that achieves state-of-the-art accuracy on this challenging task. Unlike past methods that formulate this as inverse-rendering, i.e., estimation of reflectance, illumination, and geometry from images, our key idea is to realize that reflectance and illumination need not be disentangled and instead estimated as a compound reflectance map. We introduce a novel deep reflectance map estimation network that recovers the camera-view reflectance maps from the surface normals of the current geometry estimate and the input multi-view images. The network also explicitly estimates per-pixel confidence scores to handle global light transport effects. A deep shape-from-shading network then updates the geometry estimate expressed with a signed distance function using the recovered reflectance maps. By alternating between these two, and, most important, by bypassing the ill-posed problem of reflectance and illumination decomposition, the method accurately recovers object geometry in these challenging settings. Extensive experiments on both synthetic and real-world data clearly demonstrate its state-of-the-art accuracy. Kohei Yamashita 0001, Shohei Nobuhara, Ko Nishino |
3DV | 2 |
| 2024 | SPIDeRS: Structured Polarization for Invisible Depth and Reflectance SensingabstractCan we capture shape and reflectance in stealth? Such capability would be valuable for many application domains in vision, xR, robotics, and HCI. We introduce structured polarization for invisible depth and reflectance sensing (SPIDeRS), the first depth and reflectance sensing method using patterns of polarized light. The key idea is to modulate the angle of linear polarization (AoLP) of projected light at each pixel. The use of polarization makes it invisible and lets us recover not only depth but also directly surface normals and even reflectance. We implement SPIDeRS with a liquid crystal spatial light modulator (SLM) and a polarimetric camera. We derive a novel method for robustly extracting the projected structured polarization pattern from the polarimetric object appearance. We evaluate the effectiveness of SPIDeRS by applying it to a number of real-world objects. The results show that our method successfully reconstructs object shapes of various materials and is robust to diffuse reflection and ambient light. We also demonstrate relighting using recovered surface normals and reflectance. We believe SPIDeRS opens a new avenue of polarization use in visual sensing. Tomoki Ichikawa, Shohei Nobuhara, Ko Nishino |
CVPR | 2 |
| 2024 | Fooling Polarization-Based Vision Using Locally Controllable Polarizing ProjectionabstractPolarization is a fundamental property of light that encodes abundant information regarding surface shape, material, illumination and viewing geometry. The computer vision community has witnessed a blossom of polarization-based vision applications, such as reflection removal, shape-from-polarization (SfP), transparent object segmentation and color constancy, partially due to the emergence of single-chip mono/color polarization sensors that make polarization data acquisition easier than ever. However, is polarization-based vision vulnerable to adversarial attacks? If so, is that possible to realize these adversarial attacks in the physical world, without being perceived by human eyes? In this paper, we warn the community of the vulnerability of polarization-based vision, which can be more serious than RGB-based vision. By adapting a commercial LCD projector, we achieve locally controllable polarizing projection, which is successfully utilized to fool state-of-the-art polarization-based vision algorithms for glass segmentation and SfP. Compared with existing physical attacks on RGB-based vision, which always suffer from the trade-off between attack efficacy and eye conceivability, the adversarial attackers based on polarizing projection are contact-free and visually imperceptible, since naked human eyes can rarely perceive the difference of viciously manipulated polarizing light and ordinary illumination. This poses unprecedented risks on polarization-based vision, for which due attentions should be paid and counter measures be considered. Zhuoxiao Li, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
CVPR | 3 |
| 2024 | KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter
Yifan Zhan, Zhuoxiao Li, Muyao Niu, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
ECCV (45) | 5 |
| 2024 | RGB road scene material segmentationabstractWe introduce RGB road scene material segmentation, i.e. , per-pixel segmentation of materials in real-world driving views with pure RGB images , as a novel computer vision task by building a benchmark dataset and by deriving a new method. Our dataset, KITTI-Materials, is based on the well-established KITTI dataset and consists of 1000 frames covering 24 different road scenes of urban/suburban landscapes, carefully annotated with one of 20 material categories for every pixel. It is the first dataset for RGB material segmentation in real driving scenes. Through careful analysis of KITTI-Materials, we identify the extraction and fusion of texture and image context as the key to accurate modeling of road scene material appearance. For this, we introduce R oad scene M aterial S egmentation Net work ( RMSNet ) as a baseline method for this challenging task. RMSNet encodes multi-scale hierarchical features with efficient Transformer layers. We construct the decoder of RMSNet based on a novel efficient self-attention model, which we refer to as SAMixer which adaptively fuses texture and context cues across multiple feature levels. Extensive experiments on KITTI-Materials validate the effectiveness of our RMSNet. We believe our work lays a solid foundation for further studies on RGB road scene material segmentation. Sudong Cai, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino |
Image Vis. Comput. | 3 |
| 2024 | SAN: Structure-aware attention network for dyadic human relation recognition in images
Kaen Kogashi, Shohei Nobuhara, Ko Nishino |
Multim. Tools Appl. | 2 |
| 2023 | Fresnel Microfacet BRDF: Unification of Polari-Radiometric Surface-Body ReflectionabstractComputer vision applications have heavily relied on the linear combination of Lambertian diffuse and microfacet specular reflection models for representing reflected radiance, which turns out to be physically incompatible and limited in applicability. In this paper, we derive a novel analytical reflectance model, which we refer to as Fresnel Microfacet BRDF model, that is physically accurate and generalizes to various real-world surfaces. Our key idea is to model the Fresnel reflection and transmission of the surface microgeometry with a collection of oriented mirror facets, both for body and surface reflections. We carefully derive the Fresnel reflection and transmission for each microfacet as well as the light transport between them in the subsurface. This physically-grounded modeling also allows us to express the polarimetric behavior of reflected light in addition to its radiometric behavior. That is, FMBRDF unifies not only body and surface reflections but also light reflection in radiometry and polarization and represents them in a single model. Experimental results demonstrate its effectiveness in accuracy, expressive power, image-based estimation, and geometry recovery. Tomoki Ichikawa, Yoshiki Fukao, Shohei Nobuhara, Ko Nishino |
CVPR | 3 |
| 2023 | Teleidoscopic Imaging System for Microscale 3D Shape ReconstructionabstractThis paper proposes a practical method of microscale 3D shape capturing by a teleidoscopic imaging system. The main challenge in microscale 3D shape reconstruction is to capture the target from multiple viewpoints with a large enough depth-of-field. Our idea is to employ a teleidoscopic measurement system consisting of three planar mirrors and monocentric lens. The planar mirrors virtually define multiple viewpoints by multiple reflections, and the monocentric lens realizes a high magnification with less blurry and surround view even in closeup imaging. Our contributions include, a structured ray-pixel camera model which handles refractive and reflective projection rays efficiently, analytical evaluations of depth of field of our teleidoscopic imaging system, and a practical calibration algorithm of the teleidoscopic imaging system. Evaluations with real images prove the concept of our measurement system. Ryo Kawahara, Meng-Yu Kuo, Shohei Nobuhara |
CVPR | 3 |
| 2023 | DeePoint: Visual Pointing Recognition and Direction EstimationabstractIn this paper, we realize automatic visual recognition and direction estimation of pointing. We introduce the first neural pointing understanding method based on two key contributions. The first is the introduction of a first-of-its-kind large-scale dataset for pointing recognition and direction estimation, which we refer to as the DP Dataset. DP Dataset consists of more than 2 million frames of 33 people pointing in various styles annotated for each frame with pointing timings and 3D directions. The second is DeePoint, a novel deep network model for joint recognition and 3D direction estimation of pointing. DeePoint is a Transformer-based network which fully leverages the spatio-temporal coordination of the body parts, not just the hands. Through extensive experiments, we demonstrate the accuracy and efficiency of DeePoint. We believe DP Dataset and DeePoint will serve as a sound foundation for visual human intention understanding. Shu Nakamura, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino |
ICCV | 3 |
| 2023 | NeRFrac: Neural Radiance Fields through Refractive SurfaceabstractNeural Radiance Fields (NeRF) is a popular neural representation for novel view synthesis. By querying spatial points and view directions, a multilayer perceptron (MLP) can be trained to output the volume density and radiance along a ray, which lets us render novel views of the scene. The original NeRF and its recent variants, however, are limited to opaque scenes dominated with diffuse reflection surfaces and cannot handle complex refractive surfaces well. We introduce NeRFrac to realize neural novel view synthesis of scenes captured through refractive surfaces, typically water surfaces. For each queried ray, an MLP-based Refractive Field is trained to estimate the distance from the ray origin to the refractive surface. A refracted ray at each intersection point is then computed by Snell’s Law, given the input ray and the approximated local normal. Points of the scene are sampled along the refracted ray and are sent to a Radiance Field for further radiance estimation. We show that from a sparse set of images, our model achieves accurate novel view synthesis of the scene underneath the refractive surface and simultaneously reconstructs the refractive surface. We evaluate the effectiveness of our method with synthetic and real scenes seen through water surfaces. Experimental results demonstrate the accuracy of NeRFrac for modeling scenes seen through wavy refractive surfaces. Github page: https://github.com/Yifever20002/NeRFrac. Yifan Zhan, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
ICCV | 2 |
| 2023 | nLMVS-Net: Deep Non-Lambertian Multi-View StereoabstractWe introduce a novel multi-view stereo (MVS) method that can simultaneously recover not just per-pixel depth but also surface normals, together with the reflectance of textureless, complex non-Lambertian surfaces captured under known but natural illumination. Our key idea is to formulate MVS as an end-to-end learnable network, which we refer to as nLMVS-Net, that seamlessly integrates radiometric cues to leverage surface normals as view-independent surface features for learned cost volume construction and filtering. It first estimates surface normals as pixel-wise probability densities for each view with a novel shape-from-shading network. These per-pixel surface normal densities and the input multi-view images are then input to a novel cost volume filtering network that learns to recover per-pixel depth and surface normal. The reflectance is also explicitly estimated by alternating with geometry reconstruction. Extensive quantitative evaluations on newly established synthetic and real-world datasets show that nLMVS-Net can robustly and accurately recover the shape and reflectance of complex objects in natural settings. Kohei Yamashita 0001, Yuto Enyo, Shohei Nobuhara, Ko Nishino |
WACV | 3 |
| 2023 | View Birdification in the Crowd: Ground-Plane Localization from Perceived Movements
Mai Nishimura, Shohei Nobuhara, Ko Nishino |
Int. J. Comput. Vis. | 2 |
| 2023 | Video Region Annotation with Sparse Bounding Boxes
Yuzheng Xu, Yang Wu 0001, Nur Sabrina binti Zuraimi, Shohei Nobuhara, Ko Nishino |
Int. J. Comput. Vis. | 4 |
| 2022 | RGB Road Scene Material Segmentation
Sudong Cai, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino |
ACCV (2) | 3 |
| 2022 | Multimodal Material SegmentationabstractRecognition of materials from their visual appearance is essential for computer vision tasks, especially those that involve interaction with the real world. Material segmentation, i.e., dense per-pixel recognition of materials, remains challenging as, unlike objects, materials do not exhibit clearly discernible visual signatures in their regular RGB appearances. Different materials, however, do lead to different radiometric behaviors, which can often be captured with non-RGB imaging modalities. We realize multimodal material segmentation from RGB, polarization, and near-infrared images. We introduce the MCubeS dataset (from MultiModal Material Segmentation) which contains 500 sets of multimodal images capturing 42 street scenes. Ground truth material segmentation as well as semantic segmentation are annotated for every image and pixel. We also derive a novel deep neural network, MCubeSNet, which learns to focus on the most informative combinations of imaging modalities for each material class with a newly derived region-guided filter selection (RGFS) layer. We use semantic segmentation as a prior to “guide” this filter selection. To the best of our knowledge, our work is the first comprehensive study on truly multimodal material segmentation. We believe our work opens new avenues of practical use of material information in safety critical applications. Yupeng Liang, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino |
CVPR | 3 |
| 2022 | Dynamic 3D Gaze from Afar: Deep Gaze Estimation from Temporal Eye-Head-Body CoordinationabstractWe introduce a novel method and dataset for 3D gaze estimation of a freely moving person from a distance, typically in surveillance views. Eyes cannot be clearly seen in such cases due to occlusion and lacking resolution. Existing gaze estimation methods suffer or fall back to approximating gaze with head pose as they primarily rely on clear, close-up views of the eyes. Our key idea is to instead leverage the intrinsic gaze, head, and body coordination of people. Our method formulates gaze estimation as Bayesian prediction given temporal estimates of head and body orientations which can be reliably estimated from a far. We model the head and body orientation likelihoods and the conditional prior of gaze direction on those with separate neural networks which are then cascaded to output the 3D gaze direction. We introduce an extensive new dataset that consists of surveillance videos annotated with 3D gaze directions captured in 5 indoor and outdoor scenes. Experimental results on this and other datasets validate the accuracy of our method and demonstrate that gaze can be accurately estimated from a typical surveillance distance even when the person's face is not visible to the camera. Soma Nonaka, Shohei Nobuhara, Ko Nishino |
CVPR | 2 |
| 2022 | Dyadic Human Relation RecognitionabstractWe introduce a new dataset and novel method for dyadic human relation (DHR) recognition. DHR recognition is a new image understanding task that concerns the recognition of the type (i.e., verb) and roles of a two-person interaction. Unlike past human action and dyadic interaction detection, our goal is to extract richer information regarding the roles of actors, i.e., which person is acting on whom. For this, we introduce the DHR-WebImages dataset which consists of a total of 22,046 images of 51 verb classes of DHR with per-image annotation of the verb and roles. We derive a novel model for DHR recognition by consolidating information from three streams corresponding to the visual, pose, and spatial relation features of the two people. Experimental results show that our model achieves accurate DHR recognition robust to severe occlusions and lacking visibility of actors. We believe these results will lead to richer visual information extraction. Kaen Kogashi, Shohei Nobuhara, Ko Nishino |
ICME | 2 |
| 2022 | Invertible Neural BRDF for Object Inverse RenderingabstractWe introduce a novel neural network-based BRDF model and a Bayesian framework for object inverse rendering, i.e., joint estimation of reflectance and natural illumination from a single image of an object of known geometry. The BRDF is expressed with an invertible neural network, namely, normalizing flow, which provides the expressive power of a high-dimensional representation, computational simplicity of a compact analytical model, and physical plausibility of a real-world BRDF. We extract the latent space of real-world reflectance by conditioning this model, which directly results in a strong reflectance prior. We refer to this model as the invertible neural BRDF model (iBRDF). We also devise a deep illumination prior by leveraging the structural bias of deep neural networks. By integrating this novel BRDF model and reflectance and illumination priors in a MAP estimation formulation, we show that this joint estimation can be computed efficiently with stochastic gradient descent. We experimentally validate the accuracy of the invertible neural BRDF model on a large number of measured data and demonstrate its use in object inverse rendering on a number of synthetic and real images. The results show new ways in which deep neural networks can help solve challenging radiometric inverse problems. Shohei Nobuhara, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Surface Normals and Shape From WaterabstractIn this paper, we introduce a novel method for reconstructing surface normals and depth of dynamic objects in water. Past shape recovery methods have leveraged various visual cues for estimating shape (e.g., depth) or surface normals. Methods that estimate both compute one from the other. We show that these two geometric surface properties can be simultaneously recovered for each pixel when the object is observed underwater. Our key idea is to leverage multi-wavelength near-infrared light absorption along different underwater light paths in conjunction with surface shading. Our method can handle both Lambertian and non-Lambertian surfaces. We derive a principled theory for this surface normals and shape from water method and a practical calibration method for determining its imaging parameters values. By construction, the method can be implemented as a one-shot imaging system. We prototype both an off-line and a video-rate imaging system and demonstrate the effectiveness of the method on a number of real-world static and dynamic objects. The results show that the method can recover intricate surface features that are otherwise inaccessible. Meng-Yu Kuo, Satoshi Murai, Ryo Kawahara, Shohei Nobuhara, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Structure of Multiple Mirror System From Kaleidoscopic Projections of Single 3D PointabstractThis paper proposes a novel algorithm of discovering the structure of a kaleidoscopic imaging system that consists of multiple planar mirrors and a camera. The kaleidoscopic imaging system can be recognized as the virtual multi-camera system and has strong advantages in that the virtual cameras are strictly synchronized and have the same intrinsic parameters. In this paper, we focus on the extrinsic calibration of the virtual multi-camera system. The problems to be solved in this paper are two-fold. The first problem is to identify to which mirror chamber each of the 2D projections of mirrored 3D points belongs. The second problem is to estimate all mirror parameters, i.e., normals, and distances of the mirrors. The key contribution of this paper is to propose novel algorithms for these problems using a single 3D point of unknown geometry by utilizing a kaleidoscopic projection constraint, which is an epipolar constraint on mirror reflections. We demonstrate the performance of the proposed algorithm of chamber assignment and estimation of mirror parameters with qualitative and quantitative evaluations using synthesized and real data. Kosuke Takahashi, Shohei Nobuhara |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | View Birdification in the Crowd: Ground-Plane Localization from Perceived Movements
Mai Nishimura, Shohei Nobuhara, Ko Nishino |
BMVC | 2 |
| 2021 | Polarimetric Normal StereoabstractWe introduce a novel method for recovering per-pixel surface normals from a pair of polarization cameras. Unlike past methods that use polarimetric observations as auxiliary features for correspondence matching, we fully integrate them in cost volume construction and filtering to directly recover per-pixel surface normals, not as byproducts of recovered disparities. Our key idea is to introduce a polarimetric cost volume of distance defined on the polarimetric observations and the polarization state computed from the surface normal. We adapt a belief propagation algorithm to filter this cost volume. The filtering algorithm simultaneously estimates the disparities and surface normals as separate entities, while effectively denoising the original noisy polarimetric observations of a quad-Bayer polarization camera. In addition, in contrast to past methods, we model polarimetric light reflection of mesoscopic surface roughness, which is essential to account for its illumination-dependency. We demonstrate the effectiveness of our method on a number of complex, real objects. Our method offers a simple and detailed 3D sensing capability for complex, non-Lambertian surfaces. Yoshiki Fukao, Ryo Kawahara, Shohei Nobuhara, Ko Nishino |
CVPR | 3 |
| 2021 | Shape From Sky: Polarimetric Normal Recovery Under the SkyabstractThe sky exhibits a unique spatial polarization pattern by scattering the unpolarized sun light. Just like insects use this unique angular pattern to navigate, we use it to map pixels to directions on the sky. That is, we show that the unique polarization pattern encoded in the polarimetric appearance of an object captured under the sky can be decoded to reveal the surface normal at each pixel. We derive a polarimetric reflection model of a diffuse plus mirror surface lit by the sun and a clear sky. This model is used to recover the per-pixel surface normal of an object from a single polarimetric image or from multiple polarimetric images captured under the sky at different times of the day. We experimentally evaluate the accuracy of our shape-from-sky method on a number of real objects of different surface compositions. The results clearly show that this passive approach to fine-geometry recovery that fully leverages the unique illumination made by nature is a viable option for 3D sensing. With the advent of quad-Bayer polarization chips, we believe the implications of our method span a wide range of domains. Tomoki Ichikawa, Matthew Purri, Ryo Kawahara, Shohei Nobuhara, Kristin J. Dana, Ko Nishino |
CVPR | 4 |
| 2021 | Human-object interaction detection with missing objects
Kaen Kogashi, Yang Wu 0001, Shohei Nobuhara, Ko Nishino |
Image Vis. Comput. | 3 |
| 2021 | Non-Rigid Shape From WaterabstractWe introduce a novel 3D sensing method for recovering a consistent, dense 3D shape of a dynamic, non-rigid object in water. The method reconstructs a complete (or fuller) 3D surface of the target object in a canonical frame (e.g., rest shape) as it freely deforms and moves between frames by estimating underwater 3D scene flow and using it to integrate per-frame depth estimates recovered from two near-infrared observations. The reconstructed shape is refined in the course of this global non-rigid shape recovery by leveraging both geometric and radiometric constraints. We implement our method with a single camera and a light source without the orthographic assumption on either by deriving a practical calibration method that estimates the point source position with respect to the camera. Our reconstruction method also accounts for scattering by water. We prototype a video-rate imaging system and show 3D shape reconstruction results on a number of real-world static, deformable, and dynamic objects and creatures in real-world water. The results demonstrate the effectiveness of the method in recovering complete shapes of complex, non-rigid objects in water, which opens new avenues of application for underwater 3D sensing in the sub-meter range. Meng-Yu Kuo, Ryo Kawahara, Shohei Nobuhara, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Video Region Annotation with Sparse Bounding Boxes
Yuzheng Xu, Yang Wu 0001, Nur Sabrina binti Zuraimi, Shohei Nobuhara, Ko Nishino |
BMVC | 4 |
| 2020 | 3D-GMNet: Single-View 3D Shape Recovery as A Gaussian Mixture
Kohei Yamashita 0001, Shohei Nobuhara, Ko Nishino |
BMVC | 2 |
| 2020 | Invertible Neural BRDF for Object Inverse Rendering
Shohei Nobuhara, Ko Nishino |
ECCV (5) | 2 |
| 2020 | Appearance and Shape from Water Reflection
Ryo Kawahara, Meng-Yu Kuo, Shohei Nobuhara, Ko Nishino |
WACV | 3 |
| 2019 | Surface Normals and Shape From WaterabstractIn this paper, we introduce a novel method for reconstructing surface normals and depth of dynamic objects in water. Past shape recovery methods have leveraged various visual cues for estimating shape (e.g., depth) or surface normals. Methods that estimate both compute one from the other. We show that these two geometric surface properties can be simultaneously recovered for each pixel when the object is observed underwater. Our key idea is to leverage multi-wavelength near-infrared light absorption along different underwater light paths in conjunction with surface shading. We derive a principled theory for this surface normals and shape from water method and a practical calibration method for determining its imaging parameters values. By construction, the method can be implemented as a one-shot imaging system. We prototype both an off-line and a video-rate imaging system and demonstrate the effectiveness of the method on a number of real-world static and dynamic objects. The results show that the method can recover intricate surface features that are otherwise inaccessible. Satoshi Murai, Meng-Yu Kuo, Ryo Kawahara, Shohei Nobuhara, Ko Nishino |
ICCV | 4 |
| 2019 | Panoptic Studio: A Massively Multiview System for Social Interaction CaptureabstractWe present an approach to capture the 3D motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent; (2) subtle motion needs to be measured over a space large enough to host a social group; (3) human appearance and configuration variation is immense; and (4) attaching markers to the body may prime the nature of interactions. The Panoptic Studio is a system organized around the thesis that social interactions should be measured through the integration of perceptual analyses over a large variety of view points. We present a modularized system designed around this principle, consisting of integrated structural, hardware, and software innovations. The system takes, as input, 480 synchronized video streams of multiple people engaged in social activities, and produces, as output, the labeled time-varying 3D structure of anatomical landmarks on individuals in the space. Our algorithm is designed to fuse the "weak" perceptual processes in the large number of views by progressively generating skeletal proposals from low-level appearance cues, and a framework for temporal refinement is also presented by associating body parts to reconstructed dense 3D trajectory stream. Our system and method are the first in reconstructing full body motion of more than five people engaged in social interactions without using markers. We also empirically demonstrate the impact of the number of views in achieving this goal. Hanbyul Joo, Tomas Simon, Xulong Li 0001, Hao Liu 0125, Sean Banerjee, Timothy Godisart, Bart C. Nabbe, Iain A. Matthews, Takeo Kanade, Shohei Nobuhara, Yaser Sheikh |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2017 | A Linear Extrinsic Calibration of Kaleidoscopic Imaging System from Single 3D PointabstractThis paper proposes a new extrinsic calibration of kaleidoscopic imaging system by estimating normals and distances of the mirrors. The problem to be solved in this paper is a simultaneous estimation of all mirror parameters consistent throughout multiple reflections. Unlike conventional methods utilizing a pair of direct and mirrored images of a reference 3D object to estimate the parameters on a per-mirror basis, our method renders the simultaneous estimation problem into solving a linear set of equations. The key contribution of this paper is to introduce a linear estimation of multiple mirror parameters from kaleidoscopic 2D projections of a single 3D point of unknown geometry. Evaluations with synthesized and real images demonstrate the performance of the proposed algorithm in comparison with conventional methods. Kosuke Takahashi, Akihiro Miyata, Shohei Nobuhara, Takashi Matsuyama |
CVPR | 3 |
| 2016 | A Single-Shot Multi-Path Interference Resolution for Mirror-Based Full 3D Shape Measurement with a Correlation-Based ToF CameraabstractThis paper is aimed at presenting a new algorithm for multi-path interference resolutions under mirror-based full 3D capture using a single correlation-based ToF camera. Our algorithm does not require additional captures or device modifications, and resolves the interference using a single ToF sensing that is also used for the 3D reconstruction as well. Evaluations with real images prove the concept of the proposed algorithm qualitatively and quantitatively. Shohei Nobuhara, Takashi Kashino, Takashi Matsuyama, Kouta Takeuchi, Kensaku Fujii |
3DV | 1 |
| 2015 | Interference-Free Epipole-Centered Structured Light Pattern for Mirror-Based Multi-view Active StereoabstractThis paper is aimed at proposing a new structured light pattern for mirror-based multi-view active stereo so that the patterns cast onto the object surface do not interfere even where the object is illuminated by the projector directly and indirectly via mirror. The key idea of our interference-free projection is to encode the projector pixel locations so that they do not collide with the code from other projector pixels by exploiting the epipolar geometry defined by the real and the virtual projectors. We prove that our new encoding does not generate code collisions between the direct and indirect patterns from the real and the virtual projectors respectively. Evaluations using real and synthesized datasets demonstrate that our approach can realize an interference-free projection without using specialized equipment such as orthographic projectors used in the state-of-the-art methods. Tomu Tahara, Ryo Kawahara, Shohei Nobuhara, Takashi Matsuyama |
3DV | 3 |
| 2015 | Panoptic Studio: A Massively Multiview System for Social Motion CaptureabstractWe present an approach to capture the 3D structure and motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent, (2) subtle motion needs to be measured over a space large enough to host a social group, and (3) human appearance and configuration variation is immense. The Panoptic Studio is a system organized around the thesis that social interactions should be measured through the perceptual integration of a large variety of view points. We present a modularized system designed around this principle, consisting of integrated structural, hardware, and software innovations. The system takes, as input, 480 synchronized video streams of multiple people engaged in social activities, and produces, as output, the labeled time-varying 3D structure of anatomical landmarks on individuals in the space. The algorithmic contributions include a hierarchical approach for generating skeletal trajectory proposals, and an optimization framework for skeletal reconstruction with trajectory re-association. Hanbyul Joo, Hao Liu 0125, Bart C. Nabbe, Iain A. Matthews, Takeo Kanade, Shohei Nobuhara, Yaser Sheikh |
ICCV | 8 |
| 2015 | A Linear Generalized Camera Calibration from Three Intersecting Reference PlanesabstractThis paper presents a new generalized (or ray-pixel, raxel) camera calibration algorithm for camera systems involving distortions by unknown refraction and reflection processes. The key idea is use of intersections of calibration planes, while conventional methods utilized collinearity constraints of points on the planes. We show that intersections of calibration planes can realize a simple linear algorithm, and that our method can be applied to any ray-distributions while conventional methods require knowing the ray-distribution class in advance. Evaluations using synthesized and real datasets demonstrate the performance of our method quantitatively and qualitatively. Mai Nishimura, Shohei Nobuhara, Takashi Matsuyama, Shinya Shimizu, Kensaku Fujii |
ICCV | 2 |
| 2015 | Cell-based visual surveillance with active cameras for 3D human gaze computation
Zhaozheng Hu, Takashi Matsuyama, Shohei Nobuhara |
Multim. Tools Appl. | 3 |
| 2015 | Augmented Motion History Volume for Spatiotemporal Editing of 3-D Video in Multiparty Interaction ScenesabstractWe present a novel method that performs spatio temporal editing of 3-D multiparty interaction scenes for free-viewpoint browsing, from separately captured 3-D video data. The main idea is to first propose the augmented motion history volume (aMHV) for individual motion representation. Then, by modeling the correlations between different aMHVs, we can define a multiparty interaction dictionary, describing the spatiotemporal constraints for different types of multi-party interaction events. Finally, a constraint satisfaction and a global optimization method synthesize natural and continuous 3-D multiparty interaction scenes. Evaluations with real data demonstrate the effectiveness of our method. Qun Shi, Shohei Nobuhara, Takashi Matsuyama |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | A 3D Shape Descriptor for Segmentation of Unstructured Meshes into Segment-Wise Coherent Mesh SeriesabstractThis paper presents a novel shape descriptor for topology-based segmentation of 3D video sequence. 3D video is a series of 3D meshes without temporal correspondences which benefit for applications including compression, motion analysis, and kinematic editing. In 3D video, both 3D mesh connectivities and the global surface topology can change frame by frame. This characteristic prevents from making accurate temporal correspondences through the entire 3D mesh series. To overcome this difficulty, we propose a two-step strategy which decomposes the entire sequence into a series of topologically coherent segments using our new shape descriptor, and then estimates temporal correspondences on a per-segment basis. We demonstrate the robustness and accuracy of the shape descriptor on real data which consist of large non-rigid motion and reconstruction errors. Tomoyuki Mukasa, Shohei Nobuhara, Tony Tung, Takashi Matsuyama |
3DV | 2 |
| 2014 | A Real-Time View-Dependent Shape Optimization for High Quality Free-Viewpoint Rendering of 3D VideoabstractThis paper is aimed at proposing a new high quality free-viewpoint rendering algorithm of 3D video. The main challenge on visualizing 3D video is how to utilize the original multi-view images used to estimate the 3D surface, and how to manage the mismatches between them due to calibration and reconstruction errors. The key idea to solve this problem is to optimize the 3D shape on a per-viewpoint basis on the fly. Given a virtual viewpoint for visualization, our algorithm optimizes the 3D shape so as to maximize the photo-consistency over the surface visible from the virtual viewpoint. An evaluation demonstrates that our method outperforms the state-of-the-art rendering qualitatively and quantitatively. Shohei Nobuhara, Wei Ning, Takashi Matsuyama |
3DV | 1 |
| 2013 | Augmented Motion History Volume for Spatiotemporal Editing of 3D Video in Multi-party Interaction ScenesabstractIn this paper we present a novel method that performs spatiotemporal editing of 3D multi-party interaction scenes for free-viewpoint browsing, from separately captured data. The main idea is to first propose the augmented Motion History Volume (aMHV) for motion representation. Then by modeling the correlations between different aMHVs we can define a multi-party interaction dictionary, describing the spatiotemporal constraints for different types of multiparty interaction events. Finally, constraint satisfaction and global optimization methods are proposed to synthesize natural and continuous multi-party interaction scenes. Experiments with real data illustrate the effectiveness of our method. Qun Shi, Shohei Nobuhara, Takashi Matsuyama |
3DV | 2 |
| 2012 | A new mirror-based extrinsic camera calibration using an orthogonality constraintabstractThis paper is aimed at calibrating the relative posture and position, i.e. extrinsic parameters, of a stationary camera against a 3D reference object which is not directly visible from the camera. We capture the reference object via a mirror under three different unknown poses, and then calibrate the extrinsic parameters from 2D appearances of reflections of the reference object in the mirrors. The key contribution of this paper is to present a new algorithm which returns a unique solution of three P3P problems from three mirrored images. While each P3P problem has up to four solutions and therefore a set of three P3P problems has up to 64 solutions, our method can select a solution based on an orthogonality constraint which should be satisfied by all families of reflections of a single reference object. In addition we propose a new scheme to compute the extrinsic parameters by solving a large system of linear equations. These two points enable us to provide a unique and robust solution. We demonstrate the advantages of the proposed method against a state-of-the-art by qualitative and quantitative evaluations using synthesized and real data. Kosuke Takahashi, Shohei Nobuhara, Takashi Matsuyama |
CVPR | 2 |
| 2011 | Structure from motion blur in low lightabstractIn theory, the precision of structure from motion estimation is known to increase as camera motion increases. In practice, larger camera motions induce motion blur, particularly in low light where longer exposures are needed. If the camera center moves during exposure, the trajectory traces in a motion-blurred image encode the underlying 3D structure of points and the motion of the camera. In this paper, we propose an algorithm to explicitly estimate the 3D structure of point light sources and camera motion from a motion-blurred image in a low light scene with point light sources. The algorithm identifies extremal points of the traces mapped out by the point sources in the image and classifies them into start and end sets. Each trace is charted out incrementally using local curvature, providing correspondences between start and end points. We use these correspondences to obtain an initial estimate of the epipolar geometry embedded in a motion-blurred image. The reconstruction and the 2D traces are used to estimate the motion of the camera during the interval of capture, and multiple view bundle adjustment is applied to refine the estimates. Shohei Nobuhara, Yaser Sheikh |
CVPR | 2 |
| 2009 | Complete multi-view reconstruction of dynamic scenes from probabilistic fusion of narrow and wide baseline stereoabstractThis paper presents a novel approach to achieve accurate and complete multi-view reconstruction of dynamic scenes (or 3D videos). 3D videos consist in sequences of 3D models in motion captured by a surrounding set of video cameras. To date 3D videos are reconstructed using multiview wide baseline stereo (MVS) reconstruction techniques. However it is still tedious to solve stereo correspondence problems: reconstruction accuracy falls when stereo photo-consistency is weak, and completeness is limited by self-occlusions. Most MVS techniques were indeed designed to deal with static objects in a controlled environment and therefore cannot solve these issues. Hence we propose to take advantage of the image content stability provided by each single-view video to recover any surface regions visible by at least one camera. In particular we present an original probabilistic framework to derive and predict the true surface of models. We propose to fuse multi-view structure-from-motion with robust 3D features obtained by MVS in order to significantly improve reconstruction completeness and accuracy. A min-cut problem where all exact features serve as priors is solved in a final step to reconstruct the 3D models. In addition, experimental results were conducted on synthetic and challenging real world datasets to illustrate the robustness and accuracy of our method. Tony Tung, Shohei Nobuhara, Takashi Matsuyama |
ICCV | 2 |
| 2009 | The Multiple-Camera 3-D Production StudioabstractMultiple-camera systems are currently widely used in research and development as a means of capturing and synthesizing realistic 3-D video content. Studio systems for 3-D production of human performance are reviewed from the literature, and the practical experience gained in developing prototype studios is reported across two research laboratories. System design should consider the studio backdrop for foreground matting, lighting for ambient illumination, camera acquisition hardware, the camera configuration for scene capture, and accurate geometric and photometric camera calibration. A ground-truth evaluation is performed to quantify the effect of different constraints on the multiple-camera system in terms of geometric accuracy and the requirement for high-quality view synthesis. As changing camera height has only a limited influence on surface visibility, multiple-camera sets or an active vision system may be required for wide area capture, and accurate reconstruction requires a camera baseline of 25deg, and the achievable accuracy is 5-10-mm at current camera resolutions. Accuracy is inherently limited, and view-dependent rendering is required for view synthesis with sub-pixel accuracy where display resolutions match camera resolutions. The two prototype studios are contrasted and state-of-the-art techniques for 3-D content production demonstrated. Jonathan Starck, Atsuto Maki, Shohei Nobuhara, Adrian Hilton 0001, Takashi Matsuyama |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Simultaneous super-resolution and 3D video using graph-cutsabstractThis paper presents a new method to increase the quality of 3D video, a new media developed to represent 3D objects in motion. This representation is obtained from multi-view reconstruction techniques that require images recorded simultaneously by several video cameras. All cameras are calibrated and placed around a dedicated studio to fully surround the models. The limited quality and quantity of cameras may produce inaccurate 3D model reconstruction with low quality texture. To overcome this issue, first we propose super-resolution (SR) techniques for 3D video: SR on multi-view images and SR on single-view video frames. Second, we propose to combine both super-resolution and dynamic 3D shape reconstruction problems into a unique Markov Random Field (MRF) energy formulation. The MRF minimization is performed using graph-cuts. Thus, we jointly compute the optimal solution for super-resolved texture and 3D shape model reconstruction. Moreover, we propose a coarse-to-fine strategy to iteratively produce 3D video with increasing quality. Our experiments show the accuracy and robustness of the proposed technique on challenging 3D video sequences. Tony Tung, Shohei Nobuhara, Takashi Matsuyama |
CVPR | 2 |
| 2008 | Complex human motion estimation using visibilityabstractThis paper presents a novel algorithm for estimating complex human motion from 3D video. We base our algorithm on a model-based approach which uses a complete surface mesh of a 3D human model to be matched with 3D video data. This type of method usually works well against partly incomplete input data. However, it fails to estimate what we call ldquocomplex motionrdquo: where some parts of the body touch each other for a long period. It is because the touching deteriorates the visibility of the neighbouring surface, which causes matching failures. In order to solve this problem, we introduce a ldquovisibilityrdquo measure for each mesh vertex that represents how it is occluded or missed on the observed surface. Using the ldquovisibilityrdquo we selectively suppress the outliers caused by low observability while traditional surface matching algorithms try to find corresponding area for the entire surface and cannot converge to the real posture by definition. Our algorithm shows improvements over naive surface matching algorithm on both synthesized and real 3D video. Tomoyuki Mukasa, Arata Miyamoto, Shohei Nobuhara, Atsuto Maki, Takashi Matsuyama |
FG | 3 |
| 2004 | Real-time 3D shape reconstruction, dynamic 3D mesh deformation, and high fidelity visualization for 3D video
Takashi Matsuyama, Takeshi Takai, Shohei Nobuhara |
Comput. Vis. Image Underst. | 4 |