VLDB 2026 Research / reviewers in the wild / expert
Ko Nishino
dblp:07/2706
· DBLP profile ↗
100ranked-venue papers
12as first author
36since 2021 · last 2026
0000-0002-3534-3447ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 11 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 67 · 8 first-author · 23 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REACH: Hand Pose Estimation from Room Corners
Shu Nakamura, Ryo Kawahara, Genki Kinoshita, Ryosuke Hirai, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino |
FG | 7 |
| 2025 | MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse ViewsabstractWe present a novel appearance model that simultaneously realizes explicit high-quality 3D surface mesh recovery and photorealistic novel view synthesis from sparse view samples. Our key idea is to model the underlying scene geometry Mesh as an Atlas of Charts which we render with 2D Gaussian surfels (MAtCha Gaussians). MAtCha distills high-frequency scene surface details from an off-the-shelf monocular depth estimator and refines it through Gaussian surfel rendering. The Gaussian surfels are attached to the charts on the fly, satisfying photorealism of neural volumetric rendering and crisp geometry of a mesh model, i.e., two seemingly contradicting goals in a single model. At the core of MAtCha lies a novel neural deformation model and a structure loss that preserve the fine surface details distilled from learned monocular depths while addressing their fundamental scale ambiguities. Results of extensive experimental validation demonstrate MAtCha’s state-of-the-art quality of surface reconstruction and photorealism on-par with top contenders but with dramatic reduction in the number of input views and computational time. We believe MAtCha will serve as a foundational tool for any visual application in vision, graphics, and robotics that require explicit geometry in addition to photorealism. Antoine Guédon, Tomoki Ichikawa, Kohei Yamashita 0001, Ko Nishino |
CVPR | 4 |
| 2025 | HeatFormer: A Neural Optimizer for Multiview Human Mesh RecoveryabstractWe introduce a novel method for human shape and pose recovery that can fully leverage multiple static views. We target fixed-multiview people monitoring, including elderly care and safety monitoring, in which cameras can be installed at the corners of a room or an open space but whose configuration may vary depending on the environment. Our key idea is to formulate it as neural optimization. We achieve this with HeatFormer, a neural optimizer that iteratively refines the SMPL parameters given multiview images. HeatFormer realizes this SMPL parameter estimation as heatmap generation and alignment with a novel transformer encoder and decoder. Our formulation makes HeatFormer fundamentally agnostic to the number of cameras, their configuration, and calibration. We demonstrate the effectiveness of HeatFormer including its accuracy, robustness to occlusion, and generalizability through an extensive set of experiments. We believe HeatFormer can serve a key role in passive human behavior modeling. Yuto Matsubara, Ko Nishino |
CVPR | 2 |
| 2025 | SHeaP: Self-Supervised Head Geometry Predictor Learned via 2D GaussiansabstractAccurate, real-time 3D reconstruction of human heads from monocular images and videos underlies numerous visual applications. As 3D ground truth data is hard to come by at scale, previous methods have sought to learn from abundant 2D videos in a self-supervised manner. Typically, this involves the use of differentiable mesh rendering, which is effective but faces limitations. To improve on this, we propose SHeaP (Self-supervised Head Geometry Predictor Learned via 2D Gaussians). Given a source image, we predict a 3DMM mesh and a set of Gaussians that are rigged to this mesh. We then reanimate this rigged head avatar to match a target frame, and backpropagate photometric losses to both the 3DMM and Gaussian prediction networks. We find that using Gaussians for rendering substantially improves the effectiveness of this self-supervised approach. Training solely on 2D data, our method surpasses existing self-supervised approaches in geometric evaluations on the NoW benchmark for neutral faces and a new benchmark for non-neutral expressions. Our method also produces highly expressive meshes, outperforming state-of-the-art in emotion classification. Liam Schoneveld, Davide Davoli 0002, Jiapeng Tang, Saimon Terazawa, Ko Nishino, Matthias Nießner |
ICCV | 6 |
| 2025 | Generative Perception of Shape and Material from Differential MotionabstractPerceiving the shape and material of an object from a single image is inherently ambiguous, especially when lighting is unknown and unconstrained. Despite this, humans can often disentangle shape and material, and when they are uncertain, they often move their head slightly or rotate the object to help resolve the ambiguities. Inspired by this behavior, we introduce a novel conditional denoising-diffusion model that generates samples of shape-and-material maps from a short video of an object undergoing differential motions. Our parameter-efficient architecture allows training directly in pixel-space, and it generates many disentangled attributes of an object simultaneously. Trained on a modest number of synthetic object-motion videos with supervision on shape and material, the model exhibits compelling emergent behavior: For static observations, it produces diverse, multimodal predictions of plausible shape-and-material maps that capture the inherent ambiguities; and when objects move, the distributions converge to more accurate explanations. The model also produces high-quality shape-and-material estimates for less ambiguous, real-world objects.
By moving beyond single-view to continuous motion observations, and by using generative perception to capture visual ambiguities, our work suggests ways to improve visual reasoning in physically-embodied systems. Xinran Nicole Han, Ko Nishino, Todd E. Zickler |
NeurIPS | 2 |
| 2024 | DeepShaRM: Multi-View Shape and Reflectance Map Recovery Under Unknown LightingabstractGeometry reconstruction of textureless, non-Lambertian objects under unknown natural illumination (i.e., in the wild) remains challenging as correspondences cannot be established and the reflectance cannot be expressed in simple analytical forms. We derive a novel multi-view method, DeepShaRM, that achieves state-of-the-art accuracy on this challenging task. Unlike past methods that formulate this as inverse-rendering, i.e., estimation of reflectance, illumination, and geometry from images, our key idea is to realize that reflectance and illumination need not be disentangled and instead estimated as a compound reflectance map. We introduce a novel deep reflectance map estimation network that recovers the camera-view reflectance maps from the surface normals of the current geometry estimate and the input multi-view images. The network also explicitly estimates per-pixel confidence scores to handle global light transport effects. A deep shape-from-shading network then updates the geometry estimate expressed with a signed distance function using the recovered reflectance maps. By alternating between these two, and, most important, by bypassing the ill-posed problem of reflectance and illumination decomposition, the method accurately recovers object geometry in these challenging settings. Extensive experiments on both synthetic and real-world data clearly demonstrate its state-of-the-art accuracy. Kohei Yamashita 0001, Shohei Nobuhara, Ko Nishino |
3DV | 3 |
| 2024 | Diffusion Reflectance Map: Single-Image Stochastic Inverse Rendering of Illumination and ReflectanceabstractReflectance bounds the frequency spectrum of illumination in the object appearance. In this paper, we introduce the first stochastic inverse rendering method, which recovers the attenuated frequency spectrum of an illumination jointly with the reflectance of an object of known geometry from a single image. Our key idea is to solve this blind in-verse problem in the reflectance map, an appearance representation invariant to the underlying geometry, by learning to reverse the image formation with a novel diffusion model which we refer to as the Diffusion Reflectance Map Net-work (DRMNet). Given an observed reflectance map converted and completed from the single input image, DRM-Net generates a reflectance map corresponding to a perfect mirror sphere while jointly estimating the reflectance. The forward process can be understood as gradually filtering a natural illumination with lower and lower frequency re-flectance and additive Gaussian noise. DRMNet learns to invert this process with two subnetworks, IllNet and RefNet, which work in concert towards this joint estimation. The network is trained on an extensive synthetic dataset and is demonstrated to generalize to real images, showing state-of-the-art accuracy on established datasets. Yuto Enyo, Ko Nishino |
CVPR | 2 |
| 2024 | SPIDeRS: Structured Polarization for Invisible Depth and Reflectance SensingabstractCan we capture shape and reflectance in stealth? Such capability would be valuable for many application domains in vision, xR, robotics, and HCI. We introduce structured polarization for invisible depth and reflectance sensing (SPIDeRS), the first depth and reflectance sensing method using patterns of polarized light. The key idea is to modulate the angle of linear polarization (AoLP) of projected light at each pixel. The use of polarization makes it invisible and lets us recover not only depth but also directly surface normals and even reflectance. We implement SPIDeRS with a liquid crystal spatial light modulator (SLM) and a polarimetric camera. We derive a novel method for robustly extracting the projected structured polarization pattern from the polarimetric object appearance. We evaluate the effectiveness of SPIDeRS by applying it to a number of real-world objects. The results show that our method successfully reconstructs object shapes of various materials and is robust to diffuse reflection and ambient light. We also demonstrate relighting using recovered surface normals and reflectance. We believe SPIDeRS opens a new avenue of polarization use in visual sensing. Tomoki Ichikawa, Shohei Nobuhara, Ko Nishino |
CVPR | 3 |
| 2024 | Fooling Polarization-Based Vision Using Locally Controllable Polarizing ProjectionabstractPolarization is a fundamental property of light that encodes abundant information regarding surface shape, material, illumination and viewing geometry. The computer vision community has witnessed a blossom of polarization-based vision applications, such as reflection removal, shape-from-polarization (SfP), transparent object segmentation and color constancy, partially due to the emergence of single-chip mono/color polarization sensors that make polarization data acquisition easier than ever. However, is polarization-based vision vulnerable to adversarial attacks? If so, is that possible to realize these adversarial attacks in the physical world, without being perceived by human eyes? In this paper, we warn the community of the vulnerability of polarization-based vision, which can be more serious than RGB-based vision. By adapting a commercial LCD projector, we achieve locally controllable polarizing projection, which is successfully utilized to fool state-of-the-art polarization-based vision algorithms for glass segmentation and SfP. Compared with existing physical attacks on RGB-based vision, which always suffer from the trade-off between attack efficacy and eye conceivability, the adversarial attackers based on polarizing projection are contact-free and visually imperceptible, since naked human eyes can rarely perceive the difference of viciously manipulated polarizing light and ordinary illumination. This poses unprecedented risks on polarization-based vision, for which due attentions should be paid and counter measures be considered. Zhuoxiao Li, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
CVPR | 4 |
| 2024 | Camera Height Doesn't Change: Unsupervised Training for Metric Monocular Road-Scene Depth Estimation
Genki Kinoshita, Ko Nishino |
ECCV (23) | 2 |
| 2024 | Correspondences of the Third Kind: Camera Pose Estimation from Object Reflection
Kohei Yamashita 0001, Vincent Lepetit, Ko Nishino |
ECCV (65) | 3 |
| 2024 | KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter
Yifan Zhan, Zhuoxiao Li, Muyao Niu, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
ECCV (45) | 6 |
| 2024 | Multistable Shape from Shading Emerges from Patch DiffusionabstractModels for inferring monocular shape of surfaces with diffuse reflection---shape from shading---ought to produce distributions of outputs, because there are fundamental mathematical ambiguities of both continuous (e.g., bas-relief) and discrete (e.g., convex/concave) types that are also experienced by humans. Yet, the outputs of current models are limited to point estimates or tight distributions around single modes, which prevent them from capturing these effects. We introduce a model that reconstructs a multimodal distribution of shapes from a single shading image, which aligns with the human experience of multistable perception. We train a small denoising diffusion process to generate surface normal fields from $16\times 16$ patches of synthetic images of everyday 3D objects. We deploy this model patch-wise at multiple scales, with guidance from inter-patch shape consistency constraints. Despite its relatively small parameter count and predominantly bottom-up structure, we show that multistable shape explanations emerge from this model for ambiguous test images that humans experience as being multistable. At the same time, the model produces veridical shape estimates for object-like images that include distinctive occluding contours and appear less ambiguous. This may inspire new architectures for stochastic 3D shape perception that are more efficient and better aligned with human experience. Xinran Nicole Han, Todd E. Zickler, Ko Nishino |
NeurIPS | 3 |
| 2024 | RGB road scene material segmentationabstractWe introduce RGB road scene material segmentation, i.e. , per-pixel segmentation of materials in real-world driving views with pure RGB images , as a novel computer vision task by building a benchmark dataset and by deriving a new method. Our dataset, KITTI-Materials, is based on the well-established KITTI dataset and consists of 1000 frames covering 24 different road scenes of urban/suburban landscapes, carefully annotated with one of 20 material categories for every pixel. It is the first dataset for RGB material segmentation in real driving scenes. Through careful analysis of KITTI-Materials, we identify the extraction and fusion of texture and image context as the key to accurate modeling of road scene material appearance. For this, we introduce R oad scene M aterial S egmentation Net work ( RMSNet ) as a baseline method for this challenging task. RMSNet encodes multi-scale hierarchical features with efficient Transformer layers. We construct the decoder of RMSNet based on a novel efficient self-attention model, which we refer to as SAMixer which adaptively fuses texture and context cues across multiple feature levels. Extensive experiments on KITTI-Materials validate the effectiveness of our RMSNet. We believe our work lays a solid foundation for further studies on RGB road scene material segmentation. Sudong Cai, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino |
Image Vis. Comput. | 4 |
| 2024 | SAN: Structure-aware attention network for dyadic human relation recognition in images
Kaen Kogashi, Shohei Nobuhara, Ko Nishino |
Multim. Tools Appl. | 3 |
| 2023 | Fresnel Microfacet BRDF: Unification of Polari-Radiometric Surface-Body ReflectionabstractComputer vision applications have heavily relied on the linear combination of Lambertian diffuse and microfacet specular reflection models for representing reflected radiance, which turns out to be physically incompatible and limited in applicability. In this paper, we derive a novel analytical reflectance model, which we refer to as Fresnel Microfacet BRDF model, that is physically accurate and generalizes to various real-world surfaces. Our key idea is to model the Fresnel reflection and transmission of the surface microgeometry with a collection of oriented mirror facets, both for body and surface reflections. We carefully derive the Fresnel reflection and transmission for each microfacet as well as the light transport between them in the subsurface. This physically-grounded modeling also allows us to express the polarimetric behavior of reflected light in addition to its radiometric behavior. That is, FMBRDF unifies not only body and surface reflections but also light reflection in radiometry and polarization and represents them in a single model. Experimental results demonstrate its effectiveness in accuracy, expressive power, image-based estimation, and geometry recovery. Tomoki Ichikawa, Yoshiki Fukao, Shohei Nobuhara, Ko Nishino |
CVPR | 4 |
| 2023 | DeePoint: Visual Pointing Recognition and Direction EstimationabstractIn this paper, we realize automatic visual recognition and direction estimation of pointing. We introduce the first neural pointing understanding method based on two key contributions. The first is the introduction of a first-of-its-kind large-scale dataset for pointing recognition and direction estimation, which we refer to as the DP Dataset. DP Dataset consists of more than 2 million frames of 33 people pointing in various styles annotated for each frame with pointing timings and 3D directions. The second is DeePoint, a novel deep network model for joint recognition and 3D direction estimation of pointing. DeePoint is a Transformer-based network which fully leverages the spatio-temporal coordination of the body parts, not just the hands. Through extensive experiments, we demonstrate the accuracy and efficiency of DeePoint. We believe DP Dataset and DeePoint will serve as a sound foundation for visual human intention understanding. Shu Nakamura, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino |
ICCV | 4 |
| 2023 | NeRFrac: Neural Radiance Fields through Refractive SurfaceabstractNeural Radiance Fields (NeRF) is a popular neural representation for novel view synthesis. By querying spatial points and view directions, a multilayer perceptron (MLP) can be trained to output the volume density and radiance along a ray, which lets us render novel views of the scene. The original NeRF and its recent variants, however, are limited to opaque scenes dominated with diffuse reflection surfaces and cannot handle complex refractive surfaces well. We introduce NeRFrac to realize neural novel view synthesis of scenes captured through refractive surfaces, typically water surfaces. For each queried ray, an MLP-based Refractive Field is trained to estimate the distance from the ray origin to the refractive surface. A refracted ray at each intersection point is then computed by Snell’s Law, given the input ray and the approximated local normal. Points of the scene are sampled along the refracted ray and are sent to a Radiance Field for further radiance estimation. We show that from a sparse set of images, our model achieves accurate novel view synthesis of the scene underneath the refractive surface and simultaneously reconstructs the refractive surface. We evaluate the effectiveness of our method with synthetic and real scenes seen through water surfaces. Experimental results demonstrate the accuracy of NeRFrac for modeling scenes seen through wavy refractive surfaces. Github page: https://github.com/Yifever20002/NeRFrac. Yifan Zhan, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
ICCV | 3 |
| 2023 | nLMVS-Net: Deep Non-Lambertian Multi-View StereoabstractWe introduce a novel multi-view stereo (MVS) method that can simultaneously recover not just per-pixel depth but also surface normals, together with the reflectance of textureless, complex non-Lambertian surfaces captured under known but natural illumination. Our key idea is to formulate MVS as an end-to-end learnable network, which we refer to as nLMVS-Net, that seamlessly integrates radiometric cues to leverage surface normals as view-independent surface features for learned cost volume construction and filtering. It first estimates surface normals as pixel-wise probability densities for each view with a novel shape-from-shading network. These per-pixel surface normal densities and the input multi-view images are then input to a novel cost volume filtering network that learns to recover per-pixel depth and surface normal. The reflectance is also explicitly estimated by alternating with geometry reconstruction. Extensive quantitative evaluations on newly established synthetic and real-world datasets show that nLMVS-Net can robustly and accurately recover the shape and reflectance of complex objects in natural settings. Kohei Yamashita 0001, Yuto Enyo, Shohei Nobuhara, Ko Nishino |
WACV | 4 |
| 2023 | View Birdification in the Crowd: Ground-Plane Localization from Perceived Movements
Mai Nishimura, Shohei Nobuhara, Ko Nishino |
Int. J. Comput. Vis. | 3 |
| 2023 | Video Region Annotation with Sparse Bounding Boxes
Yuzheng Xu, Yang Wu 0001, Nur Sabrina binti Zuraimi, Shohei Nobuhara, Ko Nishino |
Int. J. Comput. Vis. | 5 |
| 2022 | RGB Road Scene Material Segmentation
Sudong Cai, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino |
ACCV (2) | 4 |
| 2022 | Multimodal Material SegmentationabstractRecognition of materials from their visual appearance is essential for computer vision tasks, especially those that involve interaction with the real world. Material segmentation, i.e., dense per-pixel recognition of materials, remains challenging as, unlike objects, materials do not exhibit clearly discernible visual signatures in their regular RGB appearances. Different materials, however, do lead to different radiometric behaviors, which can often be captured with non-RGB imaging modalities. We realize multimodal material segmentation from RGB, polarization, and near-infrared images. We introduce the MCubeS dataset (from MultiModal Material Segmentation) which contains 500 sets of multimodal images capturing 42 street scenes. Ground truth material segmentation as well as semantic segmentation are annotated for every image and pixel. We also derive a novel deep neural network, MCubeSNet, which learns to focus on the most informative combinations of imaging modalities for each material class with a newly derived region-guided filter selection (RGFS) layer. We use semantic segmentation as a prior to “guide” this filter selection. To the best of our knowledge, our work is the first comprehensive study on truly multimodal material segmentation. We believe our work opens new avenues of practical use of material information in safety critical applications. Yupeng Liang, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino |
CVPR | 4 |
| 2022 | Dynamic 3D Gaze from Afar: Deep Gaze Estimation from Temporal Eye-Head-Body CoordinationabstractWe introduce a novel method and dataset for 3D gaze estimation of a freely moving person from a distance, typically in surveillance views. Eyes cannot be clearly seen in such cases due to occlusion and lacking resolution. Existing gaze estimation methods suffer or fall back to approximating gaze with head pose as they primarily rely on clear, close-up views of the eyes. Our key idea is to instead leverage the intrinsic gaze, head, and body coordination of people. Our method formulates gaze estimation as Bayesian prediction given temporal estimates of head and body orientations which can be reliably estimated from a far. We model the head and body orientation likelihoods and the conditional prior of gaze direction on those with separate neural networks which are then cascaded to output the 3D gaze direction. We introduce an extensive new dataset that consists of surveillance videos annotated with 3D gaze directions captured in 5 indoor and outdoor scenes. Experimental results on this and other datasets validate the accuracy of our method and demonstrate that gaze can be accurately estimated from a typical surveillance distance even when the person's face is not visible to the camera. Soma Nonaka, Shohei Nobuhara, Ko Nishino |
CVPR | 3 |
| 2022 | Dyadic Human Relation RecognitionabstractWe introduce a new dataset and novel method for dyadic human relation (DHR) recognition. DHR recognition is a new image understanding task that concerns the recognition of the type (i.e., verb) and roles of a two-person interaction. Unlike past human action and dyadic interaction detection, our goal is to extract richer information regarding the roles of actors, i.e., which person is acting on whom. For this, we introduce the DHR-WebImages dataset which consists of a total of 22,046 images of 51 verb classes of DHR with per-image annotation of the verb and roles. We derive a novel model for DHR recognition by consolidating information from three streams corresponding to the visual, pose, and spatial relation features of the two people. Experimental results show that our model achieves accurate DHR recognition robust to severe occlusions and lacking visibility of actors. We believe these results will lead to richer visual information extraction. Kaen Kogashi, Shohei Nobuhara, Ko Nishino |
ICME | 3 |
| 2022 | Invertible Neural BRDF for Object Inverse RenderingabstractWe introduce a novel neural network-based BRDF model and a Bayesian framework for object inverse rendering, i.e., joint estimation of reflectance and natural illumination from a single image of an object of known geometry. The BRDF is expressed with an invertible neural network, namely, normalizing flow, which provides the expressive power of a high-dimensional representation, computational simplicity of a compact analytical model, and physical plausibility of a real-world BRDF. We extract the latent space of real-world reflectance by conditioning this model, which directly results in a strong reflectance prior. We refer to this model as the invertible neural BRDF model (iBRDF). We also devise a deep illumination prior by leveraging the structural bias of deep neural networks. By integrating this novel BRDF model and reflectance and illumination priors in a MAP estimation formulation, we show that this joint estimation can be computed efficiently with stochastic gradient descent. We experimentally validate the accuracy of the invertible neural BRDF model on a large number of measured data and demonstrate its use in object inverse rendering on a number of synthetic and real images. The results show new ways in which deep neural networks can help solve challenging radiometric inverse problems. Shohei Nobuhara, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Surface Normals and Shape From WaterabstractIn this paper, we introduce a novel method for reconstructing surface normals and depth of dynamic objects in water. Past shape recovery methods have leveraged various visual cues for estimating shape (e.g., depth) or surface normals. Methods that estimate both compute one from the other. We show that these two geometric surface properties can be simultaneously recovered for each pixel when the object is observed underwater. Our key idea is to leverage multi-wavelength near-infrared light absorption along different underwater light paths in conjunction with surface shading. Our method can handle both Lambertian and non-Lambertian surfaces. We derive a principled theory for this surface normals and shape from water method and a practical calibration method for determining its imaging parameters values. By construction, the method can be implemented as a one-shot imaging system. We prototype both an off-line and a video-rate imaging system and demonstrate the effectiveness of the method on a number of real-world static and dynamic objects. The results show that the method can recover intricate surface features that are otherwise inaccessible. Meng-Yu Kuo, Satoshi Murai, Ryo Kawahara, Shohei Nobuhara, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Estimation of Wetness and Color from a Single Multispectral ImageabstractRecognizing wet surfaces and their degrees of wetness is essential for many computer vision applications. Surface wetness can inform us slippery spots on a road to autonomous vehicles, muddy areas of a trail to humanoid robots, and the freshness of groceries to us. The fact that surfaces darken when wet, i.e., monochromatic appearance change, has been modeled to recognize wet surfaces in the past. In this paper, we show that color change, particularly in its spectral behavior, carries rich information about surface wetness. We first derive an analytical spectral appearance model of wet surfaces that expresses the characteristic spectral sharpening due to multiple scattering and absorption in the surface. We present a novel method for estimating key parameters of this spectral appearance model, which enables the recovery of the original surface color and the degree of wetness from a single multispectral image. Applied to a multispectral image, the method estimates the spatial map of wetness together with the dry spectral distribution of the surface. To our knowledge, this is the first work to model and leverage the spectral characteristics of wet surfaces to decipher its appearance. We conduct comprehensive experimental validation with a number of wet real surfaces. The results demonstrate the accuracy of our model and the effectiveness of our method for surface wetness and color estimation. Hiroki Okawa, Mihoko Shimano, Yuta Asano, Ryoma Bise, Ko Nishino, Imari Sato |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Differential Viewpoints for Ground Terrain Material RecognitionabstractComputational surface modeling that underlies material recognition has transitioned from reflectance modeling using in-lab controlled radiometric measurements to image-based representations based on internet-mined single-view images captured in the scene. We take a middle-ground approach for material recognition that takes advantage of both rich radiometric cues and flexible image capture. A key concept is differential angular imaging, where small angular variations in image capture enables angular-gradient features for an enhanced appearance representation that improves recognition. We build a large-scale material database, Ground Terrain in Outdoor Scenes (GTOS) database, to support ground terrain recognition for applications such as autonomous driving and robot navigation. The database consists of over 30,000 images covering 40 classes of outdoor ground terrain under varying weather and lighting conditions. We develop a novel approach for material recognition called texture-encoded angular network (TEAN) that combines deep encoding pooling of RGB information and differential angular images for angular-gradient features to fully leverage this large dataset. With this novel network architecture, we extract characteristics of materials encoded in the angular and spatial gradients of their appearance. Our results show that TEAN achieves recognition performance that surpasses single view performance and standard (non-differential/large-angle sampling) multiview performance. Jia Xue, Hang Zhang 0005, Ko Nishino, Kristin J. Dana |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | View Birdification in the Crowd: Ground-Plane Localization from Perceived Movements
Mai Nishimura, Shohei Nobuhara, Ko Nishino |
BMVC | 3 |
| 2021 | Polarimetric Normal StereoabstractWe introduce a novel method for recovering per-pixel surface normals from a pair of polarization cameras. Unlike past methods that use polarimetric observations as auxiliary features for correspondence matching, we fully integrate them in cost volume construction and filtering to directly recover per-pixel surface normals, not as byproducts of recovered disparities. Our key idea is to introduce a polarimetric cost volume of distance defined on the polarimetric observations and the polarization state computed from the surface normal. We adapt a belief propagation algorithm to filter this cost volume. The filtering algorithm simultaneously estimates the disparities and surface normals as separate entities, while effectively denoising the original noisy polarimetric observations of a quad-Bayer polarization camera. In addition, in contrast to past methods, we model polarimetric light reflection of mesoscopic surface roughness, which is essential to account for its illumination-dependency. We demonstrate the effectiveness of our method on a number of complex, real objects. Our method offers a simple and detailed 3D sensing capability for complex, non-Lambertian surfaces. Yoshiki Fukao, Ryo Kawahara, Shohei Nobuhara, Ko Nishino |
CVPR | 4 |
| 2021 | Shape From Sky: Polarimetric Normal Recovery Under the SkyabstractThe sky exhibits a unique spatial polarization pattern by scattering the unpolarized sun light. Just like insects use this unique angular pattern to navigate, we use it to map pixels to directions on the sky. That is, we show that the unique polarization pattern encoded in the polarimetric appearance of an object captured under the sky can be decoded to reveal the surface normal at each pixel. We derive a polarimetric reflection model of a diffuse plus mirror surface lit by the sun and a clear sky. This model is used to recover the per-pixel surface normal of an object from a single polarimetric image or from multiple polarimetric images captured under the sky at different times of the day. We experimentally evaluate the accuracy of our shape-from-sky method on a number of real objects of different surface compositions. The results clearly show that this passive approach to fine-geometry recovery that fully leverages the unique illumination made by nature is a viable option for 3D sensing. With the advent of quad-Bayer polarization chips, we believe the implications of our method span a wide range of domains. Tomoki Ichikawa, Matthew Purri, Ryo Kawahara, Shohei Nobuhara, Kristin J. Dana, Ko Nishino |
CVPR | 6 |
| 2021 | Human-object interaction detection with missing objects
Kaen Kogashi, Yang Wu 0001, Shohei Nobuhara, Ko Nishino |
Image Vis. Comput. | 4 |
| 2021 | Depth Sensing by Near-Infrared Light Absorption in WaterabstractThis paper introduces a novel depth recovery method based on light absorption in water. Water absorbs light at almost all wavelengths whose absorption coefficient is related to the wavelength. Based on the Beer-Lambert model, we introduce a bispectral depth recovery method that leverages the light absorption difference between two near-infrared wavelengths captured with a distant point source and orthographic cameras. Through extensive analysis, we show that accurate depth can be recovered irrespective of the surface texture and reflectance, and introduce algorithms to correct for nonidealities of a practical implementation including tilted light source and camera placement, nonideal bandpass filters and the perspective effect of the camera with a diverging point light source. We construct a coaxial bispectral depth imaging system using low-cost off-the-shelf hardware and demonstrate its use for recovering the shapes of complex and dynamic objects in water. We also present a trispectral variant to further improve robustness to extremely challenging surface reflectance. Experimental results validate the theory and practical implementation of this novel depth recovery paradigm, which we refer to as shape from water. Yuta Asano, Yinqiang Zheng, Ko Nishino, Imari Sato |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Non-Rigid Shape From WaterabstractWe introduce a novel 3D sensing method for recovering a consistent, dense 3D shape of a dynamic, non-rigid object in water. The method reconstructs a complete (or fuller) 3D surface of the target object in a canonical frame (e.g., rest shape) as it freely deforms and moves between frames by estimating underwater 3D scene flow and using it to integrate per-frame depth estimates recovered from two near-infrared observations. The reconstructed shape is refined in the course of this global non-rigid shape recovery by leveraging both geometric and radiometric constraints. We implement our method with a single camera and a light source without the orthographic assumption on either by deriving a practical calibration method that estimates the point source position with respect to the camera. Our reconstruction method also accounts for scattering by water. We prototype a video-rate imaging system and show 3D shape reconstruction results on a number of real-world static, deformable, and dynamic objects and creatures in real-world water. The results demonstrate the effectiveness of the method in recovering complete shapes of complex, non-rigid objects in water, which opens new avenues of application for underwater 3D sensing in the sub-meter range. Meng-Yu Kuo, Ryo Kawahara, Shohei Nobuhara, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Guest Editorial: Introduction to the Special Section on Computational PhotographyabstractThe papers in this special section focus on computational photography. The past year has been significant in many ways. For the scientific community of computational photography, we have a hybrid in-person conference, one of the first in the vision/optics/graphics community post-pandemic. Further, we expanded our community to better include the physical optics community. This move was made consciously to strengthen and expand the span of computational photography and make these different, yet closely related communities, have a common venue to share ideas. This expansion we believe has significantly enriched the papers submitted to this special issue, through the IEEE International Conference on Computational Photography (ICCP’2021). Yoav Y. Schechner, Kavita Bala, Ori Katz, Kalyan Sunkavalli, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Video Region Annotation with Sparse Bounding Boxes
Yuzheng Xu, Yang Wu 0001, Nur Sabrina binti Zuraimi, Shohei Nobuhara, Ko Nishino |
BMVC | 5 |
| 2020 | 3D-GMNet: Single-View 3D Shape Recovery as A Gaussian Mixture
Kohei Yamashita 0001, Shohei Nobuhara, Ko Nishino |
BMVC | 3 |
| 2020 | Invertible Neural BRDF for Object Inverse Rendering
Shohei Nobuhara, Ko Nishino |
ECCV (5) | 3 |
| 2020 | Appearance and Shape from Water Reflection
Ryo Kawahara, Meng-Yu Kuo, Shohei Nobuhara, Ko Nishino |
WACV | 4 |
| 2020 | Guest Editors' Introduction to the Special Issue on RGB-D Vision: Methods and ApplicationsabstractThe twenty-six papers in this special issue focus on Red Blue Green (RBG)-D vision, an emerging research topic in computer vision, with a number of applications in robotics, entertainment, biometrics and multimedia. Compared to 2D images and 3D data (including depth images, point clouds and meshes), RGB-D images represent both the photometric and geometric information of a scene. Moreover, low-cost consumer depth cameras (e.g., Microsoft Kinect v2, Intel Realsense, Orbbec Astra) can enable realtime applications due to their high acquisition frame-rate. In the last few years, a large number of RGB-D datasets have also been publicly released to tackle various vision tasks. Although remarkable progress has been achieved, several critical problems still remain open. The aim of this special issue is to stimulate researchers from different fields to present their state-of-the-art work, and to provide a cross-fertilization ground for discussions on the next steps in this important research area. Mohammed Bennamoun, Yulan Guo, Federico Tombari, Kamal Youcef-Toumi, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Recognizing Material Properties from ImagesabstractHumans implicitly rely on the properties of materials to guide our interactions. Grasping smooth materials, for example, requires more care than rough ones. We may even visually infer non-visual properties (e.g., softness is a physical material property). We refer to visually-recognizable material properties as visual material attributes. Recognizing these attributes in images can provide valuable information for scene understanding and material recognition. Unlike typical object and scene attributes, however, visual material attributes are local (i.e., "fuzziness" does not have a shape). Given full supervision, we may accurately recognize such attributes from purely local information (small image patches). Obtaining consistent full supervision at scale, however, is challenging. To solve this problem, we probe the human visual perception of materials. By asking simple yes/no questions comparing pairs of image patches, we obtain the weak supervision required to build a set of classifiers for attributes that, while unnamed, function similarly to the attributes with which we describe materials. Furthermore, we integrate this method in the end-to-end learning of a CNN that simultaneously recognizes materials and their visual attributes. Experiments show that visual material attributes serve as both a useful representation for known material categories and as a basis for transfer learning. Gabriel Schwartz, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Surface Normals and Shape From WaterabstractIn this paper, we introduce a novel method for reconstructing surface normals and depth of dynamic objects in water. Past shape recovery methods have leveraged various visual cues for estimating shape (e.g., depth) or surface normals. Methods that estimate both compute one from the other. We show that these two geometric surface properties can be simultaneously recovered for each pixel when the object is observed underwater. Our key idea is to leverage multi-wavelength near-infrared light absorption along different underwater light paths in conjunction with surface shading. We derive a principled theory for this surface normals and shape from water method and a practical calibration method for determining its imaging parameters values. By construction, the method can be implemented as a one-shot imaging system. We prototype both an off-line and a video-rate imaging system and demonstrate the effectiveness of the method on a number of real-world static and dynamic objects. The results show that the method can recover intricate surface features that are otherwise inaccessible. Satoshi Murai, Meng-Yu Kuo, Ryo Kawahara, Shohei Nobuhara, Ko Nishino |
ICCV | 5 |
| 2019 | Reflectance and Shape Estimation with a Light Field Camera Under Natural Illumination
Trung Ngo Thanh, Hajime Nagahara, Ko Nishino, Rin-Ichiro Taniguchi, Yasushi Yagi |
Int. J. Comput. Vis. | 3 |
| 2018 | A Data-Driven Approach for Direct and Global Component Separation from a Single Image
Shijie Nie, Lin Gu 0003, Art Subpa-Asa, Ilyes Kacher, Ko Nishino, Imari Sato |
ACCV (6) | 5 |
| 2018 | Variable Ring Light Imaging: Capturing Transient Subsurface Scattering with an Ordinary Camera
Ko Nishino, Art Subpa-Asa, Yuta Asano, Mihoko Shimano, Imari Sato |
ECCV (11) | 1 |
| 2018 | Editorial for ACCV'16 award papers
Shang-Hong Lai, Vincent Lepetit, Ko Nishino, Yoichi Sato 0001 |
Comput. Vis. Image Underst. | 3 |
| 2017 | Reflectance and Shape Estimation with a Light Field Camera under Natural Illumination
Trung Ngo Thanh, Hajime Nagahara, Ko Nishino, Rin-Ichiro Taniguchi, Yasushi Yagi |
BMVC | 3 |
| 2017 | Wetness and Color from a Single Multispectral ImageabstractVisual recognition of wet surfaces and their degrees of wetness is important for many computer vision applications. It can inform slippery spots on a road to autonomous vehicles, muddy areas of a trail to humanoid robots, and the freshness of groceries to us. In the past, monochromatic appearance change, the fact that surfaces darken when wet, has been modeled to recognize wet surfaces. In this paper, we show that color change, particularly in its spectral behavior, carries rich information about a wet surface. We derive an analytical spectral appearance model of wet surfaces that expresses the characteristic spectral sharpening due to multiple scattering and absorption in the surface. We derive a novel method for estimating key parameters of this spectral appearance model, which enables the recovery of the original surface color and the degree of wetness from a single observation. Applied to a multispectral image, the method estimates the spatial map of wetness together with the dry spectral distribution of the surface. To our knowledge, this work is the first to model and leverage the spectral characteristics of wet surfaces to revert its appearance. We conduct comprehensive experimental validation with a number of wet real surfaces. The results demonstrate the accuracy of our model and the effectiveness of our method for surface wetness and color estimation. Mihoko Shimano, Hiroki Okawa, Yuta Asano, Ryoma Bise, Ko Nishino, Imari Sato |
CVPR | 5 |
| 2017 | Differential Angular Imaging for Material RecognitionabstractMaterial recognition for real-world outdoor surfaces has become increasingly important for computer vision to support its operation in the wild. Computational surface modeling that underlies material recognition has transitioned from reflectance modeling using in-lab controlled radiometric measurements to image-based representations based on internet-mined images of materials captured in the scene. We propose to take a middle-ground approach for material recognition that takes advantage of both rich radiometric cues and flexible image capture. We realize this by developing a framework for differential angular imaging, where small angular variations in image capture provide an enhanced appearance representation and significant recognition improvement. We build a large-scale material database, Ground Terrain in Outdoor Scenes (GTOS) database, geared towards real use for autonomous agents. The database consists of over 30,000 images covering 40 classes of outdoor ground terrain under varying weather and lighting conditions. We develop a novel approach for material recognition called a Differential Angular Imaging Network (DAIN) to fully leverage this large dataset. With this novel network architecture, we extract characteristics of materials encoded in the angular and spatial gradients of their appearance. Our results show that DAIN achieves recognition performance that surpasses single view or coarsely quantized multiview images. These results demonstrate the effectiveness of differential angular imaging as a means for flexible, in-place material recognition. Jia Xue, Hang Zhang 0005, Kristin J. Dana, Ko Nishino |
CVPR | 4 |
| 2016 | Radiometric Scene Decomposition: Scene Reflectance, Illumination, and Geometry from RGB-D ImagesabstractRecovering the radiometric properties of a scene (i.e., the reflectance, illumination, and geometry) is a long-sought ability of computer vision that can provide invaluable information for a wide range of applications. Deciphering the radiometric ingredients from the appearance of a real-world scene, as opposed to a single isolated object, is particularly challenging as it generally consists of various objects with different material compositions exhibiting complex reflectance and light interactions that are also part of the illumination. We introduce the first method for radiometric decomposition of real-world scenes that handles those intricacies. We use RGB-D images to bootstrap geometry recovery and simultaneously recover the complex reflectance and natural illumination while refining the noisy initial geometry and segmenting the scene into different material regions. Most important, we handle real-world scenes consisting of multiple objects of unknown materials, which necessitates the modeling of spatially-varying complex reflectance, natural illumination, texture, interreflection and shadows. We systematically evaluate the effectiveness of our method on synthetic scenes and demonstrate its application to real-world scenes. The results show that rich radiometric information can be recovered from RGB-D images and demonstrate a new role RGB-D sensors can play for general scene understanding tasks. Stephen Lombardi, Ko Nishino |
3DV | 2 |
| 2016 | Shape from Water: Bispectral Light Absorption for Depth Recovery
Yuta Asano, Yinqiang Zheng, Ko Nishino, Imari Sato |
ECCV (6) | 3 |
| 2016 | Friction from Reflectance: Deep Reflectance Codes for Predicting Physical Surface Properties from One-Shot In-Field Reflectance
Hang Zhang 0005, Kristin J. Dana, Ko Nishino |
ECCV (4) | 3 |
| 2016 | Reflectance and Illumination Recovery in the WildabstractThe appearance of an object in an image encodes invaluable information about that object and the surrounding scene. Inferring object reflectance and scene illumination from an image would help us decode this information: reflectance can reveal important properties about the materials composing an object; the illumination can tell us, for instance, whether the scene is indoors or outdoors. Recovering reflectance and illumination from a single image in the real world, however, is a difficult task. Real scenes illuminate objects from every visible direction and real objects vary greatly in reflectance behavior. In addition, the image formation process introduces ambiguities, like color constancy, that make reversing the process ill-posed. To address this problem, we propose a Bayesian framework for joint reflectance and illumination inference in the real world. We develop a reflectance model and priors that precisely capture the space of real-world object reflectance and a flexible illumination model that can represent real-world illumination with priors that combat the deleterious effects of image formation. We analyze the performance of our approach on a set of synthetic data and demonstrate results on real-world scenes. These contributions enable reliable reflectance and illumination inference in the real world. Stephen Lombardi, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Shape and Reflectance Estimation in the WildabstractOur world is full of objects with complex reflectances situated in rich illumination environments. Though stunning, the diversity of appearance that arises from this complexity is also daunting. For this reason, past work on geometry recovery has tried to frame the problem into simplistic models of reflectance (such as Lambertian, mirrored, or dichromatic) or illumination (one or more distant point light sources). In this work, we directly tackle the problem of joint reflectance and geometry estimation under known but uncontrolled natural illumination by fully exploiting the surface orientation cues that become embedded in the appearance of the object. Intuitively, salient scene features (such as the sun or stained glass windows) act analogously to the point light sources of traditional geometry estimation frameworks by strongly constraining the possible orientations of the surface patches reflecting them. By jointly estimating the reflectance of the object, which modulates the illumination, the appearance of a surface patch can be used to derive a nonparametric distribution of its possible orientations. If only a single image exists, these strongly constrained surface patches may then be used to anchor the geometry estimation and give context to the less-descriptive regions. When multiple images exist, the distribution of possible surface orientations becomes tighter as additional context is given, though integrating the separate views poses additional challenges. In this paper we introduce two methods, one for the single image case, and another for the case of multiple images. The effectiveness of our methods is evaluated extensively on synthetic and real-world data sets that span the wide range of real-world environments and reflectances that lies between the extremes that have been the focus of past work. Geoffrey Oxholm, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Automatically discovering local visual material attributesabstractShape cues play an important role in computer vision, but shape is not the only information available in images. Materials, such as fabric and plastic, are discernible in images even when shapes, such as those of an object, are not. We argue that it would be ideal to recognize materials without relying on object cues such as shape. This would allow us to use materials as a context for other vision tasks, such as object recognition. Humans are intuitively able to find visual cues that describe materials. Previous frameworks attempt to recognize these cues (as visual material traits) using fully-supervised learning. This requirement is not feasible when multiple annotators and large quantities of images are involved. In this paper, we derive a framework that allows us to discover locally-recognizable material attributes from crowdsourced perceptual material distances. We show that the attributes we discover do in fact separate material categories. Our learned attributes exhibit the same desirable properties as material traits, despite the fact that they are discovered using only partial supervision. Gabriel Schwartz, Ko Nishino |
CVPR | 2 |
| 2015 | Reflectance hashing for material recognitionabstractWe introduce a novel method for using reflectance to identify materials. Reflectance offers a unique signature of the material but is challenging to measure and use for recognizing materials due to its high-dimensionality. In this work, one-shot reflectance of a material surface which we refer to as a reflectance disk is capturing using a unique optical camera. The pixel coordinates of these reflectance disks correspond to the surface viewing angles. The reflectance has class-specific stucture and angular gradients computed in this reflectance space reveal the material class. These reflectance disks encode discriminative information for efficient and accurate material recognition. We introduce a framework called reflectance hashing that models the reflectance disks with dictionary learning and binary hashing. We demonstrate the effectiveness of reflectance hashing for material recognition with a number of real-world materials. Hang Zhang 0005, Kristin J. Dana, Ko Nishino |
CVPR | 3 |
| 2014 | Multiview Shape and Reflectance from Natural IlluminationabstractThe world is full of objects with complex reflectances, situated in complex illumination environments. Past work on full 3D geometry recovery, however, has tried to handle this complexity by framing it into simplistic models of reflectance (Lambetian, mirrored, or diffuse plus specular) or illumination (one or more point light sources). Though there has been some recent progress in directly utilizing such complexities for recovering a single view geometry, it is not clear how such single-view methods can be extended to reconstruct the full geometry. To this end, we derive a probabilistic geometry estimation method that fully exploits the rich signal embedded in complex appearance. Though each observation provides partial and unreliable information, we show how to estimate the reflectance responsible for the diverse appearance, and unite the orientation cues embedded in each observation to reconstruct the underlying geometry. We demonstrate the effectiveness of our method on synthetic and real-world objects. The results show that our method performs accurately across a wide range of real-world environments and reflectances that lies between the extremes that have been the focus of past work. Geoffrey Oxholm, Ko Nishino |
CVPR | 2 |
| 2013 | Two-Point Gait: Decoupling Gait from Body ShapeabstractHuman gait modeling (e.g., for person identification) largely relies on image-based representations that muddle gait with body shape. Silhouettes, for instance, inherently entangle body shape and gait. For gait analysis and recognition, decoupling these two factors is desirable. Most important, once decoupled, they can be combined for the task at hand, but not if left entangled in the first place. In this paper, we introduce Two-Point Gait, a gait representation that encodes the limb motions regardless of the body shape. Two-Point Gait is directly computed on the image sequence based on the two point statistics of optical flow fields. We demonstrate its use for exploring the space of human gait and gait recognition under large clothing variation. The results show that we can achieve state-of-the-art person recognition accuracy on a challenging dataset. Stephen Lombardi, Ko Nishino, Yasushi Makihara, Yasushi Yagi |
ICCV | 2 |
| 2012 | Single image multimaterial estimationabstractEstimating the reflectance and illumination from a single image becomes particularly challenging when the object surface consists of multiple materials. The key difficulty lies in recovering the reflectance from sparse angular samples while correctly assigning them to different materials. We tackle this problem by extracting and fully leveraging reflectance priors. The idea is to strongly constrain the possible solutions so that the recovered reflectance conform with those of real-world materials. We achieve this by modeling the parameter space of a directional statistics BRDF model and by extracting an analytical distribution of the subspace that real-world materials span. This is used, with other priors, in a layered MRF-based formulation that models material regions and their spatially varying reflectance with continuous latent layers. The material regions and their reflectance, and the direction and strength of a single point source are jointly estimated. We demonstrate the effectiveness of the method on real and synthetic images. Stephen Lombardi, Ko Nishino |
CVPR | 2 |
| 2012 | Going with the Flow: Pedestrian Efficiency in Crowded Scenes
Louis Kratz, Ko Nishino |
ECCV (4) | 2 |
| 2012 | Reflectance and Natural Illumination from a Single Image
Stephen Lombardi, Ko Nishino |
ECCV (6) | 2 |
| 2012 | The Scale of Geometric Texture
Geoffrey Oxholm, Prabin Bariya, Ko Nishino |
ECCV (1) | 3 |
| 2012 | Shape and Reflectance from Natural Illumination
Geoffrey Oxholm, Ko Nishino |
ECCV (1) | 2 |
| 2012 | 3D Geometric Scale Variability in Range Images: Features and Descriptors
Prabin Bariya, John Novatnack, Gabriel Schwartz, Ko Nishino |
Int. J. Comput. Vis. | 4 |
| 2012 | Bayesian Defogging
Ko Nishino, Louis Kratz, Stephen Lombardi |
Int. J. Comput. Vis. | 1 |
| 2012 | Tracking Pedestrians Using Local Spatio-Temporal Motion Patterns in Extremely Crowded ScenesabstractTracking pedestrians is a vital component of many computer vision applications, including surveillance, scene understanding, and behavior analysis. Videos of crowded scenes present significant challenges to tracking due to the large number of pedestrians and the frequent partial occlusions that they produce. The movement of each pedestrian, however, contributes to the overall crowd motion (i.e., the collective motions of the scene's constituents over the entire video) that exhibits an underlying spatially and temporally varying structured pattern. In this paper, we present a novel Bayesian framework for tracking pedestrians in videos of crowded scenes using a space-time model of the crowd motion. We represent the crowd motion with a collection of hidden Markov models trained on local spatio-temporal motion patterns, i.e., the motion patterns exhibited by pedestrians as they move through local space-time regions of the video. Using this unique representation, we predict the next local spatio-temporal motion pattern a tracked pedestrian will exhibit based on the observed frames of the video. We then use this prediction as a prior for tracking the movement of an individual in videos of extremely crowded scenes. We show that our approach of leveraging the crowd motion enables tracking in videos of complex scenes that present unique difficulty to other approaches. Louis Kratz, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Locally rigid globally non-rigid surface registrationabstractWe present a novel non-rigid surface registration method that achieves high accuracy and matches characteristic features without manual intervention. The key insight is to consider the entire shape as a collection of local structures that individually undergo rigid transformations to collectively deform the global structure. We realize this locally rigid but globally non-rigid surface registration with a newly derived dual-grid Free-form Deformation (FFD) framework. We first represent the source and target shapes with their signed distance fields (SDF). We then superimpose a sampling grid onto a conventional FFD grid that is dual to the control points. Each control point is then iteratively translated by a rigid transformation that minimizes the difference between two SDFs within the corresponding sampling region. The translated control points then interpolate the embedding space within the FFD grid and determine the overall deformation. The experimental results clearly demonstrate that our method is capable of overcoming the difficulty of preserving and matching local features. Kent Fujiwara, Ko Nishino, Jun Takamatsu, Bo Zheng 0001, Katsushi Ikeuchi |
ICCV | 2 |
| 2011 | Aligning surfaces without aligning surfacesabstractWe introduce a novel method for matching and aligning 3D surfaces that do not have any overlapping surface information. When two matching surfaces do not overlap, all that remains in common between them is a thin strip along their borders. Aligning such fragments is challenging but crucial for various applications, such as reassembly of thin-shell ceramics from their broken pieces. Past work approach this problem by heavily relying on simplistic assumptions about the shape of the object, or its texture. Our method makes no such assumptions; instead, we leverage the geometric and photometric similarity of the matching surfaces along the break-line. We first encode the shape and color of the boundary contour of each fragment at various scales in a novel 2D representation. Reformulating contour matching as 2D image registration based on these scale-space images enables efficient and accurate break-line matching. We then align the fragments by estimating the rotation around the break-line through maximizing the geometric continuity across it with a least-squares minimization. We evaluate our method on real-word colonial artifacts recently excavated in Philadelphia, Pennsylvania. Our system dramatically increases the ease and efficiency at which users reassemble artifacts as we demonstrate on three different vessels. Geoffrey Oxholm, Ko Nishino |
WACV | 2 |
| 2010 | Scale-hierarchical 3D object recognition in cluttered scenesabstract3D object recognition in scenes with occlusion and clutter is a difficult task. In this paper, we introduce a method that exploits the geometric scale-variability to aid in this task. Our key insight is to leverage the rich discriminative information provided by the scale variation of local geometric structures to constrain the massive search space of potential correspondences between model and scene points. In particular, we exploit the geometric scale variability in the form of the intrinsic geometric scale of each computed feature, the hierarchy induced within the set of these intrinsic geometric scales, and the discriminative power of the local scale-dependent/invariant 3D shape descriptors. The method exploits the added information in a hierarchical coarse-to-fine manner that lets it cull the space of all potential correspondences effectively. We experimentally evaluate the accuracy of our method on an extensive set of real scenes with varying amounts of partial occlusion and achieve recognition rates higher than the state-of-the-art. Furthermore, for the first time we systematically demonstrate the method's ability to accurately localize objects despite changes in their global scales. Prabin Bariya, Ko Nishino |
CVPR | 2 |
| 2010 | Tracking with local spatio-temporal motion patterns in extremely crowded scenesabstractTracking individuals in extremely crowded scenes is a challenging task, primarily due to the motion and appearance variability produced by the large number of people within the scene. The individual pedestrians, however, collectively form a crowd that exhibits a spatially and temporally structured pattern within the scene. In this paper, we extract this steady-state but dynamically evolving motion of the crowd and leverage it to track individuals in videos of the same scene. We capture the spatial and temporal variations in the crowd's motion by training a collection of hidden Markov models on the motion patterns within the scene. Using these models, we predict the local spatio-temporal motion patterns that describe the pedestrian movement at each space-time location in the video. Based on these predictions, we hypothesize the target's movement between frames as it travels through the local space-time volume. In addition, we robustly model the individual's unique motion and appearance to discern them from surrounding pedestrians. The results show that we may track individuals in scenes that present extreme difficulty to previous techniques. Louis Kratz, Ko Nishino |
CVPR | 2 |
| 2010 | Membrane Nonrigid Image Registration
Geoffrey Oxholm, Ko Nishino |
ECCV (2) | 2 |
| 2009 | Illumination and spatially varying specular reflectance from a single viewabstractEstimating the illumination and the reflectance properties of an object surface from a sparse set of images is an important but inherently ill-posed problem. The problem becomes even harder if we wish to account for the spatial variation of material properties on the surface. In this paper, we derive a novel method for estimating the spatially varying specular reflectance properties, of a surface of known geometry, as well as the illumination distribution from a specular-only image, for instance, captured using polarization to separate reflection components. Unlike previous work, we do not assume the illumination to be a single point light source. We model specular reflection with a spherical statistical distribution and encode the spatial variation with radial basis functions of its parameters. This allows us to formulate the simultaneous estimation of spatially varying specular reflectance and illumination as a sound probabilistic inference problem, in particular, using Csiszar's I-divergence measure. To solve it, we derive an iterative algorithm similar to expectation maximization. We demonstrate the effectiveness of the method on synthetic and real-world scenes. Kenji Hara, Ko Nishino |
CVPR | 2 |
| 2009 | Anomaly detection in extremely crowded scenes using spatio-temporal motion pattern modelsabstractExtremely crowded scenes present unique challenges to video analysis that cannot be addressed with conventional approaches. We present a novel statistical framework for modeling the local spatio-temporal motion pattern behavior of extremely crowded scenes. Our key insight is to exploit the dense activity of the crowded scene by modeling the rich motion patterns in local areas, effectively capturing the underlying intrinsic structure they form in the video. In other words, we model the motion variation of local space-time volumes and their spatial-temporal statistical behaviors to characterize the overall behavior of the scene. We demonstrate that by capturing the steady-state motion behavior with these spatio-temporal motion pattern models, we can naturally detect unusual activity as statistical deviations. Our experiments show that local spatio-temporal motion pattern modeling offers promising results in real-world scenes with complex activities that are hard for even human observers to analyze. Louis Kratz, Ko Nishino |
CVPR | 2 |
| 2009 | Factorizing Scene Albedo and Depth from a Single Foggy ImageabstractAtmospheric conditions induced by suspended particles, such as fog and haze, severely degrade image quality. Restoring the true scene colors (clear day image) from a single image of a weather-degraded scene remains a challenging task due to the inherent ambiguity between scene albedo and depth. In this paper, we introduce a novel probabilistic method that fully leverages natural statistics of both the albedo and depth of the scene to resolve this ambiguity. Our key idea is to model the image with a factorial Markov random field in which the. scene albedo and depth are. two statistically independent latent layers. We. show that we may exploit natural image and depth statistics as priors on these hidden layers and factorize a single foggy image via a canonical Expectation Maximization algorithm with alternating minimization. Experimental results show that the proposed method achieves more accurate restoration compared to state-of-the-art methods that focus on only recovering scene albedo or depth individually. Louis Kratz, Ko Nishino |
ICCV | 2 |
| 2009 | Directional statistics BRDF modelabstractWe introduce a novel parametric BRDF model that can accurately encode a wide variety of real-world isotropic BRDFs with a small number of parameters. The key observation we make is that a BRDF may be viewed as a statistical distribution on a unit hemisphere. We derive a novel directional statistics distribution, which we refer to as the hemispherical exponential power distribution, and model an isotropic BRDF with a mixture of it. The novel directional statistics BRDF model allows us to derive a canonical probabilistic method for estimating its parameters including the number of components. We show that the model captures the full spectrum of real-world isotropic BRDFs with accuracy comparable to non-parametric models but with a much more compact representation. We also experimentally show that the model achieves better accuracy with less measurements compared with such non-parametric models. We further demonstrate the advantages of the novel BRDF model by showing its use for reflection component separation and for exploring the space of isotropic BRDFs. Ko Nishino |
ICCV | 1 |
| 2008 | Scale-Dependent/Invariant Local 3D Shape Descriptors for Fully Automatic Registration of Multiple Sets of Range Images
John Novatnack, Ko Nishino |
ECCV (3) | 2 |
| 2008 | Canonical subsets of image features
Trip Denton, Ali Shokoufandeh, John Novatnack, Ko Nishino |
Comput. Vis. Image Underst. | 4 |
| 2008 | Mixture of Spherical Distributions for Single-View RelightingabstractWe present a method for simultaneously estimating the illumination of a scene and the reflectance property of an object from single view images - a single image or a small number of images taken from the same viewpoint. We assume that the illumination consists of multiple point light sources and the shape of the object is known. First, we represent the illumination on the surface of a unit sphere as a finite mixture of von Mises-Fisher distributions based on a novel spherical specular reflection model that well approximates the Torrance-Sparrow reflection model. Next, we estimate the parameters of this mixture model including the number of its component distributions and the standard deviation of them, which correspond to the number of light sources and the surface roughness, respectively. Finally, using these results as the initial estimates, we iteratively refine the estimates based on the original Torrance-Sparrow reflection model. The final estimates can be used to relight single-view images such as altering the intensities and directions of the individual light sources. The proposed method provides a unified framework based on directional statistics for simultaneously estimating the intensities and directions of an unknown number of light sources as well as the specular reflection parameter of the object in the scene. Kenji Hara, Ko Nishino, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Scale-Dependent 3D Geometric FeaturesabstractThree-dimensional geometric data play fundamental roles in many computer vision applications. However, their scale-dependent nature, i.e. the relative variation in the spatial extents of local geometric structures, is often overlooked. In this paper we present a comprehensive framework for exploiting this 3D geometric scale variability. Specifically, we focus on detecting scale-dependent geometric features on triangular mesh models of arbitrary topology. The key idea of our approach is to analyze the geometric scale variability of a given 3D model in the scale-space of a dense and regular 2D representation of its surface geometry encoded by the surface normals. We derive novel corner and edge detectors, as well as an automatic scale selection method, that acts upon this representation to detect salient geometric features and determine their intrinsic scales. We evaluate the effectiveness and robustness of our method on a number of models of different topology. The results show that the resulting scale-dependent geometric feature set provides a reliable basis for constructing a rich but concise representation of the geometric structure at hand. John Novatnack, Ko Nishino |
ICCV | 2 |
| 2007 | The Great Buddha Project: Digitally Archiving, Restoring, and Analyzing Cultural Heritage Objects
Katsushi Ikeuchi, Takeshi Oishi, Jun Takamatsu, Ryusuke Sagawa, Atsushi Nakazawa, Ryo Kurazume, Ko Nishino, Mawo Kamakura, Yasuhide Okamoto |
Int. J. Comput. Vis. | 7 |
| 2006 | Corneal Imaging System: Environment from Eyes
Ko Nishino, Shree K. Nayar |
Int. J. Comput. Vis. | 1 |
| 2005 | Multiple Light Sources and Reflectance Property Estimation Based on a Mixture of Spherical DistributionsabstractIn tins paper we propose a new method for simultaneously estimating the illumination of the scene and the reflectance property of an object from a single image. We assume that the illumination consists of multiple point sources and the shape of the object is known. Unlike previous methods, we will recover not only the direction and intensity of the light sources, but also the number of light sources and the specular reflection parameter of the object. First, we represent the illumination on the surface of a unit sphere as a finite mixture of von Mises-Fisher distributions by deriving a spherical specular reflection model. Next, we estimate this mixture and the number of distributions. Finally, using this result as initial estimates, we refine the estimates using the original specular reflection model. We can use the results to render the object under novel lighting conditions Kenji Hara, Ko Nishino, Katsushi Ikeuchi |
ICCV | 2 |
| 2005 | Using Eye Reflections for Face Recognition Under Varying IlluminationabstractFace recognition under varying illumination remains a challenging problem. Much progress has been made toward a solution through methods that require multiple gallery images of each subject under varying illumination. Yet for many applications, this requirement is too severe. In this paper, we propose a novel method that requires only a single gallery image per subject taken under unknown lighting. The method builds upon two contributions. We first estimate the lighting from its reflection in the eyes. This allows us to explicitly recover the illumination in the single gallery images as well as the probe image. Next, we exploit the local linearity of face appearance variation across different people. We represent the gallery images as locally linear montages of images of many different faces taken under the same lighting (bootstrap images). Then, we transfer the estimated combination of bootstrap images to synthesize each subject's face under tile probe lighting to accomplish recognition. Finally, we show through tests on the CMU PIE database that we can achieve better recognition results using our lighting estimation method and locally linear montages than the current state-of-the-art. Ko Nishino, Peter N. Belhumeur, Shree K. Nayar |
ICCV | 1 |
| 2005 | Light Source Position and Reflectance Estimation from a Single View without the Distant Illumination AssumptionabstractSeveral techniques have been developed for recovering reflectance properties of real surfaces under unknown illumination. However, in most cases, those techniques assume that the light sources are located at inifinity, which cannot be applied safely to, for example, reflectance modeling of indoor environments. In this paper, we propose two types of methods to estimate the surface reflectance property of an object, as well as the position of a light source from a single view without the distant illumination assumption, thus relaxing the conditions in the previous methods. Given a real image and a 3D geometric model of an object with specular reflection as inputs, the first method estimates the light source position by fitting to the Lambertian diffuse component, while separating the specular and diffuse components by using an iterative relaxation scheme. Our second method extends that first method by using as input a specular component image, which is acquired by analyzing multiple polarization images taken from a single view, thus removing its constraints on the diffuse reflectance property. This method simultaneously recovers the reflectance properties and the light source positions by optimizing the linearity of a log-transformed Torrance-Sparrow model. By estimating the object's reflectance property and the light source position, we can freely generate synthetic images of the target object under arbitrary lighting conditions with not only source direction modification but also source-surface distance modification. Experimental results show the accuracy of our estimation framework. Kenji Hara, Ko Nishino, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Clustered Blockwise PCA for Representing Visual DataabstractPrincipal Component Analysis (PCA) is extensively used in computer vision and image processing. Since it provides the optimal linear subspace in a least-square sense, it has been used for dimensionality reduction and subspace analysis in various domains. However, its scalability is very limited because of its inherent computational complexity. We introduce a new framework for applying PCA to visual data which takes advantage of the spatio-temporal correlation and localized frequency variations that are typically found in such data. Instead of applying PCA to the whole volume of data (complete set of images), we partition the volume into a set of blocks and apply PCA to each block. Then, we group the subspaces corresponding to the blocks and merge them together. As a result, we not only achieve greater efficiency in the resulting representation of the visual data, but also successfully scale PCA to handle large data sets. We present a thorough analysis of the computational complexity and storage benefits of our approach. We apply our algorithm to several types of videos. We show that, in addition to its storage and speed benefits, the algorithm results in a useful representation of the visual data. Ko Nishino, Shree K. Nayar, Tony Jebara |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Adaptively Merging Large-Scale Range Data with Reflectance PropertiesabstractIn this paper, we tackle the problem of geometric and photometric modeling of large intricately shaped objects. Typical target objects we consider are cultural heritage objects. When constructing models of such objects, we are faced with several important issues that have not been addressed in the past-issues that mainly arise due to the large amount of data that has to be handled. We propose two novel approaches to efficiently handle such large amounts of data: A highly adaptive algorithm for merging range images and an adaptive nearest-neighbor search to be used with the algorithm. We construct an integrated mesh model of the target object in adaptive resolution, taking into account the geometric and/or photometric attributes associated with the range images. We use surface curvature for the geometric attributes and (laser) reflectance values for the photometric attributes. This adaptive merging framework leads to a significant reduction in the necessary amount of computational resources. Furthermore, the resulting adaptive mesh models can be of great use for applications such as texture mapping, as we will briefly demonstrate. Additionally, we propose an additional test for the k-d tree nearest-neighbor search algorithm. Our approach successfully omits back-tracking, which is controlled adaptively depending on the distance to the nearest neighbor. Since the main consumption of computational cost lies in the nearest-neighbor search, the proposed algorithm leads to a significant speed-up of the whole merging process. In this paper, we present the theories and algorithms of our approaches with pseudo code and apply them to several real objects, including large-scale cultural assets. Ryusuke Sagawa, Ko Nishino, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | The World in an Eye
Ko Nishino, Shree K. Nayar |
CVPR (1) | 1 |
| 2004 | Illumination Normalization with Time-Dependent Intrinsic Images for Video SurveillanceabstractVariation in illumination conditions caused by weather, time of day, etc., makes the task difficult when building video surveillance systems of real world scenes. Especially, cast shadows produce troublesome effects, typically for object tracking from a fixed viewpoint, since it yields appearance variations of objects depending on whether they are inside or outside the shadow. In this paper, we handle such appearance variations by removing shadows in the image sequence. This can be considered as a preprocessing stage which leads to robust video surveillance. To achieve this, we propose a framework based on the idea of intrinsic images. Unlike previous methods of deriving intrinsic images, we derive time-varying reflectance images and corresponding illumination images from a sequence of images instead of assuming a single reflectance image. Using obtained illumination images, we normalize the input image sequence in terms of incident lighting distribution to eliminate shadowing effects. We also propose an illumination normalization scheme which can potentially run in real time, utilizing the illumination eigenspace, which captures the illumination variation due to weather, time of day, etc., and a shadow interpolation method based on shadow hulls. This paper describes the theory of the framework with simulation results and shows its effectiveness with object tracking results on real scene data sets. Yasuyuki Matsushita, Ko Nishino, Katsushi Ikeuchi, Masao Sakauchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Separating Reflection Components Based on Chromaticity and Noise AnalysisabstractMany algorithms in computer vision assume diffuse only reflections and deem specular reflections to be outliers. However, in the real world, the presence of specular reflections is inevitable since there are many dielectric inhomogeneous objects which have both diffuse and specular reflections. To resolve this problem, we present a method to separate the two reflection components. The method is principally based on the distribution of specular and diffuse points in a two-dimensional maximum chromaticity-intensity space. We found that, by utilizing the space and known illumination color, the problem of reflection component separation can be simplified into the problem of identifying diffuse maximum chromaticity. To be able to identify the diffuse maximum chromaticity correctly, an analysis of the noise is required since most real images suffer from it. Unlike existing methods, the proposed method can separate the reflection components robustly for any kind of surface roughness and light direction. Robby T. Tan, Ko Nishino, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Eyes for relightingabstractThe combination of the cornea of an eye and a camera viewing the eye form a catadioptric (mirror + lens) imaging system with a very wide field of view. We present a detailed analysis of the characteristics of this corneal imaging system. Anatomical studies have shown that the shape of a normal cornea (without major defects) can be approximated with an ellipsoid of fixed eccentricity and size. Using this shape model, we can determine the geometric parameters of the corneal imaging system from the image. Then, an environment map of the scene with a large field of view can be computed from the image. The environment map represents the illumination of the scene with respect to the eye. This use of an eye as a natural light probe is advantageous in many relighting scenarios. For instance, it enables us to insert virtual objects into an image such that they appear consistent with the illumination of the scene. The eye is a particularly useful probe when relighting faces. It allows us to reconstruct the geometry of a face by simply waving a light source in front of the face. Finally, in the case of an already captured image, eyes could be the only direct means for obtaining illumination information. We show how illumination computed from eyes can be used to replace a face in an image with another one. We believe that the eye not only serves as a useful tool for relighting but also makes relighting possible in situations where current approaches are hard to use. Ko Nishino, Shree K. Nayar |
ACM Trans. Graph. | 1 |
| 2003 | Illumination Normalization with Time-dependent Intrinsic Images for Video SurveillanceabstractCast shadows produce troublesome effects for video surveillance systems, typically for object tracking from a fixed viewpoint, since it yields appearance variations of objects depending on whether they are inside or outside the shadow. To robustly eliminate these shadows from image sequences as a preprocessing stage for robust video surveillance, we propose a framework based on the idea of intrinsic images. Unlike previous methods for deriving intrinsic images, we derive time-varying reflectance images and corresponding illumination images from a sequence of images. Using obtained illumination images, we normalize the input image sequence in terms of incident lighting distribution to eliminate shadow effects. We also propose an illumination normalization scheme, which can potentially run in real time, utilizing the illumination eigenspace, which captures the illumination variation due to weather, time of day etc., and a shadow interpolation method based on shadow hulls. This paper describes the theory of the framework with simulation results, and shows its effectiveness with object tracking results on real scene data sets for traffic monitoring. Yasuyuki Matsushita, Ko Nishino, Katsushi Ikeuchi, Masao Sakauchi |
CVPR (1) | 2 |
| 2003 | Illumination Chromaticity Estimation using Inverse-Intensity Chromaticity SpaceabstractExisting color constancy methods cannot handle both uniform colored surfaces and highly textured surfaces in a single integrated framework. Statistics-based methods require many surface colors, and become error prone when there are only few surface colors. In contrast, dichromatic-based methods can successfully handle uniformly colored surfaces, but cannot be applied to highly textured surfaces since they require precise color segmentation. In this paper, we present a single integrated method to estimate illumination chromaticity from single/multi-colored surfaces. Unlike the existing dichromatic-based methods, the proposed method requires only rough highlight regions, without segmenting the colors inside them. We show that, by analyzing highlights, a direct correlation between illumination chromaticity and image chromaticity can be obtained. This correlation is clearly described in "inverse-intensity chromaticity space", a new two-dimensional space we introduce. In addition, by utilizing the Hough transform and histogram analysis in this space, illumination chromaticity can be estimated robustly, even for a highly textured surface. Experimental results on real images show the effectiveness of the method. Robby T. Tan, Ko Nishino, Katsushi Ikeuchi |
CVPR (1) | 2 |
| 2003 | Determining Reflectance and Light Position from a Single Image Without Distant Illumination AssumptionabstractSeveral techniques have been developed for recovering reflectance properties of real surfaces under unknown illumination conditions. However, in most cases, those techniques assume that the light sources are located at infinity, which cannot be applied to, for example, photometric modelling of indoor environments. We propose two methods to estimate the surface reflectance property of an object, as well as the position of a light source from a single image without the distant illumination assumption. Given a color image of an object with specular reflection as an input, the first method estimates the light source position by fitting to the Lambertian diffuse component, while separating the specular and diffuse components by using an iterative relaxation scheme. Moreover, we extend the above method by using a single specular image as an input, thus removing its constraints on the diffuse reflectance property and the number of light sources. This method simultaneously recovers the reflectance properties and the light source positions by optimizing the linearity of a log-transformed Torrance-Sparrow model. By estimating the object's reflectance property and the light source position, we can freely generate synthetic images of the target object under arbitrary source directions and source-surface distances. Kenji Hara, Ko Nishino, Katsushi Ikeuchi |
ICCV | 2 |
| 2001 | Robust and Adaptive Integration of Multiple Range Images with Photometric AttributesabstractIntegration of multiple range images is important to make use of 3D data acquired from stereo systems, laser range finders, etc. We propose a new range image integration method based on volumetric representation. Unlike other volume-based integration methods, we adaptively subdivide voxels depending on the curvature of the surface to be reconstructed, providing efficient representation of the underlying geometry and efficient use of computational resources. In our range image merging framework, additional attributes, e.g., color, laser reflectance power, etc., can be taken into account as well as 3D geometric information. This ability allows us to generate 3D models preserving sharp edges around texture boundaries, thereby providing a good basis for efficient rendering and texture mapping. The overall framework is designed to be robust against noise, taking consensus carefully in both geometry and color, which could be suitable for 3D model reconstruction from noisy stereo images. In this paper, we describe the system, and present several results of applying our framework to real data. We also present some other future applications based on our framework. Ryusuke Sagawa, Ko Nishino, Katsushi Ikeuchi |
CVPR (2) | 2 |
| 2001 | Determining Reflectance Parameters and Illumination Distribution from a Sparse Set of Images for View-dependent Image Synthesis
Ko Nishino, Zhengyou Zhang, Katsushi Ikeuchi |
ICCV | 1 |
| 2001 | Parallel processing of range data mergingabstractThis paper describes a volumetric view-merging algorithm that generates a consensus surface of an object from its range images. Our original method merges a set of range images into a volumetric implicit-surface representation, which is converted to a surface mesh by using a variant of the marching-cubes algorithm. We propose a method that increases the computation and memory efficiency for computing signed distances and the method of parallel computing on a PC cluster Since our method permits a reduction in the data amount allocated in memory, the closest point is searched efficiently; this allows us to increase the number of parallel traversals and to reduce the computation time. In this paper, we describe the following two algorithms which are complementary in terms of the efficiency of CPU and memory usage: distributed allocation of range data and parallel traversal of partial octrees. By adjusting them according to the system specifications, we can build the model efficiently by a PC cluster We have implemented this system and evaluated its performance. Ryusuke Sagawa, Ko Nishino, Mark D. Wheeler, Katsushi Ikeuchi |
IROS | 2 |
| 2001 | Eigen-Texture Method: Appearance Compression and Synthesis Based on a 3D ModelabstractImage-based and model-based methods are two representative rendering methods for generating virtual images of objects from their real images. However, both methods still have several drawbacks when we attempt to apply them to mixed reality where we integrate virtual images with real background images. To overcome these difficulties, we propose a new method, which we refer to as the Eigen-Texture method. The proposed method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface generated from a sequence of range images. The Eigen-Texture method is an example of a view-dependent texturing approach which combines the advantages of image-based and model-based approaches. No reflectance analysis of the object surface is needed, while an accurate 3D geometric model facilitates integration with other scenes. The paper describes the method and reports on its implementation. Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Eigen-Texture Method: Appearance Compression Based on 3D ModelabstractImage-based and model-based methods are two representative rendering methods for generating virtual images of objects from their real images. Extensive research on these two methods has been made in CV and CG communities. However, both methods still have several drawbacks when it comes to applying them to the mixed reality where we integrate such virtual images with real background images. To overcome these difficulties, we propose a new method which we refer to as the Eigen-Texture method. The proposed method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface. The 3D model is generated from a sequence of range images. The Eigen-Texture method is practical because it does not require any detailed reflectance analysis of the object surface, and has great advantages due to the accurate 3D geometric models. This paper describes the method, and reports on its implementation. Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi |
CVPR | 1 |
| 1999 | Appearance Compression and Synthesis based on 3D Model for Mixed RealityabstractRendering photorealistic virtual objects from their real images is one of the main research issues in mixed reality systems. We previously proposed the Eigen-Texture method (K. Nishino et al., 1999), a new rendering method for generating virtual images of objects from their real images to deal with the problems posed by past work in image based methods and model based methods. Eigen-Texture method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface. However, we had a serious limitation in our system, due to the alignment problem of the 3D model and color images. We deal with this limitation by solving the alignment problem; we do this by using the method originally designed by P. Viola (1995). The paper describes the method and reports on how we implement it. Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi |
ICCV | 1 |