Satoshi Ikehata

dblp:117/4857 · DBLP profile ↗
← Back
33ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0002-6061-7956ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 10 first-author · 23 since 2021Artificial intelligence and machine learning · 20 · 9 first-author · 14 since 2021
YearPublicationVenuePosition
2026 Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo Under Limited Multi-Illumination Cues
abstract
Universal Photometric Stereo is a promising approach for recovering surface normals without strict lighting assumptions. However, it struggles when multi-illumination cues are unreliable, such as under biased lighting or in shadows or self-occluded regions of complex in-the-wild scenes. We propose GeoUniPS, a universal photometric stereo network that integrates synthetic supervision with high-level geometric priors from large-scale 3D reconstruction models pretrained on massive in-the-wild data. Our key insight is that these 3D reconstruction models serve as visual-geometry foundation models, inherently encoding rich geometric knowledge of real scenes. To leverage this, we design a Light-Geometry Dual-Branch Encoder that extracts both multi-illumination cues and geometric priors from the frozen 3D reconstruction model. We also address the limitations of the conventional orthographic projection assumption by introducing the PS-Perp dataset with realistic perspective projection to enable learning of spatially varying view directions. Extensive experiments demonstrate that GeoUniPS delivers state-of-the-arts performance across multiple datasets, both quantitatively and qualitatively, especially in the complex in-the-wild scenes.
King-Man Tam, Satoshi Ikehata, Yuta Asano, Zhaoyi An, Rei Kawakami
AAAI2
2026 Building and evaluating a realistic virtual world for large scale urban exploration from 360° videos
Mizuki Takenawa, Naoki Sugimoto, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
Multim. Tools Appl.4
2026 360CityGML: Realistic and Interactive Urban Visualization System Integrating CityGML Model and 360$^{\circ }$ Videos
abstract
We introduce a novel urban visualization system that integrates 3D urban model (CityGML) and 360$^{\circ }$∘ walkthrough videos. By aligning the videos with the model and dynamically projecting relevant video frames onto the geometries, our system creates photorealistic urban visualizations, allowing users to intuitively interpret geospatial data from a pedestrian view.
Tatsuro Banno, Mizuki Takenawa, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
IEEE Trans. Vis. Comput. Graph.4
2025 Rectified Lagrangian for Out-of-Distribution Detection in Modern Hopfield Networks
abstract
Modern Hopfield networks (MHNs) have recently gained significant attention in the field of artificial intelligence because they can store and retrieve a large set of patterns with an exponentially large memory capacity. A MHN is generally a dynamical system defined with Lagrangians of memory and feature neurons,where memories associated with in-distribution (ID) samples are represented by attractors in the feature space. One major problem in existing MHNs lies in managing out-of-distribution (OOD) samples because it was originally assumed that all samples are ID samples. To address this, we propose the rectified Lagrangian (RegLag), a new Lagrangian for memory neurons that explicitly incorporates an attractor for OOD samples in the dynamical system of MHNs. RecLag creates a trivial point attractor for any interaction matrix, enabling OOD detection by identifying samples that fall into this attractor as OOD. The interaction matrix is optimized so that the probability densities can be estimated to identify ID/OOD. We demonstrate the effectiveness of RecLag-based MHNs compared to energy-based OOD detection methods, including those using state-of-the-art Hopfield energies, across nine image datasets.
Ryo Moriai, Nakamasa Inoue, Masayuki Tanaka 0001, Rei Kawakami, Satoshi Ikehata, Ikuro Sato
AAAI5
2025 PS-EIP: Robust Photometric Stereo Based on Event Interval Profile
abstract
Recently, the energy-efficient photometric stereo method using an event camera (EventPS [67]) has been proposed to recover surface normals from events triggered by changes in logarithmic Lambertian reflections under a moving directional light source. However, EventPS treats each event interval independently, making it sensitive to noise, shadows, and non-Lambertian reflections. This paper proposes Photometric Stereo based on Event Interval Profile (PS-EIP), a robust method that recovers pixelwise surface normals from a time-series profile of event intervals. By exploiting the continuity of the profile and introducing an outlier detection method based on profile shape, our approach enhances robustness against outliers from shadows and specular reflections. Experiments using real event data from 3D-printed objects demonstrate that PS-EIP significantly improves robustness to outliers compared to EventPS’s deep-learning variant, EventPS-FCN, without relying on deep learning.
Kazuma Kitazawa, Takahito Aoto 0002, Satoshi Ikehata, Tsuyoshi Takatani
CVPR3
2025 PINO: Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized Groups
abstract
Generating realistic group interactions involving multiple characters remains challenging due to increasing complexity as group size expands. While existing conditional diffusion models incrementally generate motions by conditioning on previously generated characters, they rely on single shared prompts, limiting nuanced control and leading to overly simplified interactions. In this paper, we introduce Person-Interaction Noise Optimization (PINO), a novel, training-free framework designed for generating realistic and customizable interactions among groups of arbitrary size. PINO decomposes complex group interactions into semantically relevant pairwise interactions, and leverages pretrained two-person interaction diffusion models to incrementally compose group interactions. To ensure physical plausibility and avoid common artifacts such as overlapping or penetration between characters, PINO employs physics-based penalties during noise optimization. This approach allows precise user control over character orientation, speed, and spatial relationships without additional training. Comprehensive evaluations demonstrate that PINO generates visually realistic, physically coherent, and adaptable multi-person interactions suitable for diverse animation, gaming, and robotics applications.
Sakuya Ota, Kent Fujiwara, Satoshi Ikehata, Ikuro Sato
ICCV4
2025 Perface: Metric Learning in Perceptual Facial Similarity for Enhanced Face Anonymization
abstract
In response to rising societal awareness of privacy concerns, face anonymization techniques have advanced, including the emergence of face-swapping methods that replace one identity with another. Achieving a balance between anonymity and naturalness in face swapping requires careful selection of identities: overly similar faces compromise anonymity, while dissimilar ones reduce naturalness. Existing models, however, focus on binary identity classification "the same person or not", making it difficult to measure nuanced similarities such as "completely different" versus "highly similar but different." This paper proposes a human-perception-based face similarity metric, creating a dataset of 6,400 triplet annotations and metric learning to predict the similarity. Experimental results demonstrate significant improvements in both face similarity prediction and attribute-based face classification tasks over existing methods. Our dataset is available at https://github.com/kumanotanin/PerFace.
Haruka Kumagai, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
ICIP3
2025 Measuring Distortion Strength with Dewarping Diffusion Models in Anomaly Detection
abstract
Surface anomaly detection is a task to localize abnormal regions in a given image, typically used for product inspection. A representative approach is the reconstruction-based method, which detects defects using reconstruction errors computed by generative models, such as diffusion models trained exclusively on normal images. In reality, however, for products largely comprised of metal or resin, it is common to identify abnormalities based on local physical distortion levels; i.e., regions exceeding a predefined tolerance are regarded as defective. Reconstruction error-based approaches cannot directly estimate this metric. To address this issue, we propose DiffuDewarp, a novel method that directly estimates local distortions. Our approach defines a pseudo-deformation defect generation process as a new diffusion process based on localized warping. Experiments on the MVTec dataset demonstrate that our method outperforms state-of-the-art techniques in categories where local deformations are the primary cause of defects. Code is released at https://github.com/UCHIDA-AKIRA018/DiffuDewarp.
Akira Uchida, Satoshi Ikehata, Yuichi Yoshida, Ikuro Sato
ICIP2
2025 Field-of-View IoU for Object Detection in 360° Images
abstract
360°cameras have gained popularity over the last few years. In this paper, we propose two fundamental techniques-Field-of-View IoU (FoV-IoU) and 360Augmentation for object detection in 360° images. Although most object detection neural networks designed for perspective images are applicable to 360° images in equirectangular projection (ERP) format, their performance deteriorates owing to the distortion in ERP images. Our method can be readily integrated with existing perspective object detectors and significantly improves the performance. The FoV-IoU computes the intersection-over-union of two Field-of-View bounding boxes in a spherical image which could be used for training, inference, and evaluation while 360Augmentation is a data augmentation technique specific to 360° object detection task which randomly rotates a spherical image and solves the bias due to the sphere-to-plane projection. We conduct extensive experiments on the 360° indoor dataset with different types of perspective object detectors and show the consistent effectiveness of our method.
Satoshi Ikehata, Kiyoharu Aizawa
IEEE Trans. Image Process.2
2024 A Simple Finetuning Strategy Based on Bias-Variance Ratios of Layer-Wise Gradients
Mao Tomita, Ikuro Sato, Rei Kawakami, Nakamasa Inoue, Satoshi Ikehata, Masayuki Tanaka 0001
ACCV (8)5
2024 Entity-NeRF: Detecting and Removing Moving Entities in Urban Scenes
abstract
Recent advancements in the study of Neural Radiance Fields (NeRF) for dynamic scenes often involve explicit modeling of scene dynamics. However, this approach faces challenges in modeling scene dynamics in urban environments, where moving objects of various categories and scales are present. In such settings, it becomes crucial to effectively eliminate moving objects to accurately reconstruct static backgrounds. Our research introduces an innovative method, termed here as Entity-NeRF, which combines the strengths of knowledge-based and statistical strategies. This approach utilizes entity-wise statistics, leveraging entity segmentation and stationary entity classification through thing/stuff segmentation. To assess our methodology, we created an urban scene dataset masked with moving objects. Our comprehensive experiments demonstrate that Entity-NeRF notably outperforms existing techniques in removing moving objects and reconstructing static urban backgrounds, both quantitatively and qualitatively.11Our project page is available at https://otonari726.github.io/entitynerf/
Takashi Otonari, Satoshi Ikehata, Kiyoharu Aizawa
CVPR2
2024 SpectraM-PS: Spectrally Multiplexed Photometric Stereo Under Unknown Spectral Composition
Satoshi Ikehata, Yuta Asano
ECCV (2)1
2024 MERLiN: Single-Shot Material Estimation and Relighting for Photometric Stereo
Ashish Tiwari 0005, Satoshi Ikehata, Shanmuganathan Raman
ECCV (12)2
2024 Gumbel-NeRF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields
abstract
We propose Gumbel-NeRF, a mixture-of-expert (MoE) neural radiance fields (NeRF) model with a hindsight expert selection mechanism for synthesizing novel views of unseen objects. Previous studies have shown that the MoE structure provides high-quality representations of a given large-scale scene consisting of many objects. However, we observe that such a MoE NeRF model often produces low-quality representations in the vicinity of experts’ boundaries when applied to the task of novel view synthesis of an unseen object from one/few-shot input. We find that this deterioration is primarily caused by the foresight expert selection mechanism, which may leave an unnatural discontinuity in the object shape near the experts’ boundaries. Gumbel-NeRF adopts a hindsight expert selection mechanism, which guarantees continuity in the density field even near the experts’ boundaries. Experiments using the SRN cars dataset demonstrate the superiority of Gumbel-NeRF over the baselines in terms of various image quality metrics. The code will be available upon acceptance.
Yusuke Sekikawa, Chingwei Hsu, Satoshi Ikehata, Rei Kawakami, Ikuro Sato
ICIP3
2024 Investigating the Perception of Facial Anonymization Techniques in 360° Videos
abstract
In this work, we investigate facial anonymization techniques in 360° videos and assess their influence on the perceived realism, anonymization effect, and presence of participants. In comparison to traditional footage, 360° videos can convey engaging, immersive experiences that accurately represent the atmosphere of real-world locations. As the entire environment is captured simultaneously, it is necessary to anonymize the faces of bystanders in recordings of public spaces. Since this alters the video content, the perceived realism and immersion could be reduced. To understand these effects, we compare non-anonymized and anonymized 360° videos using blurring, black boxes, and face-swapping shown either on a regular screen or in a head-mounted display (HMD). Our results indicate significant differences in the perception of the anonymization techniques. We find that face-swapping is the most realistic and least disruptive; however, participants raised concerns regarding the effectiveness of the anonymization. Furthermore, we observe that presence is affected by facial anonymization in HMD condition. Overall, the results underscore the need for facial anonymization techniques that balance both photo-realism and a sense of privacy.
Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
ACM Trans. Appl. Percept.2
2023 Scalable, Detailed and Mask-Free Universal Photometric Stereo
abstract
In this paper, we introduce SDM-UniPS, a groundbreaking Scalable, Detailed, Mask-free, and Universal Photometric Stereo network. Our approach can recover astonishingly intricate surface normal maps, rivaling the quality of 3D scanners, even when images are captured under unknown, spatially-varying lighting conditions in uncontrolled environments. We have extended previous universal photometric stereo networks to extract spatial-light features, utilizing all available information in high-resolution input images and accounting for non-local interactions among surface points. Moreover, we present a new synthetic training dataset that encompasses a diverse range of shapes, materials, and illumination scenarios found in real-world scenes. Through extensive evaluation, we demonstrate that our method not only surpasses calibrated, lighting-specific techniques on public benchmarks, but also excels with a significantly smaller number of input images even without object masks.
Satoshi Ikehata
CVPR1
2023 360RVW: Fusing Real 360° Videos and Interactive Virtual Worlds
abstract
We propose a system to generate 360° realistic virtual worlds (360RVW) for the interactive spatial exploration of omnidirectional street-view videos. Our 360RVW enables users to explore photorealistic scenes with digital avatars, and interact with others. To create the virtual worlds our system only requires 360° videos with annotations of the start and end camera coordinate as input. We first detect street intersections to divide the input videos and remove the camera operator from the recordings using a video completion technique. Next, we analyze the 3D structure of the scene using semantic segmentation to define walkable areas. Finally, we render the environment using an ellipsoid projection surface to achieve a more realistic integration of the avatar into real-world 360° videos. The whole process is largely automated, enabling users to produce realistic and interactive virtual worlds without specialized skills or time-consuming manual interventions.
Mizuki Takenawa, Naoki Sugimoto, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
ACM Multimedia4
2022 Non-uniform Sampling Strategies for NeRF on 360° images
Takashi Otonari, Satoshi Ikehata, Kiyoharu Aizawa
BMVC2
2022 Universal Photometric Stereo Network using Global Lighting Contexts
abstract
This paper tackles a new photometric stereo task, named universal photometric stereo. Unlike existing tasks that assumed specific physical lighting models; hence, drastically limited their usability, a solution algorithm of this task is supposed to work for objects with diverse shapes and materials under arbitrary lighting variations without assum-ig any specific models. To solve this extremely challenging task, we present a purely data-driven method, which eliminates the prior assumption of lighting by replacing the recovery of physical lighting parameters with the extraction of the generic lighting representation, named global lighting contexts. We use them like lighting parameters in a calibrated photometric stereo network to recover surface normal vectors pixelwisely. To adapt our network to a wide variety of shapes, materials and lightings, it is trained on a new synthetic dataset which simulates the appearance of objects in the wild. Our method is compared with other state-of-the-art uncalibrated photometric stereo methods on our test data to demonstrate the significance of our method.
Satoshi Ikehata
CVPR1
2022 Dual-Erp Representation for Object Detection in 360° Images
abstract
Object detection has achieved good performance on perspective images. However, a general object detector does not maintain this performance when applied to a 360° image in a single equirectangular projection (ERP) or multi-projection representation because of the distortion in the high-latitude region or discontinuity at the boundaries. In this paper, we proposed dual-ERP, which is a multi-view ERP representation, as the network input for 360° object detection in training and inference. Dual-ERP combines the advantages of single ERP and multi-projection representations, and it can easily be integrated with existing object detectors. The experimental results showed that compared to other representations, dual-ERP significantly improved the performance of different baseline object detectors.
Satoshi Ikehata, Kiyoharu Aizawa
ICIP2
2022 Does Physical Interpretability of Observation Map Improve Photometric Stereo Networks?
abstract
In this paper, we revisit observation map which is the input representation for the deep photometric stereo networks where pixelwise observations under different lights are protectively integrated to handle an arbitrary number of input images. Based on the hypothesis that the physical interpretability of observation map contributes to its performance, we empirically validate it by proposing two novel ideas; one is a pixelwise unified inverse rendering framework which accounts the physical reasoning to recover the surface normals and the other is the network architecture that is equivariant/invariant to the view-axis-around rotation of the pixelwise observation map. By introducing these two ideas, our experimental evaluation on the public dataset indicated that more explicit physical reasoning of observation map improves the performance of the photometric stereo task.
Satoshi Ikehata
ICIP1
2021 PS-Transformer: Learning Sparse Photometric Stereo Network using Self-Attention Mechanism
Satoshi Ikehata
BMVC1
2021 Intersection Prediction from Single 360° Image via Deep Detection of Possible Direction of Travel
Naoki Sugimoto, Satoshi Ikehata, Kiyoharu Aizawa
BMVC2
2021 360° Single Image Super Resolution via Distortion-Aware Network and Distorted Perspective Images
abstract
Effective 360° imaging requires a very high resolution because the field of view is extraordinarily high. Single-image super-resolution (SISR) applied to 360° imaging has the potential to solve the resolution/quality problem in this modality. In this paper, we exploit existing perspective SISR networks to address this problem by (1) introducing a distortion map as an additional input with the $360^{\circ}-$ distortion-aware loss function, and (2) augmenting the training 360° images by distorting the perspective images. We also present a new 360° image dataset from YouTube for training. Our extensive experiments show that how each component contributes to the better transfer from the perspective domain to the 360° domain and merging all the ideas leads to the best performance in quantitative and qualitative ways for the 360° SISR task.
Akito Nishiyama, Satoshi Ikehata, Kiyoharu Aizawa
ICIP2
2018 CNN-PS: CNN-Based Photometric Stereo for General Non-convex Surfaces
Satoshi Ikehata
ECCV (15)1
2018 Efficiency-enhanced cost-volume filtering featuring coarse-to-fine strategy
abstract
Cost-volume filtering (CVF) is one of the most widely used techniques for solving general multi-labeling problems based on a Markov random field (MRF). However it is inefficient when the label space size (i.e., the number of labels) is large. This paper presents a coarse-to-fine strategy for cost-volume filtering that efficiently and accurately addresses multi-labeling problems with a large label space size. Based on the observation that true labels at the same coordinates in images of different scales are highly correlated, we truncate unimportant labels for cost-volume filtering by leveraging the labeling output of lower scales. Experimental results show that our algorithm achieves much higher efficiency than the original CVF method while maintaining a comparable level of accuracy. Although we performed experiments that deal with only stereo matching and optical flow estimation, the proposed method can be employed in many other applications because of the applicability of CVF to general discrete pixel-labeling problems based on an MRF.
Ryosuke Furuta, Satoshi Ikehata, Toshihiko Yamasaki, Kiyoharu Aizawa
Multim. Tools Appl.2
2017 From Bayesian Sparsity to Gated Recurrent Nets
abstract
The iterations of many first-order algorithms, when applied to minimizing common regularized regression functions, often resemble neural network layers with pre-specified weights. This observation has prompted the development of learning-based approaches that purport to replace these iterations with enhanced surrogates forged as DNN models from available training data. For example, important NP-hard sparse estimation problems have recently benefitted from this genre of upgrade, with simple feedforward or recurrent networks ousting proximal gradient-based iterations. Analogously, this paper demonstrates that more powerful Bayesian algorithms for promoting sparsity, which rely on complex multi-loop majorization-minimization techniques, mirror the structure of more sophisticated long short-term memory (LSTM) networks, or alternative gated feedback networks previously designed for sequence prediction. As part of this development, we examine the parallels between latent variable trajectories operating across multiple time-scales during optimization, and the activations within deep network structures designed to adaptively model such characteristic sequences. The resulting insights lead to a novel sparse estimation system that, when granted training data, can estimate optimal solutions efficiently in regimes where other algorithms fail, including practical direction-of-arrival (DOA) and 3D geometry recovery problems. The underlying principles we expose are also suggestive of a learning process for a richer class of multi-loop algorithms in other domains.
Hao He 0011, Bo Xin, Satoshi Ikehata, David P. Wipf
NIPS3
2015 Structured Indoor Modeling
abstract
This paper presents a novel 3D modeling framework that reconstructs an indoor scene as a structured model from panorama RGBD images. A scene geometry is represented as a graph, where nodes correspond to structural elements such as rooms, walls, and objects. The approach devises a structure grammar that defines how a scene graph can be manipulated. The grammar then drives a principled new reconstruction algorithm, where the grammar rules are sequentially applied to recover a structured model. The paper also proposes a new room segmentation algorithm and an offset-map reconstruction algorithm that are used in the framework and can enforce architectural shape priors far beyond existing state-of-the-art. The structured scene representation enables a variety of novel applications, ranging from indoor scene visualization, automated floorplan generation, Inverse-CAD, and more. We have tested our framework and algorithms on six synthetic and five real datasets with qualitative and quantitative evaluations. The source code and the data are available at the project website [15].
Satoshi Ikehata, Yasutaka Furukawa
ICCV1
2014 Photometric Stereo Using Constrained Bivariate Regression for General Isotropic Surfaces
abstract
This paper presents a photometric stereo method that is purely pixelwise and handles general isotropic surfaces in a stable manner. Following the recently proposed sum-of-lobes representation of the isotropic reflectance function, we constructed a constrained bivariate regression problem where the regression function is approximated by smooth, bivariate Bernstein polynomials. The unknown normal vector was separated from the unknown reflectance function by considering the inverse representation of the image formation process, and then we could accurately compute the unknown surface normals by solving a simple and efficient quadratic programming problem. Extensive evaluations that showed the state-of-the-art performance using both synthetic and real-world images were performed.
Satoshi Ikehata, Kiyoharu Aizawa
CVPR1
2014 Coarse-to-fine strategy for efficient cost-volume filtering
abstract
Cost-volume filtering is one of the most widely known techniques to solve general multi-label problems, however it is problematically inefficient when the label space size is extremely large. This paper presents a coarse-to-fine strategy of the cost-volume filtering that handles efficiently and accurately multi-label problems with a large label space size. Based upon the observation that true labels at the same image coordinate of different scales are highly correlated, we truncate unimportant labels for the cost-volume filtering by leveraging the labeling output of lower scales. Experimental results show that our algorithm achieves much higher efficiency than the original cost-volume filtering while enjoying the comparable accuracy to it.
Ryosuke Furuta, Satoshi Ikehata, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP2
2014 Photometric Stereo Using Sparse Bayesian Regression for General Diffuse Surfaces
abstract
Most conventional algorithms for non-Lambertian photometric stereo can be partitioned into two categories. The first category is built upon stable outlier rejection techniques while assuming a dense Lambertian structure for the inliers, and thus performance degrades when general diffuse regions are present. The second utilizes complex reflectance representations and non-linear optimization over pixels to handle non-Lambertian surfaces, but does not explicitly account for shadows or other forms of corrupting outliers. In this paper, we present a purely pixel-wise photometric stereo method that stably and efficiently handles various non-Lambertian effects by assuming that appearances can be decomposed into a sparse, non-diffuse component (e.g., shadows, specularities, etc.) and a diffuse component represented by a monotonic function of the surface normal and lighting dot-product. This function is constructed using a piecewise linear approximation to the inverse diffuse model, leading to closed-form estimates of the surface normals and model parameters in the absence of non-diffuse corruptions. The latter are modeled as latent variables embedded within a hierarchical Bayesian model such that we may accurately compute the unknown surface normals while simultaneously separating diffuse from non-diffuse components. Extensive evaluations are performed that show state-of-the-art performance using both synthetic and real-world images.
Satoshi Ikehata, David P. Wipf, Yasuyuki Matsushita, Kiyoharu Aizawa
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Depth map inpainting and super-resolution based on internal statistics of geometry and appearance
abstract
Depth maps captured by multiple sensors often suffer from poor resolution and missing pixels caused by low reflectivity and occlusions in the scene. To address these problems, we propose a combined framework of patch-based inpainting and super-resolution. Unlike previous works, which relied solely on depth information, we explicitly take advantage of the internal statistics of a depth map and a registered highresolution texture image that capture the same scene. We account these statistics to locate non-local patches for hole filling and constrain the sparse coding-based super-resolution problem. Extensive evaluations are performed and show the state-of-the-art performance when using real-world datasets.
Satoshi Ikehata, Ji-Ho Cho, Kiyoharu Aizawa
ICIP1
2012 Robust photometric stereo using sparse regression
abstract
This paper presents a robust photometric stereo method that effectively compensates for various non-Lambertian corruptions such as specularities, shadows, and image noise. We construct a constrained sparse regression problem that enforces both Lambertian, rank-3 structure and sparse, additive corruptions. A solution method is derived using a hierarchical Bayesian approximation to accurately estimate the surface normals while simultaneously separating the non-Lambertian corruptions. Extensive evaluations are performed that show state-of-the-art performance using both synthetic and real-world images.
Satoshi Ikehata, David P. Wipf, Yasuyuki Matsushita, Kiyoharu Aizawa
CVPR1