VLDB 2026 Research / reviewers in the wild / expert
Yasuyuki Matsushita
dblp:11/3619
· DBLP profile ↗
150ranked-venue papers
9as first author
33since 2021 · last 2026
0000-0002-1935-4752ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 126 · 8 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 107 · 6 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DP-SfM: Dual-Pixel Structure-From-Motion Without Scale AmbiguityabstractMulti-view 3D reconstruction, namely, structure-from-motion followed by multi-view stereo, is a fundamental component of 3D computer vision. In general, multi-view 3D reconstruction suffers from an unknown scale ambiguity unless a reference object of known size is present in the scene. In this article, we show that multi-view images captured using a dual-pixel (DP) sensor can automatically resolve the scale ambiguity, without requiring a reference object or prior calibration. Specifically, the defocus blur observed in DP images provides sufficient information to determine the absolute scale when paired with depth maps (up to scale) recovered from multi-view 3D reconstruction. Based on this observation, we develop a simple yet effective linear method to estimate the absolute scale, followed by the intensity-based optimization stage that aligns the left and right DP images by shifting them back toward each other using cross-view blur kernels. Experiments demonstrate the effectiveness of the proposed approach across diverse scenes captured with different cameras and lenses. Lilika Makabe, Kohei Ashida, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Instance-wise distribution control of text-to-image diffusion modelsabstractText-to-image diffusion models are increasingly used to generate synthetic datasets for downstream vision tasks. However, they often inherit biases from large-scale training data, which can result in unbalanced attribute distributions in the generated images. While prior efforts have attempted to mitigate these biases, most focus on single-object images and struggle to control attributes across object instances in multi-instance generations. To address this limitation, we propose an instance-wise control of the attribute distribution by fine-tuning diffusion models with guidance from a pre-trained object detector and an attribute classifier. Our approach aligns the attribute distribution over object instances in generated images with a user-defined distribution, which enables precise control over attribute proportions at the instance level. Experiments across various objects and attributes demonstrate that our method generates high-quality, multi-instance images that match the specified distribution, supporting the scalable creation of distribution-aware synthetic datasets for in-the-wild vision tasks. Weng Ian Chan, Hiroaki Santo, Yasuyuki Matsushita, Fumio Okura |
Pattern Recognit. | 3 |
| 2025 | HoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian SplattingabstractNovel view synthesis has demonstrated impressive progress recently, with 3D Gaussian splatting (3DGS) offering efficient training time and photorealistic real-time rendering. However, reliance on Cartesian coordinates limits 3DGS’s performance on distant objects, which is important for reconstructing unbounded outdoor environments. We found that, despite its ultimate simplicity, using homogeneous coordinates, a concept on the projective geometry, for the 3DGS pipeline remarkably improves the rendering accuracies of distant objects. We therefore propose Homogeneous Gaussian Splatting (HoGS) incorporating homogeneous coordinates into the 3DGS framework, providing a unified representation for enhancing near and distant objects. HoGS effectively manages both expansive spatial positions and scales particularly in outdoor unbounded environments by adopting projective geometry principles. Experiments show that HoGS significantly enhances accuracy in reconstructing distant objects while maintaining high-quality rendering of nearby objects, along with fast training speed and real-time rendering capability. Our implementations are available on our project page https://kh129.github.io/hogs/. Xinpeng Liu 0007, Zeyi Huang, Fumio Okura, Yasuyuki Matsushita |
CVPR | 4 |
| 2025 | Spectral Sensitivity Estimation with an Uncalibrated Diffraction GratingabstractThis paper introduces a practical and accurate calibration method for camera spectral sensitivity using a diffraction grating. Accurate calibration of camera spectral sensitivity is crucial for various computer vision tasks, including color correction, illumination estimation, and material analysis. Unlike existing approaches that require specialized narrow-band filters or reference targets with known spectral reflectances, our method only requires an uncalibrated diffraction grating sheet, readily available off-the-shelf. By capturing images of the direct illumination and its diffracted pattern through the grating sheet, our method estimates both the camera spectral sensitivity and the diffraction grating parameters in a closed-form manner. Experiments on synthetic and real-world data demonstrate that our method outperforms conventional reference target-based methods, underscoring its effectiveness and practicality. Lilika Makabe, Hiroaki Santo, Fumio Okura, Michael S. Brown, Yasuyuki Matsushita |
ICCV | 5 |
| 2025 | NeuraLeaf: Neural Parametric Leaf Models with Shape and Deformation Disentanglement
Yang Yang 0124, Zhendong Mao 0001, Hiroaki Santo, Yasuyuki Matsushita, Fumio Okura |
ICCV | 4 |
| 2024 | DiLiGenRT: A Photometric Stereo Dataset with Quantified Roughness and TranslucencyabstractPhotometric stereo faces challenges from non-Lambertian reflectance in real-world scenarios. Systematically measuring the reliability of photometric stereo methods in handling such complex reflectance necessitates a real-world dataset with quantitatively controlled reflectances. This paper introduces DiLiGenRT, the first real-world dataset for evaluating photometric stereo methods under quantified reflectances by manufacturing 54 hemispheres with varying degrees of two reflectance properties: Roughness and Transluency, Unlike qualitative and semantic labels, such as “diffuse” and “specular,” that have been used in previous datasets, our quantified dataset allows comprehensive and systematic benchmark evaluations. In addition, it facilitates selecting best-fit photometric stereo methods based on the quantitative reflectance properties. Our dataset and benchmark results are available at https://photometricstereo.github.io/diligentrt.html. Heng Guo 0003, Jieji Ren, Feishi Wang, Boxin Shi, Ming Jun Ren, Yasuyuki Matsushita |
CVPR | 6 |
| 2024 | MVCPS-NeuS: Multi-View Constrained Photometric Stereo for Neural Surface ReconstructionabstractMulti-view photometric stereo (MVPS) recovers a high-fidelity 3D shape of a scene by benefiting from both multi-view stereo and photometric stereo. While photometric stereo boosts detailed shape reconstruction, it necessitates recording images under various light conditions for each viewpoint. In particular, calibrating the light directions for each view significantly increases the cost of acquiring images. To make MVPS more accessible, we introduce a practical and easy-to-implement setup, multi-view constrained photometric stereo (MVCPS), where the light directions are unknown but constrained to move together with the camera. Unlike con-ventional multi-view uncalibrated photometric stereo, our constrained setting reduces the ambiguities of surface normal estimates from per-view linear ambiguities to a single and global linear one, thereby simplifying the disambiguation process. The proposed method integrates the ambiguous surface normal into neural surface reconstruction (NeuS) to simultaneously resolve the global ambiguity and estimate the detailed 3D shape. Experiments demonstrate that our method estimates accurate shapes under sparse viewpoints using only a few multi-view constrained light sources. Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
CVPR | 3 |
| 2024 | Resolving Scale Ambiguity in Multi-view 3D Reconstruction Using Dual-Pixel Sensors
Kohei Ashida, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ECCV (50) | 4 |
| 2024 | In Memoriam: Xiaoou Tang
Yasuyuki Matsushita, Svetlana Lazebnik, Jiri Matas |
Int. J. Comput. Vis. | 1 |
| 2023 | Multi-View Azimuth Stereo via Tangent Space ConsistencyabstractWe present a method for 3D reconstruction only using calibrated multi-view surface azimuth maps. Our method, multi-view azimuth stereo, is effective for textureless or specular surfaces, which are difficult for conventional multi-view stereo methods. We introduce the concept of tangent space consistency: Multi-view azimuth observations of a surface point should be lifted to the same tangent space. Leveraging this consistency, we recover the shape by optimizing a neural implicit surface representation. Our method harnesses the robust azimuth estimation capabilities of photometric stereo methods or polarization imaging while bypassing potentially complex zenith angle estimation. Experiments using azimuth maps from various sources validate the accurate shape recovery with our method, even without zenith angles. Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
CVPR | 4 |
| 2023 | Learning to Synthesize Photorealistic Dual-pixel Images from RGBD framesabstractAs a special sensor that implicitly provides ordinal depth information, dual-pixel (DP) appears to be beneficial for various tasks such as defocus deblurring and monocular depth estimation. Recent advances in data-driven dual-pixel (DP) research are bottlenecked by the difficulties in reaching large-scale DP datasets, and a photorealistic image synthesis approach appears to be a credible solution. To benchmark the accuracy of various existing DP image simulators and facilitate data-driven DP image synthesis, this work presents a real-world DP dataset consisting of approximately 5000 high-quality pairs of sharp images, DP defocus blur images, detailed imaging parameters, and accurate depth maps. Based on this large-scale dataset, we also propose a holistic data-driven framework to synthesize photorealistic DP images, where a neural network replaces conventional handcrafted imaging models. Experiments show that our neural DP simulator can generate more photorealistic DP images than existing state-of-the-art methods and effectively benefit data-driven DP-related tasks. Our code and dataset are released at https://github.com/SILI1994/Dual-Pixel-Simulator. Feiran Li, Heng Guo 0003, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ICCP | 5 |
| 2023 | Near-light Photometric Stereo with Symmetric LightsabstractThis paper describes a linear solution method for near-light photometric stereo by exploiting symmetric light source arrangements. Unlike conventional non-convex optimization approaches, by arranging multiple sets of symmetric nearby light source pairs, our method derives a closed-form solution for surface normal and depth without requiring initialization. In addition, our method works as long as the light sources are symmetrically distributed about an arbitrary point even when the entire spatial offset is uncalibrated. Experiments showcase the accuracy of shape recovery accuracy of our method, achieving comparable results to the state-of-the-art calibrated near-light photometric stereo method while significantly reducing requirements of careful depth initialization and light calibration. Lilika Makabe, Heng Guo 0003, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ICCP | 5 |
| 2023 | A Compact BRDF Scanner with Multi-conjugate OpticsabstractReflectance, represented as Bidirectional reflectance distribution functions (BRDFs), is an important scene property that we wish to obtain from the real world together with the scene’s 3D shape. BRDF acquisition, however, remains a difficult task because it is time-consuming and requires huge and costly devices. To make the BRDF acquisition easier and accessible to everyone, we develop a compact projector-camera-based BRDF scanner with coaxial multi-conjugate optics, where the sensor/source is conjugate with the Fourier transform plane of the objective lens, and the camera/projector pupil is conjugate with the sample plane. It consists of only off-the-shelf components, and the carefully designed conjugate optics make the entire system compact. BRDFs of outdoor objects are scanned to show the validity of our device. Kensuke Uchida, Hajime Nagahara, Yasuyuki Matsushita |
ICCP | 3 |
| 2023 | Shuffled Linear Regression with Outliers in Both Covariates and Responses
Feiran Li, Kent Fujiwara, Fumio Okura, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 4 |
| 2023 | Discrete Search Photometric Stereo for Fast and Accurate Shape EstimationabstractWe consider the problem of estimating surface normals of a scene with spatially varying, general bidirectional reflectance distribution functions (BRDFs) observed by a static camera under varying distant illuminations. Unlike previous approaches that rely on continuous optimization of surface normals, we cast the problem as a discrete search problem over a set of finely discretized surface normals. In this setting, we show that the expensive processes can be precomputed in a scene-independent manner, resulting in accelerated inference. We discuss two variants of our discrete search photometric stereo (DSPS), one working with continuous linear combinations of BRDF bases and the other working with discrete BRDFs sampled from a BRDF space. Experiments show that DSPS has comparable accuracy to state-of-the-art exemplar-based photometric stereo methods while achieving 10-100x acceleration. Kenji Enomoto, Michael Waechter, Fumio Okura, Kiriakos N. Kutulakos, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Bilateral Normal Integration
Hiroaki Santo, Boxin Shi, Fumio Okura, Yasuyuki Matsushita |
ECCV (1) | 5 |
| 2022 | Shape-coded ArUco: Fiducial Marker for Bridging 2D and 3D ModalitiesabstractWe introduce a fiducial marker for the registration of two-dimensional (2D) images and untextured three-dimensional (3D) shapes that are recorded by commodity laser scanners. Specifically, we design a 3D-version of the ArUco marker that retains exactly the same appearance as its 2D counterpart from any viewpoint above the marker but contains shape information. The shape-coded ArUco can naturally work with off-the-shelf ArUco marker detectors in the 2D image domain. For the 3D domain, we develop a method for detecting the marker in an untextured 3D point cloud. Experiments demonstrate accurate 2D-3D registration using our shape-coded ArUco markers in comparison to baseline methods. Lilika Makabe, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
WACV | 4 |
| 2022 | Symmetric-light Photometric StereoabstractThis paper presents symmetric-light photometric stereo for surface normal estimation, in which directional lights are distributed symmetrically with respect to the optic center. Unlike previous studies of ring-light settings that required the information of ring radius, we show that even without the knowledge of the exact light source locations or their distances from the optic center, the symmetric configuration provides us sufficient information for recovering unique surface normals without ambiguity. Specifically, under the symmetric lights, measurements of a pair of scene points having distinct surface normals but the same albedo yield a system of constrained quadratic equations about the surface normal, which has a unique solution. Experiments demonstrate that the proposed method alleviates the need for geometric light source calibration while maintaining the accuracy of calibrated photometric stereo. Kazuma Minami, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
WACV | 4 |
| 2022 | Shape and Albedo Recovery by Your Phone using Stereoscopic Flash and No-Flash PhotographyabstractRecovering shape and albedo for the immense number of existing cultural heritage artifacts is challenging. Accurate 3D reconstruction systems are typically expensive and thus inaccessible to many and cheaper off-the-shelf 3D sensors often generate results of unsatisfactory quality. This paper presents a high-fidelity shape and albedo recovery method that only requires a stereo camera and a flashlight, a typical camera setup equipped in many off-the-shelf smartphones. The stereo camera allows us to infer rough shape from a pair of no-flash images, and a flash image is further captured for shape refinement based on our flash/no-flash image formation model. We verify the effectiveness of our method on real-world artifacts in indoor and outdoor conditions using smartphones with different camera/flashlight configurations. Comparison results demonstrate that our stereoscopic flash and no-flash photography benefits the high-fidelity shape and albedo recovery on a smartphone. Using our method, people can immediately turn their phones into high-fidelity 3D scanners, facilitating the digitization of cultural heritage artifacts. Michael Waechter, Boxin Shi, Fumio Okura, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 7 |
| 2022 | Eliminating Temporal Illumination Variations in Whisk-broom Hyperspectral ImagingabstractAbstract We propose a method for eliminating the temporal illumination variations in whisk-broom (point-scan) hyperspectral imaging. Whisk-broom scanning is useful for acquiring a spatial measurement using a pixel-based hyperspectral sensor. However, when it is applied to outdoor cultural heritages, temporal illumination variations become an issue due to the lengthy measurement time. As a result, the incoming illumination spectra vary across the measured image locations because different locations are measured at different times. To overcome this problem, in addition to the standard raster scan, we propose an additional perpendicular scan that traverses the raster scan. We show that this additional scan allows us to infer the illumination variations over the raster scan. Furthermore, the sparse structure in the illumination spectrum is exploited to robustly eliminate these variations. We quantitatively show that a hyperspectral image captured under sunlight is indeed affected by temporal illumination variations, that a Naïve mitigation method suffers from severe artifacts, and that the proposed method can robustly eliminate the illumination variations. Finally, we demonstrate the usefulness of the proposed method by capturing historic stained-glass windows of a French cathedral. Takuya Funatomi, Takehiro Ogawa, Kenichiro Tanaka, Hiroyuki Kubo, Guillaume Caron, El Mustapha Mouaddib, Yasuyuki Matsushita, Yasuhiro Mukaigawa |
Int. J. Comput. Vis. | 7 |
| 2022 | Multispectral Photometric Stereo for Spatially-Varying Spectral Reflectances
Heng Guo 0003, Fumio Okura, Boxin Shi, Takuya Funatomi, Yasuhiro Mukaigawa, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 6 |
| 2022 | Deep Photometric Stereo for Non-Lambertian SurfacesabstractThis paper addresses the problem of photometric stereo, in both calibrated and uncalibrated scenarios, for non-Lambertian surfaces based on deep learning. We first introduce a fully convolutional deep network for calibrated photometric stereo, which we call PS-FCN. Unlike traditional approaches that adopt simplified reflectance models to make the problem tractable, our method directly learns the mapping from reflectance observations to surface normal, and is able to handle surfaces with general and unknown isotropic reflectance. At test time, PS-FCN takes an arbitrary number of images and their associated light directions as input and predicts a surface normal map of the scene in a fast feed-forward pass. To deal with the uncalibrated scenario where light directions are unknown, we introduce a new convolutional network, named LCNet, to estimate light directions from input images. The estimated light directions and the input images are then fed to PS-FCN to determine the surface normals. Our method does not require a pre-defined set of light directions and can handle multiple images in an order-agnostic manner. Thorough evaluation of our approach on both synthetic and real datasets shows that it outperforms state-of-the-art methods in both calibrated and uncalibrated scenarios. Guanying Chen, Kai Han 0001, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee Kenneth Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Patch-Based Uncalibrated Photometric Stereo Under Natural IlluminationabstractThis paper presents a photometric stereo method that works with unknown natural illumination without any calibration objects or initial guess of the target shape. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary orthogonal ambiguity. We further build the patch connections by extracting consistent surface normal pairs via spatial overlaps among patches and intensity profiles. Guided by these connections, the local ambiguities are unified to a global orthogonal one through Markov Random Field optimization and rotation averaging. After applying the integrability constraint, our solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods. Heng Guo 0003, Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Ping Tan 0002, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Shell Theory: A Statistical Model of RealityabstractThe foundational assumption of machine learning is that the data under consideration is separable into classes; while intuitively reasonable, separability constraints have proven remarkably difficult to formulate mathematically. We believe this problem is rooted in the mismatch between existing statistical techniques and commonly encountered data; object representations are typically high dimensional but statistical techniques tend to treat high dimensions a degenerate case. To address this problem, we develop a dedicated statistical framework for machine learning in high dimensions. The framework derives from the observation that object relations form a natural hierarchy; this leads us to model objects as instances of a high dimensional, hierarchal generative processes. Using a distance based statistical technique, also developed in this paper, we show that in such generative processes, instances of each process in the hierarchy, are almost-always encapsulated by a distinctive-shell that excludes almost-all other instances. The result is shell theory, a statistical machine learning framework in which separability constraints (distinctive-shells) are formally derived from the assumed generative process. Wen-Yan Lin, Changhao Ren, Ngai-Man Cheung, Hongdong Li, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Deep Photometric Stereo Networks for Determining Surface Normal and ReflectancesabstractThis article presents a photometric stereo method based on deep learning. One of the major difficulties in photometric stereo is designing an appropriate reflectance model that is both capable of representing real-world reflectances and computationally tractable for deriving surface normal. Unlike previous photometric stereo methods that rely on a simplified parametric image formation model, such as the Lambert's model, the proposed method aims at establishing a flexible mapping between complex reflectance observations and surface normal using a deep neural network. In addition, the proposed method predicts the reflectance, which allows us to understand surface materials and to render the scene under arbitrary lighting conditions. As a result, we propose a deep photometric stereo network (DPSN) that takes reflectance observations under varying light directions and infers the surface normal and reflectance in a per-pixel manner. To make the DPSN applicable to real-world scenes, a dataset of measured BRDFs (MERL BRDF dataset) has been used for training the network. Evaluation using simulation and real-world scenes shows the effectiveness of the proposed approach in estimating both surface normal and reflectances. Hiroaki Santo, Masaki Samejima, Yusuke Sugano, Boxin Shi, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Normal Integration via Inverse Plane Fitting With Minimum Point-to-Plane DistanceabstractThis paper presents a surface normal integration method that solves an inverse problem of local plane fitting. Surface reconstruction from normal maps is essential in photometric shape reconstruction. To this end, we formulate normal integration in the camera coordinates and jointly solve for 3D point positions and local plane displacements. Unlike existing methods that consider the vertical distances between 3D points, we minimize the sum of squared point-to-plane distances. Our method can deal with both orthographic or perspective normal maps with arbitrary boundaries. Compared to existing normal integration methods, our method avoids the checkerboard artifact and performs more robustly against natural boundaries, sharp features, and outliers. We further provide a geometric analysis of the source of artifacts that appear in previous methods based on our plane fitting formulation. Experimental results on analytically computed, synthetic, and real-world surfaces show that our method yields accurate and stable reconstruction for both orthographic and perspective normal maps1. Boxin Shi, Fumio Okura, Yasuyuki Matsushita |
CVPR | 4 |
| 2021 | Multispectral Photometric Stereo for Spatially-Varying Spectral Reflectances: A Well Posed Problem?abstractMultispectral photometric stereo (MPS) aims at recovering the surface normal of a scene from a single-shot multi-spectral image, which is known as an ill-posed problem. To make the problem well-posed, existing MPS methods rely on restrictive assumptions, such as shape prior, surfaces having a monochromatic with uniform albedo. This paper alleviates the restrictive assumptions in existing methods. We show that the problem becomes well-posed for a surface with a uniform chromaticity but spatially-varying albedos based on our new formulation. Specifically, if at least three (or two) scene points share the same chromaticity, the proposed method uniquely recovers their surface normals and spectral reflectance with the illumination of more than or equal to four (or five) spectral lights. Besides, our method can be made robust by having many (i.e., 4 or more) spectral bands using robust estimation techniques for conventional photometric stereo. Experiments on both synthetic and real-world scenes demonstrate the effectiveness of our method. Our data and result can be found at https://github.com/GH-HOME/MultispectralPS.git. Heng Guo 0003, Fumio Okura, Boxin Shi, Takuya Funatomi, Yasuhiro Mukaigawa, Yasuyuki Matsushita |
CVPR | 6 |
| 2021 | Lighting, Reflectance and Geometry Estimation From 360deg Panoramic StereoabstractWe propose a method for estimating high-definition spatially-varying lighting, reflectance, and geometry of a scene from 360° stereo images. Our model takes advantage of the 360° input to observe the entire scene with geometric detail, then jointly estimates the scene’s properties with physical constraints. We first reconstruct a near-field environment light for predicting the lighting at any 3D location within the scene. Then we present a deep learning model that leverages the stereo information to infer the reflectance and surface normal. Lastly, we incorporate the physical constraints between lighting and geometry to refine the reflectance of the scene. Both quantitative and qualitative experiments show that our method, benefiting from the 360° observation of the scene, outperforms prior state-of-the-art methods and enables more augmented reality applications such as mirror-objects insertion. Hongdong Li, Yasuyuki Matsushita |
CVPR | 3 |
| 2021 | Generalized Shuffled Linear RegressionabstractWe consider the shuffled linear regression problem where the correspondences between covariates and responses are unknown. While the existing formulation assumes an ideal underlying bijection in which all pieces of data should match, such an assumption barely holds in real-world applications due to either missing data or outliers. Therefore, in this work, we generalize the formulation of shuffled linear regression to a broader range of conditions where only part of the data should correspond. Moreover, we present a remarkably simple yet effective optimization algorithm with guaranteed global convergence. Distinct tasks validate the effectiveness of the proposed method.1 Feiran Li, Kent Fujiwara, Fumio Okura, Yasuyuki Matsushita |
ICCV | 4 |
| 2021 | A Closer Look at Rotation-invariant Deep Point Cloud AnalysisabstractWe consider the deep point cloud analysis tasks where the inputs of the networks are randomly rotated. Recent progress in rotation-invariant point cloud analysis is mainly driven by converting point clouds into their respective canonical poses, and principal component analysis (PCA) is a practical tool to achieve this. Due to the imperfect alignment of PCA, most of the current works are devoted to developing powerful network structures and features to overcome this deficiency, without thoroughly analyzing the PCA-based canonical poses themselves. In this work, we present a detailed study w.r.t. the PCA-based canonical poses of point clouds. Our investigation reveals that the ambiguity problem associated with the PCA-based canonical poses is handled insufficiently in some recent works. To this end, we develop a simple pose selector module for disambiguation, which presents noticeable enhancement (i.e., 5.3% classification accuracy) over state-of-the-art approaches on the challenging real-world dataset.1 Feiran Li, Kent Fujiwara, Fumio Okura, Yasuyuki Matsushita |
ICCV | 4 |
| 2021 | Toward a Unified Framework for Point Set Registration
Feiran Li, Kent Fujiwara, Yasuyuki Matsushita |
ICRA | 3 |
| 2021 | Editorial for CVIU_DL for image restoration
Jinshan Pan, Deqing Sun, Jian Yang 0003, Wangmeng Zuo, Paolo Favaro, Yasuyuki Matsushita, Ming-Hsuan Yang 0001 |
Comput. Vis. Image Underst. | 6 |
| 2021 | RotationNet for Joint Object Categorization and Unsupervised Pose Estimation from Multi-View ImagesabstractWe propose a Convolutional Neural Network (CNN)-based model "RotationNet," which takes multi-view images of an object as input and jointly estimates its pose and object category. Unlike previous approaches that use known viewpoint labels for training, our method treats the viewpoint labels as latent variables, which are learned in an unsupervised manner during the training using an unaligned object dataset. RotationNet uses only a partial set of multi-view images for inference, and this property makes it useful in practical scenarios where only partial views are available. Moreover, our pose alignment strategy enables one to obtain view-specific feature representations shared across classes, which is important to maintain high accuracy in both object categorization and pose estimation. Effectiveness of RotationNet is demonstrated by its superior performance to the state-of-the-art methods of 3D object classification on 10- and 40-class ModelNet datasets. We also show that RotationNet, even trained without known poses, achieves comparable performance to the state-of-the-art methods on an object pose estimation dataset. Furthermore, our object ranking method based on classification by RotationNet achieved the first prize in two tracks of the 3D Shape Retrieval Contest (SHREC) 2017. Finally, we demonstrate the performance of real-world applications of RotationNet trained with our newly created multi-view image dataset using a moving USB camera. Asako Kanezaki, Yasuyuki Matsushita, Yoshifumi Nishida |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Descriptor-Free Multi-view Region Matching for Instance-Wise 3D Reconstruction
Takuma Doi, Fumio Okura, Toshiki Nagahara, Yasuyuki Matsushita, Yasushi Yagi |
ACCV (5) | 4 |
| 2020 | Stereoscopic Flash and No-Flash Photography for Shape and Albedo RecoveryabstractWe present a minimal imaging setup that harnesses both geometric and photometric approaches for shape and albedo recovery. We adopt a stereo camera and a flashlight to capture a stereo image pair and a flash/no-flash pair. From the stereo image pair, we recover a rough shape that captures low-frequency shape variation without high-frequency details. From the flash/no-flash pair, we derive an image formation model for Lambertian objects under natural lighting, based on which a fine normal map is obtained and fused with the rough shape. Further, we use the flash/no-flash pair for cast shadow detection and albedo canceling, making the shape recovery robust against shadows and albedo variation. We verify the effectiveness of our approach on both synthetic and real-world data. Michael Waechter, Boxin Shi, Yasuyuki Matsushita |
CVPR | 6 |
| 2020 | Photometric Stereo via Discrete Hypothesis-and-Test SearchabstractIn this paper, we consider the problem of estimating surface normals of a scene with spatially varying, general BRDFs observed by a static camera under varying, known, distant illumination. Unlike previous approaches that are mostly based on continuous local optimization, we cast the problem as a discrete hypothesis-and-test search problem over the discretized space of surface normals. While a naive search requires a significant amount of time, we show that the expensive computation block can be precomputed in a scene-independent manner, resulting in accelerated inference for new scenes. It allows us to perform a full search over the finely discretized space of surface normals to determine the globally optimal surface normal for each scene point. We show that our method can accurately estimate surface normals of scenes with spatially varying different reflectances in a reasonable amount of time. Kenji Enomoto, Michael Waechter, Kiriakos N. Kutulakos, Yasuyuki Matsushita |
CVPR | 4 |
| 2020 | What Is Learned in Deep Uncalibrated Photometric Stereo?
Guanying Chen, Michael Waechter, Boxin Shi, Kwan-Yee Kenneth Wong, Yasuyuki Matsushita |
ECCV (14) | 5 |
| 2020 | An Analysis of Sketched IRLS for Accelerated Sparse Residual Regression
Daichi Iwata, Michael Waechter, Wen-Yan Lin, Yasuyuki Matsushita |
ECCV (12) | 4 |
| 2020 | Deep Near-Light Photometric Stereo for Spatially Varying Reflectances
Hiroaki Santo, Michael Waechter, Yasuyuki Matsushita |
ECCV (8) | 3 |
| 2020 | Reducing the amount of out-of-core data access for GPU-accelerated randomized SVDabstractSummary We propose two acceleration methods, namely, Fused and Gram, for reducing out‐of‐core data access when performing randomized singular value decomposition (RSVD) on graphics processing units (GPUs). Out‐of‐core data here are data that are too large to fit into the GPU memory at once. Both methods accelerate GPU‐enabled RSVD using the following three schemes: (1) a highly tuned general matrix‐matrix multiplication (GEMM) scheme for processing out‐of‐core data on GPUs; (2) a data‐access reduction scheme based on one‐dimensional data partition; and (3) a first‐in, first‐out scheme that reduces CPU‐GPU data transfer using the reverse iteration. The Fused method further reduces the amount of out‐of‐core data access by merging two GEMM operations into a single operation. By contrast, the Gram method reduces both in‐core and out‐of‐core data access by explicitly forming the Gram matrix. According to our experimental results, the Fused and Gram methods improved the RSVD performance up to 1.7× and 5.2×, respectively, compared with a straightforward method that deploys schemes (1) and (2) on the GPU. In addition, we present a case study of deploying the Gram method for accelerating robust principal component analysis, a convex optimization problem in machine learning. Yuechao Lu, Ichitaro Yamazaki, Fumihiko Ino, Yasuyuki Matsushita, Stanimire Tomov, Jack J. Dongarra |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | Light Structure from Pin Motion: Geometric Point Light Source CalibrationabstractAbstract We present a method for geometric point light source calibration. Unlike prior works that use Lambertian spheres, mirror spheres, or mirror planes, we use a calibration target consisting of a plane and small shadow casters at unknown positions above the plane. We show that shadow observations from a moving calibration target under a fixed light follow the principles of pinhole camera geometry and epipolar geometry, allowing joint recovery of the light position and 3D shadow caster positions, equivalent to how conventional structure from motion jointly recovers camera parameters and 3D feature positions from observed 2D features. Moreover, we devised a unified light model that works with nearby point lights as well as distant light in one common framework. Our evaluation shows that our method yields light estimates that are stable and more accurate than existing techniques while having a much simpler setup and requiring less manual labor. Hiroaki Santo, Michael Waechter, Wen-Yan Lin, Yusuke Sugano, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 5 |
| 2020 | Semi-Calibrated Photometric StereoabstractWhile conventional calibrated photometric stereo methods assume that light intensities and sensor exposures are known or unknown but identical across observed images, this assumption easily breaks down in practical settings due to individual light bulb's characteristics and limited control over sensors. This paper studies the effect of unknown and possibly non-uniform light intensities and sensor exposures among observed images on the shape recovery based on photometric stereo. This leads to the development of a "semi-calibrated" photometric stereo method, where the light directions are known but light intensities (and sensor exposures) are unknown. We show that the semi-calibrated photometric stereo becomes a bilinear problem, whose general form is difficult to solve, but in the photometric stereo context, there exists a unique solution for the surface normal and light intensities (or sensor exposures). We further show that there exists a linear solution method for the problem, and develop efficient and stable solution methods. The semi-calibrated photometric stereo is advantageous over conventional calibrated photometric stereo in accurate determination of surface normal, because it relaxes the assumption of known light intensity ratios/sensor exposures. The experimental results show superior accuracy of the semi-calibrated photometric stereo in comparison to conventional methods in practical settings. Donghyeon Cho, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Ambiguity-Free Radiometric Calibration for Internet Photo CollectionsabstractRadiometrically calibrating nonlinear images from Internet photo collections makes photometric analysis applicable not only to lab data but also to big image data in the wild. However, conventional calibration methods cannot be directly applied to such photo collections. This paper presents a method to jointly perform radiometric calibration for a set of nonlinear images in Internet photo collections. By incorporating the consistency of scene reflectance of corresponding pixels across nonlinear images, the proposed method first estimates radiometric response functions of all the nonlinear images up to a unique exponential ambiguity using a rank minimization framework. The ambiguity is then resolved using the linear edge color blending constraint. Quantitative evaluation using both synthetic and real-world data shows the effectiveness of the proposed method. Zhipeng Mo, Boxin Shi, Sai-Kit Yeung, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Self-Calibrating Deep Photometric Stereo NetworksabstractThis paper proposes an uncalibrated photometric stereo method for non-Lambertian scenes based on deep learning. Unlike previous approaches that heavily rely on assumptions of specific reflectances and light source distributions, our method is able to determine both shape and light directions of a scene with unknown arbitrary reflectances observed under unknown varying light directions. To achieve this goal, we propose a two-stage deep learning architecture, called SDPS-Net, which can effectively take advantage of intermediate supervision, resulting in reduced learning difficulty compared to a single-stage model. Experiments on both synthetic and real datasets show that our proposed approach significantly outperforms previous uncalibrated photometric stereo methods. Guanying Chen, Kai Han 0001, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee Kenneth Wong |
CVPR | 4 |
| 2019 | Learning to Minify Photometric StereoabstractPhotometric stereo estimates the surface normal given a set of images acquired under different illumination conditions. To deal with diverse factors involved in the image formation process, recent photometric stereo methods demand a large number of images as input. We propose a method that can dramatically decrease the demands on the number of images by learning the most informative ones under different illumination conditions. To this end, we use a deep learning framework to automatically learn the critical illumination conditions required at input. Furthermore, we present an occlusion layer that can synthesize cast shadows, which effectively improves the estimation accuracy. We assess our method on challenging real-world conditions, where we outperform techniques elsewhere in the literature with a significantly reduced number of light conditions. Antonio Robles-Kelly, Shaodi You, Yasuyuki Matsushita |
CVPR | 4 |
| 2019 | Material Classification from Time-of-Flight DistortionsabstractThis paper presents a material classification method using an off-the-shelf Time-of-Flight (ToF) camera. The proposed method is built upon a key observation that the depth measurement by a ToF camera is distorted for objects with certain materials, especially with translucent materials. We show that this distortion is due to the variation of time domain impulse responses across materials and also due to the measurement mechanism of the ToF cameras. Specifically, we reveal that the amount of distortion varies according to the modulation frequency of the ToF camera, the object material, and the distance between the camera and object. Our method uses the depth distortion of ToF measurements as a feature for classification and achieves material classification of a scene. Effectiveness of the proposed method is demonstrated by numerical evaluations and real-world experiments, showing its capability of material classification, even for visually indistinguishable objects. Kenichiro Tanaka, Yasuhiro Mukaigawa, Takuya Funatomi, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Shape-Conditioned Image Generation by Learning Latent Appearance Representation from Unpaired Data
Yutaro Miyauchi, Yusuke Sugano, Yasuyuki Matsushita |
ACCV (6) | 3 |
| 2018 | Probabilistic Plant Modeling via Multi-View Image-to-Image TranslationabstractThis paper describes a method for inferring three-dimensional (3D) plant branch structures that are hidden under leaves from multi-view observations. Unlike previous geometric approaches that heavily rely on the visibility of the branches or use parametric branching models, our method makes statistical inferences of branch structures in a probabilistic framework. By inferring the probability of branch existence using a Bayesian extension of image-to-image translation applied to each of multi-view images, our method generates a probabilistic plant 3D model, which represents the 3D branching pattern that cannot be directly observed. Experiments demonstrate the usefulness of the proposed approach in generating convincing branch structures in comparison to prior approaches. Takahiro Isokane, Fumio Okura, Ayaka Ide, Yasuyuki Matsushita, Yasushi Yagi |
CVPR | 4 |
| 2018 | RotationNet: Joint Object Categorization and Pose Estimation Using Multiviews From Unsupervised ViewpointsabstractWe propose a Convolutional Neural Network (CNN)-based model "RotationNet," which takes multi-view images of an object as input and jointly estimates its pose and object category. Unlike previous approaches that use known viewpoint labels for training, our method treats the viewpoint labels as latent variables, which are learned in an unsupervised manner during the training using an unaligned object dataset. RotationNet is designed to use only a partial set of multi-view images for inference, and this property makes it useful in practical scenarios where only partial views are available. Moreover, our pose alignment strategy enables one to obtain view-specific feature representations shared across classes, which is important to maintain high accuracy in both object categorization and pose estimation. Effectiveness of RotationNet is demonstrated by its superior performance to the state-of-the-art methods of 3D object classification on 10- and 40-class ModelNet datasets. We also show that RotationNet, even trained without known poses, achieves the state-of-the-art performance on an object pose estimation dataset. Asako Kanezaki, Yasuyuki Matsushita, Yoshifumi Nishida |
CVPR | 2 |
| 2018 | Dimensionality's Blessing: Clustering Images by Underlying DistributionabstractMany high dimensional vector distances tend to a constant. This is typically considered a negative "contrast-loss" phenomenon that hinders clustering and other machine learning techniques. We reinterpret "contrast-loss" as a blessing. Re-deriving "contrast-loss" using the law of large numbers, we show it results in a distribution's instances concentrating on a thin "hyper-shell". The hollow center means apparently chaotically overlapping distributions are actually intrinsically separable. We use this to develop distribution-clustering, an elegant algorithm for grouping of data points by their (unknown) underlying distribution. Distribution-clustering, creates notably clean clusters from raw unlabeled data, estimates the number of clusters for itself and is inherently robust to "outliers" which form their own clusters. This enables trawling for patterns in unorganized data and may be the key to enabling machine intelligence. Wen-Yan Lin, Jian-Huang Lai, Yasuyuki Matsushita |
CVPR | 4 |
| 2018 | Uncalibrated Photometric Stereo Under Natural IlluminationabstractThis paper presents a photometric stereo method that works with unknown natural illuminations without any calibration object. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary rotation ambiguity. Our method connects the resulting patches and unifies the local ambiguities to a global rotation one through angular distance propagation defined over the whole surface. After applying the integrability constraint, our final solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods. Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Yasuyuki Matsushita |
CVPR | 5 |
| 2018 | Light Structure from Pin Motion: Simple and Accurate Point Light Calibration for Physics-Based Modeling
Hiroaki Santo, Michael Waechter, Masaki Samejima, Yusuke Sugano, Yasuyuki Matsushita |
ECCV (3) | 5 |
| 2018 | Guest Editorial: Vision and Computational Photography and Graphics
Radu Timofte, Luc Van Gool, Ming-Hsuan Yang 0001, Shai Avidan, Yasuyuki Matsushita, Qingxiong Yang |
Comput. Vis. Image Underst. | 5 |
| 2018 | Fast Randomized Singular Value Thresholding for Low-Rank OptimizationabstractRank minimization can be converted into tractable surrogate problems, such as Nuclear Norm Minimization (NNM) and Weighted NNM (WNNM). The problems related to NNM, or WNNM, can be solved iteratively by applying a closed-form proximal operator, called Singular Value Thresholding (SVT), or Weighted SVT, but they suffer from high computational cost of Singular Value Decomposition (SVD) at each iteration. We propose a fast and accurate approximation method for SVT, that we call fast randomized SVT (FRSVT), with which we avoid direct computation of SVD. The key idea is to extract an approximate basis for the range of the matrix from its compressed matrix. Given the basis, we compute partial singular values of the original matrix from the small factored matrix. In addition, by developping a range propagation method, our method further speeds up the extraction of approximate basis at each iteration. Our theoretical analysis shows the relationship between the approximation bound of SVD and its effect to NNM via SVT. Along with the analysis, our empirical results quantitatively and qualitatively show that our approximation rarely harms the convergence of the host algorithms. We assess the efficiency and accuracy of the proposed method on various computer vision problems, e.g., subspace clustering, weather artifact removal, and simultaneous multi-image alignment and rectification. Tae-Hyun Oh, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Continuous 3D Label Stereo Matching Using Local Expansion MovesabstractWe present an accurate stereo matching method using local expansion moves based on graph cuts. This new move-making scheme is used to efficiently infer per-pixel 3D plane labels on a pairwise Markov random field (MRF) that effectively combines recently proposed slanted patch matching and curvature regularization terms. The local expansion moves are presented as many -expansions defined for small grid regions. The local expansion moves extend traditional expansion moves by two ways: localization and spatial propagation. By localization, we use different candidate -labels according to the locations of local -expansions. By spatial propagation, we design our local -expansions to propagate currently assigned labels for nearby regions. With this localization and spatial propagation, our method can efficiently infer MRF models with a continuous label space using randomized search. Our method has several advantages over previous approaches that are based on fusion moves or belief propagation; it produces submodular moves deriving a subproblem optimality; it helps find good, smooth, piecewise linear disparity maps; it is suitable for parallelization; it can use cost-volume filtering techniques for accelerating the matching cost computations. Even using a simple pairwise MRF, our method is shown to have best performance in the Middlebury stereo benchmark V2 and V3. Tatsunori Taniai, Yasuyuki Matsushita, Yoichi Sato 0001, Takeshi Naemura |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Multiview Rectification of Folded DocumentsabstractDigitally unwrapping images of paper sheets is crucial for accurate document scanning and text recognition. This paper presents a method for automatically rectifying curved or folded paper sheets from a few images captured from multiple viewpoints. Prior methods either need expensive 3D scanners or model deformable surfaces using over-simplified parametric representations. In contrast, our method uses regular images and is based on general developable surface models that can represent a wide variety of paper deformations. Our main contribution is a new robust rectification method based on ridge-aware 3D reconstruction of a paper sheet and unwrapping the reconstructed surface using properties of developable surfaces via conformal mapping. We present results on several examples including book pages, folded letters and shopping receipts. Shaodi You, Yasuyuki Matsushita, Sudipta N. Sinha, Yusuke Bou, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | GMS: Grid-Based Motion Statistics for Fast, Ultra-Robust Feature CorrespondenceabstractIncorporating smoothness constraints into feature matching is known to enable ultra-robust matching. However, such formulations are both complex and slow, making them unsuitable for video applications. This paper proposes GMS (Grid-based Motion Statistics), a simple means of encapsulating motion smoothness as the statistical likelihood of a certain number of matches in a region. GMS enables translation of high match numbers into high match quality. This provides a real-time, ultra-robust correspondence system. Evaluation on videos, with low textures, blurs and wide-baselines show GMS consistently out-performs other real-time matchers and can achieve parity with more sophisticated, much slower techniques. Jiawang Bian, Wen-Yan Lin, Yasuyuki Matsushita, Sai-Kit Yeung, Tan-Dat Nguyen, Ming-Ming Cheng |
CVPR | 3 |
| 2017 | Radiometric Calibration for Internet Photo CollectionsabstractRadiometrically calibrating the images from Internet photo collections brings photometric analysis from lab data to big image data in the wild, but conventional calibration methods cannot be directly applied to such image data. This paper presents a method to jointly perform radiometric calibration for a set of images in an Internet photo collection. By incorporating the consistency of scene reflectance for corresponding pixels in multiple images, the proposed method estimates radiometric response functions of all the images using a rank minimization framework. Our calibration aligns all response functions in an image set up to the same exponential ambiguity in a robust manner. Quantitative results using both synthetic and real data show the effectiveness of the proposed method. Zhipeng Mo, Boxin Shi, Sai-Kit Yeung, Yasuyuki Matsushita |
CVPR | 4 |
| 2017 | Material Classification Using Frequency-and Depth-Dependent Time-of-Flight DistortionabstractThis paper presents a material classification method using an off-the-shelf Time-of-Flight (ToF) camera. We use a key observation that the depth measurement by a ToF camera is distorted in objects with certain materials, especially with translucent materials. We show that this distortion is caused by the variations of time domain impulse responses across materials and also by the measurement mechanism of the existing ToF cameras. Specifically, we reveal that the amount of distortion varies according to the modulation frequency of the ToF camera, the material of the object, and the distance between the camera and object. Our method uses the depth distortion of ToF measurements as features and achieves material classification of a scene. Effectiveness of the proposed method is demonstrated by numerical evaluation and real-world experiments, showing its capability of even classifying visually similar objects. Kenichiro Tanaka, Yasuhiro Mukaigawa, Takuya Funatomi, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi |
CVPR | 5 |
| 2017 | Device-free and privacy preserving indoor positioning using infrared retro-reflection imagingabstractIndoor positioning is a core technology for indoor daily life applications such as elderly care and home automation. This study presents a device-free and privacy preserving indoor positioning method using infrared (IR) cameras and retroreflectors. The proposed method uses IR cameras equipped with IR LEDs to capture retroreflections from markers attached to walls in the environment, and detects a person who passes between the camera and a marker by observing the occlusion. Because our method employs occlusion of markers, it can track a person without attaching tags to the person. Also, our camera device permits us to filter out visible light and thus the appearance of the person is not recorded. Our evaluation in real environments showed that our method achieved an average positioning error of about 0.3 meters. Hiroaki Santo, Takuya Maekawa, Yasuyuki Matsushita |
PerCom | 3 |
| 2017 | Robust Multiview Photometric Stereo Using Planar Mesh ParameterizationabstractWe propose a robust uncalibrated multiview photometric stereo method for high quality 3D shape reconstruction. In our method, a coarse initial 3D mesh obtained using a multiview stereo method is projected onto a 2D planar domain using a planar mesh parameterization technique. We describe methods for surface normal estimation that work in the parameterized 2D space that jointly incorporates all geometric and photometric cues from multiple viewpoints. Using an estimated surface normal map, a refined 3D mesh is then recovered by computing an optimal displacement map in the same 2D planar domain. Our method avoids the need of merging view-dependent surface normal maps that is often required in conventional methods. We conduct evaluation on various real-world objects containing surfaces with specular reflections, multiple albedos, and complex topologies in both controlled and uncontrolled settings and demonstrate that accurate 3D meshes with fine geometric details can be recovered by our method. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Recovering Inner Slices of Layered Translucent Objects by Multi-Frequency IlluminationabstractThis paper describes a method for recovering appearance of inner slices of translucent objects. The appearance of a layered translucent object is the summed appearance of all layers, where each layer is blurred by a depth-dependent point spread function (PSF). By exploiting the difference of low-pass characteristics of depth-dependent PSFs, we develop a multi-frequency illumination method for obtaining the appearance of individual inner slices. Specifically, by observing the target object with varying the spatial frequency of checker-pattern illumination, our method recovers the appearance of inner slices via computation. We study the effect of non-uniform transmission due to inhomogeneity of translucent objects and develop a method for recovering clear inner slices based on the pixel-wise PSF estimates under the assumption of spatial smoothness of inner slice appearances. We quantitatively evaluate the accuracy of the proposed method by simulations and qualitatively show faithful recovery using real-world scenes. Kenichiro Tanaka, Yasuhiro Mukaigawa, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | A Holistic Approach to Cross-Channel Image Noise Modeling and Its Application to Image DenoisingabstractModelling and analyzing noise in images is a fundamental task in many computer vision systems. Traditionally, noise has been modelled per color channel assuming that the color channels are independent. Although the color channels can be considered as mutually independent in camera RAW images, signals from different color channels get mixed during the imaging process inside the camera due to gamut mapping, tone-mapping, and compression. We show the influence of the in-camera imaging pipeline on noise and propose a new noise model in the 3D RGB space to accounts for the color channel mix-ups. A data-driven approach for determining the parameters of the new noise model is introduced as well as its application to image denoising. The experiments show that our noise model represents the noise in regular JPEG images more accurately compared to the previous models and is advantageous in image denoising. Seonghyeon Nam, Youngbae Hwang, Yasuyuki Matsushita, Seon Joo Kim |
CVPR | 3 |
| 2016 | Recovering Transparent Shape from Time-of-Flight DistortionabstractThis paper presents a method for recovering shape and normal of a transparent object from a single viewpoint using a Time-of-Flight (ToF) camera. Our method is built upon the fact that the speed of light varies with the refractive index of the medium and therefore the depth measurement of a transparent object with a ToF camera may be distorted. We show that, from this ToF distortion, the refractive light path can be uniquely determined by estimating a single parameter. We estimate this parameter by introducing a surface normal consistency between the one determined by a light path candidate and the other computed from the corresponding shape. The proposed method is evaluated by both simulation and real-world experiments and shows faithful transparent shape recovery. Kenichiro Tanaka, Yasuhiro Mukaigawa, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi |
CVPR | 4 |
| 2016 | Photometric Stereo Under Non-uniform Light Intensities and Exposures
Donghyeon Cho, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
ECCV (2) | 2 |
| 2016 | Predicting location semantics combining active and passive sensing with environment-independent classifierabstractThis paper presents a method for estimating a user's indoor location without using training data collected by the user in his/her environment. Specifically, we attempt to predict the user's location semantics, i.e., location classes such as restroom and meeting room. While indoor location information can be used in many real-world services, e.g., context-aware systems, lifelogging, and monitoring the elderly, estimating the location information requires training data collected in an environment of interest. In this study, we combine passive sensing and active sound probing to capture and learn inherent sensor data features for each location class using labeled training data collected in other environments. In addition, this study modifies the random forest algorithm to effectively extract inherent sensor data features for each location class. Our evaluation showed that our method achieved about 85% accuracy without using training data collected in test environments. Masaya Tachikawa, Takuya Maekawa, Yasuyuki Matsushita |
UbiComp | 3 |
| 2016 | A Pseudo-Bayesian Algorithm for Robust PCAabstractCommonly used in many applications, robust PCA represents an algorithmic attempt to reduce the sensitivity of classical PCA to outliers. The basic idea is to learn a decomposition of some data matrix of interest into low rank and sparse components, the latter representing unwanted outliers. Although the resulting problem is typically NP-hard, convex relaxations provide a computationally-expedient alternative with theoretical support. However, in practical regimes performance guarantees break down and a variety of non-convex alternatives, including Bayesian-inspired models, have been proposed to boost estimation quality. Unfortunately though, without additional a priori knowledge none of these methods can significantly expand the critical operational range such that exact principal subspace recovery is possible. Into this mix we propose a novel pseudo-Bayesian algorithm that explicitly compensates for design weaknesses in many existing non-convex approaches leading to state-of-the-art performance with a sound analytical foundation. Tae-Hyun Oh, Yasuyuki Matsushita, In-So Kweon, David P. Wipf |
NIPS | 2 |
| 2016 | Bayesian Depth-From-Defocus With Shading ConstraintsabstractWe present a method that enhances the performance of depth-from-defocus (DFD) through the use of shading information. DFD suffers from important limitations--namely coarse shape reconstruction and poor accuracy on textureless surfaces--that can be overcome with the help of shading. We integrate both forms of data within a Bayesian framework that capitalizes on their relative strengths. Shading data, however, is challenging to accurately recover from surfaces that contain texture. To address this issue, we propose an iterative technique that utilizes depth information to improve shading estimation, which in turn is used to elevate depth estimation in the presence of textures. The shading estimation can be performed in general scenes with unknown illumination using an approximate estimate of scene lighting. With this approach, we demonstrate improvements over existing DFD techniques, as well as effective shape reconstruction of textureless surfaces. Chen Li 0031, Shuochen Su, Yasuyuki Matsushita, Kun Zhou 0001, Stephen Lin 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Continuous Symmetric Stereo with Adaptive Outlier HandlingabstractWe present a method for symmetric stereo matching in which outliers from occlusions, texture-less regions, and repeated patterns are handled in a soft and adaptive manner. Rather than making binary outlier decisions, our model incorporates continuous-valued confidence weights that account for outlier likelihood, to promote robustness in disparity estimation. In contrast to previous outlier labeling techniques that fix the labels at the start of optimization, our method iteratively updates our outlier confidence weights as the matching results are gradually refined. By doing this, errors in an initial labeling can be rectified in the matching process. Our model is optimized in an Expectation-Maximization framework that efficiently produces continuous disparity estimates. This approach provides a good combination of accuracy and speed. Experiments show that our method compares favorably to prior outlier labeling techniques on the Middlebury benchmark, and that it can generate high-quality reconstruction for outdoor images with much more complex occlusions. Chen Li 0031, Lap-Fai Yu, Zhichao Lu, Yasuyuki Matsushita, Kun Zhou 0001, Stephen Lin 0001 |
3DV | 4 |
| 2015 | Efficient Large-Scale Point Cloud Registration Using Loop ClosuresabstractAlignment of many 3D point clouds, possibly captured by multiple devices at different times, is a critical step for increasingly popular applications such as 3D model construction and augmented reality. For very large data sets, traditional methods such as ICP can become computationally intractable, or produce poor results. We present an efficient method for accurately aligning very large numbers of dense 3D point clouds, and apply it to a city-scale data set. The method relies on the novel combination of 1) partitioning the point clouds based on loop structures detected across a combined network of all device capture paths, and 2) making use of the loop closure property to accurately align point clouds within each sub-problem. Final global alignment of the loop-based results is formulated as a least squares optimization with closed form solution. Experimental results are shown for aligning 3D points across the entire city of San Francisco with centimeter-scale accuracy, via an efficient parallelized architecture. Takaaki Shiratori, Jérôme Berclaz, Michael Harville, Chintan Shah, Taoyu Li, Yasuyuki Matsushita, Stephen Shiller |
3DV | 6 |
| 2015 | Fast randomized Singular Value Thresholding for Nuclear Norm MinimizationabstractRank minimization problem can be boiled down to either Nuclear Norm Minimization (NNM) or Weighted NNM (WNNM) problem. The problems related to NNM (or WNNM) can be solved iteratively by applying a closed-form proximal operator, called Singular Value Thresholding (SVT) (or Weighted SVT), but they suffer from high computational cost to compute a Singular Value Decomposition (SVD) at each iteration. In this paper, we propose an accurate and fast approximation method for SVT, called fast randomized SVT (FRSVT), where we avoid direct computation of SVD. The key idea is to extract an approximate basis for the range of a matrix from its compressed matrix. Given the basis, we compute the partial singular values of the original matrix from a small factored matrix. While the basis approximation is the bottleneck, our method is already severalfold faster than thin SVD. By adopting a range propagation technique, we can further avoid one of the bottleneck at each iteration. Our theoretical analysis provides a stepping stone between the approximation bound of SVD and its effect to NNM via SVT. Along with the analysis, our empirical results on both quantitative and qualitative studies show our approximation rarely harms the convergence behavior of the host algorithms. We apply it and validate the efficiency of our method on various vision problems, e.g. subspace clustering, weather artifact removal, simultaneous multi-image alignment and rectification. Tae-Hyun Oh, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
CVPR | 2 |
| 2015 | Recovering inner slices of translucent objects by multi-frequency illuminationabstractThis paper describes a method for recovering appearance of inner slices of translucent objects. The outer appearance of translucent objects is a summation of the appearance of slices at all depths, where each slice is blurred by depth-dependent point spread functions (PSFs). By exploiting the difference of low-pass characteristics of depth-dependent PSFs, we develop a multi-frequency illumination method for obtaining the appearance of individual inner slices using a coaxial projector-camera setup. Specifically, by measuring the target object with varying the spatial frequency of checker patterns emitted from a projector, our method recovers inner slices via a simple linear solution method. We quantitatively evaluate accuracy of the proposed method by simulations and show qualitative recovery results using real-world scenes. Kenichiro Tanaka, Yasuhiro Mukaigawa, Hiroyuki Kubo, Yasuyuki Matsushita, Yasushi Yagi |
CVPR | 4 |
| 2015 | Superdifferential cuts for binary energiesabstractWe propose an efficient and general purpose energy optimization method for binary variable energies used in various low-level vision tasks. Our method can be used for broad classes of higher-order and pairwise non-submodular functions. We first revisit a submodular-supermodular procedure (SSP) [19], which is previously studied for higher-order energy optimization. We then present our method as generalization of SSP, which is further shown to generalize several state-of-the-art techniques for higher-order and pairwise non-submodular functions [2, 9, 25]. In the experiments, we apply our method to image segmentation, deconvolution, and binarization, and show improvements over state-of-the-art methods. Tatsunori Taniai, Yasuyuki Matsushita, Takeshi Naemura |
CVPR | 2 |
| 2015 | Photometric Stereo with Small Angular VariationsabstractMost existing successful photometric stereo setups require large angular variations in illumination directions, which results in acquisition rigs that have large spatial extent. For many applications, especially involving mobile devices, it is important that the device be spatially compact. This naturally implies smaller angular variations in the illumination directions. This paper studies the effect of small angular variations in illumination directions to photometric stereo. We explore both theoretical justification and practical issues in the design of a compact and portable photometric stereo device on which a camera is surrounded by a ring of point light sources. We first derive the relationship between the estimation error of surface normal and the baseline of the point light sources. Armed with this theoretical insight, we develop a small baseline photometric stereo prototype to experimentally examine the theory and its practicality. Jian Wang 0100, Yasuyuki Matsushita, Boxin Shi, Aswin C. Sankaranarayanan |
ICCV | 2 |
| 2015 | Photometric Stereo in the WildabstractConventional photometric stereo requires to capture images or videos in a dark room to obstruct complex environment light as much as possible. This paper presents a new method that capitalizes on environment light to avail geometry reconstruction, thus bringing photometric stereo to the wild, such as an outdoor scene, with uncontrolled lighting. We do not make restrictive assumption, and only use simple capture equipments, which include a mirror sphere and a video camera. Qualitative and quantitative experiments indicate the potential and practicality of our system to generalize existing frameworks. Chun Ho Hung, Tai-Pang Wu, Yasuyuki Matsushita, Li Xu 0001, Jiaya Jia, Chi-Keung Tang |
WACV | 3 |
| 2015 | From Intensity Profile to Surface Normal: Photometric Stereo for Unknown Light Sources and Isotropic ReflectancesabstractWe propose an uncalibrated photometric stereo method that works with general and unknown isotropic reflectances. Our method uses a pixel intensity profile, which is a sequence of radiance intensities recorded at a pixel under unknown varying directional illumination. We show that for general isotropic materials and uniformly distributed light directions, the geodesic distance between intensity profiles is linearly related to the angular difference of their corresponding surface normals, and that the intensity distribution of the intensity profile reveals reflectance properties. Based on these observations, we develop two methods for surface normal estimation; one for a general setting that uses only the recorded intensity profiles, the other for the case where a BRDF database is available while the exact BRDF of the target scene is still unknown. Quantitative and qualitative evaluations are conducted using both synthetic and real-world scenes, which show the state-of-the-art accuracy of smaller than 10 degree without using reference data and 5 degree with reference data for all 100 materials in MERL database. Feng Lu 0005, Yasuyuki Matsushita, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Appearance-Based Gaze Estimation With Online Calibration From Mouse OperationsabstractThis paper presents an unconstrained gaze estimation method using an online learning algorithm. We focus on a desktop scenario, where a user operates a personal computer, and use the mouse-clicked positions to infer, where on the screen the user is looking at. Our method continuously captures the user's head pose and eye images with a monocular camera, and each mouse click triggers learning sample acquisition. In order to handle head pose variations, the samples are adaptively clustered according to the estimated head pose. Then, local reconstruction-based gaze estimation models are incrementally updated in each cluster. We conducted a prototype evaluation in real-world environments, and our method achieved an estimation accuracy of 2.9°. Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001, Hideki Koike |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2014 | Efficient Colorization of Large-Scale Point Cloud Using Multi-pass Z-OrderingabstractWe present an efficient colorization method for a large scale point cloud using multi-view images. To address the practical issues of noisy camera parameters and color inconsistencies across multi-view images, our method takes an optimization approach for achieving visually pleasing point cloud colorization. We introduce a multi-pass Z-ordering technique that efficiently defines a graph structure to a large-scale and un-ordered set of 3D points, and use the graph structure for optimizing the point colors to be assigned. Our technique is useful for defining minimal but sufficient connectivities among 3D points so that the optimization can exploit the sparsity for efficiently solving the problem. We demonstrate the effectiveness of our method using synthetic datasets and a large-scale real-world data in comparison with other graph construction techniques. Sunyoung Cho, Jizhou Yan, Yasuyuki Matsushita, Hyeran Byun |
3DV | 3 |
| 2014 | Photometric Stereo Using Internet ImagesabstractPhotometric stereo using unorganized Internet images is very challenging, because the input images are captured under unknown general illuminations, with uncontrolled cameras. We propose to solve this difficult problem by a simple yet effective approach that makes use of a coarse shape prior. The shape prior is obtained from multi-view stereo and will be useful in twofold: resolving the shape-light ambiguity in uncalibrated photometric stereo and guiding the estimated normals to produce the high quality 3D surface. By assuming the surface albedo is not highly contrasted, we also propose a novel linear approximation of the nonlinear camera responses with our normal estimation algorithm. We evaluate our method using synthetic data and demonstrate the surface improvement on real data over multi-view stereo results. Boxin Shi, Kenji Inose, Yasuyuki Matsushita, Ping Tan 0002, Sai-Kit Yeung, Katsushi Ikeuchi |
3DV | 3 |
| 2014 | Efficient Multiview Stereo by Random-Search and PropagationabstractWe present an efficient multi-view 3D reconstruction method based on randomization and propagation scheme. Our method progressively refines 3D point estimates by randomly perturbing the initial guess of 3D points and propagates photo-consistent ones to their neighbors. In contrast to previous refinement methods that perform local optimization for a better photo-consistency, our randomization approach takes lucky matchings for reducing the computational complexity. Experiments show favorable efficiency of the proposed method with the accuracy that is close to the state-of-the-art methods. Youngjung Uh, Yasuyuki Matsushita, Hyeran Byun |
3DV | 2 |
| 2014 | Calibrating a Non-isotropic Near Point Light Source Using a PlaneabstractWe show that a non-isotropic near point light source rigidly attached to a camera can be calibrated using multiple images of a weakly textured planar scene. We prove that if the radiant intensity distribution (RID) of a light source is radially symmetric with respect to its dominant direction, then the shading observed on a Lambertian scene plane is bilaterally symmetric with respect to a 2D line on the plane. The symmetry axis detected in an image provides a linear constraint for estimating the dominant light axis. The light position and RID parameters can then be estimated using a linear method. Specular highlights if available can also be used for light position estimation. We also extend our method to handle non-Lambertian reflectances which we model using a biquadratic BRDF. We have evaluated our method on synthetic data quantitavely. Our experiments on real scenes show that our method works well in practice and enables light calibration without the need of a specialized hardware. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
CVPR | 3 |
| 2014 | Learning-by-Synthesis for Appearance-Based 3D Gaze EstimationabstractInferring human gaze from low-resolution eye images is still a challenging task despite its practical importance in many application scenarios. This paper presents a learning-by-synthesis approach to accurate image-based gaze estimation that is person- and head pose-independent. Unlike existing appearance-based methods that assume person-specific training data, we use a large amount of cross-subject training data to train a 3D gaze estimator. We collect the largest and fully calibrated multi-view gaze dataset and perform a 3D reconstruction in order to generate dense training data of eye images. By using the synthesized dataset to learn a random regression forest, we show that our method outperforms existing methods that use low-resolution eye images. Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001 |
CVPR | 2 |
| 2014 | Graph Cut Based Continuous Stereo Matching Using Locally Shared LabelsabstractWe present an accurate and efficient stereo matching method using locally shared labels, a new labeling scheme that enables spatial propagation in MRF inference using graph cuts. They give each pixel and region a set of candidate disparity labels, which are randomly initialized, spatially propagated, and refined for continuous disparity estimation. We cast the selection and propagation of locally-defined disparity labels as fusion-based energy minimization. The joint use of graph cuts and locally shared labels has advantages over previous approaches based on fusion moves or belief propagation, it produces submodular moves deriving a subproblem optimality, enables powerful randomized search, helps to find good smooth, locally planar disparity maps, which are reasonable for natural scenes, allows parallel computation of both unary and pairwise costs. Our method is evaluated using the Middlebury stereo benchmark and achieves first place in sub-pixel accuracy. Tatsunori Taniai, Yasuyuki Matsushita, Takeshi Naemura |
CVPR | 2 |
| 2014 | Interreflection Removal Using Fluorescence
Ying Fu 0001, Antony Lam, Yasuyuki Matsushita, Imari Sato, Yoichi Sato 0001 |
ECCV (5) | 3 |
| 2014 | Surface Normal Deconvolution: Photometric Stereo for Optically Thick Translucent Objects
Chika Inoshita, Yasuhiro Mukaigawa, Yasuyuki Matsushita, Yasushi Yagi |
ECCV (2) | 3 |
| 2014 | Photometric Stereo Using Sparse Bayesian Regression for General Diffuse SurfacesabstractMost conventional algorithms for non-Lambertian photometric stereo can be partitioned into two categories. The first category is built upon stable outlier rejection techniques while assuming a dense Lambertian structure for the inliers, and thus performance degrades when general diffuse regions are present. The second utilizes complex reflectance representations and non-linear optimization over pixels to handle non-Lambertian surfaces, but does not explicitly account for shadows or other forms of corrupting outliers. In this paper, we present a purely pixel-wise photometric stereo method that stably and efficiently handles various non-Lambertian effects by assuming that appearances can be decomposed into a sparse, non-diffuse component (e.g., shadows, specularities, etc.) and a diffuse component represented by a monotonic function of the surface normal and lighting dot-product. This function is constructed using a piecewise linear approximation to the inverse diffuse model, leading to closed-form estimates of the surface normals and model parameters in the absence of non-diffuse corruptions. The latter are modeled as latent variables embedded within a hierarchical Bayesian model such that we may accurately compute the unknown surface normals while simultaneously separating diffuse from non-diffuse components. Extensive evaluations are performed that show state-of-the-art performance using both synthetic and real-world images. Satoshi Ikehata, David P. Wipf, Yasuyuki Matsushita, Kiyoharu Aizawa |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Bi-Polynomial Modeling of Low-Frequency ReflectancesabstractWe present a bi-polynomial reflectance model that can precisely represent the low-frequency component of reflectance. Most existing reflectance models aim at accurately representing the complete reflectance domain for photo-realistic rendering purposes. In contrast, our bi-polynomial model is developed for the purpose of accurately solving inverse problems by effectively discarding the high-frequency component while retaining nonlinear variations in the low-frequency part. The bi-polynomial reflectance model is useful for estimating reflectance and shape of an object. Experimental evaluation in comparison with other parametric reflectance models demonstrates that the proposed model achieves better performance in reflectometry and photometric stereo applications. Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Bayesian Depth-from-Defocus with Shading ConstraintsabstractWe present a method that enhances the performance of depth-from-defocus (DFD) through the use of shading information. DFD suffers from important limitations - namely coarse shape reconstruction and poor accuracy on texture less surfaces - that can be overcome with the help of shading. We integrate both forms of data within a Bayesian framework that capitalizes on their relative strengths. Shading data, however, is challenging to recover accurately from surfaces that contain texture. To address this issue, we propose an iterative technique that utilizes depth information to improve shading estimation, which in turn is used to elevate depth estimation in the presence of textures. With this approach, we demonstrate improvements over existing DFD techniques, as well as effective shape reconstruction of texture less surfaces. Chen Li 0031, Shuochen Su, Yasuyuki Matsushita, Kun Zhou 0001, Stephen Lin 0001 |
CVPR | 3 |
| 2013 | Uncalibrated Photometric Stereo for Unknown Isotropic ReflectancesabstractWe propose an uncalibrated photometric stereo method that works with general and unknown isotropic reflectances. Our method uses a pixel intensity profile, which is a sequence of radiance intensities recorded at a pixel across multi-illuminance images. We show that for general isotropic materials, the geodesic distance between intensity profiles is linearly related to the angular difference of their surface normals, and that the intensity distribution of an intensity profile conveys information about the reflectance properties, when the intensity profile is obtained under uniformly distributed directional lightings. Based on these observations, we show that surface normals can be estimated up to a convex/concave ambiguity. A solution method based on matrix decomposition with missing data is developed for a reliable estimation. Quantitative and qualitative evaluations of our method are performed using both synthetic and real-world scenes. Feng Lu 0005, Yasuyuki Matsushita, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
CVPR | 2 |
| 2013 | Descattering of transmissive observation using Parallel High-Frequency IlluminationabstractThe inner structures of an object can be measured by capturing transmissive images. However, the recorded images of a translucent object tend to be unclear due to strong scattering of light inside the object. In this paper, we propose a descattering approach based on Parallel High-frequency Illumination. We show in this paper that the original high-frequency illumination method and the various extended techniques can be uniformly defined as a separation of overlapped and non-overlapped light rays. Also, we show that transmissive light rays do not overlap each other by constructing a parallel projection/measurement system for performing both illumination and observation. We have developed a measurement system that consists of a camera and projector with telecentric lenses and have evaluated descattering effects by extracting transmissive light rays. Kenichiro Tanaka, Yasuhiro Mukaigawa, Yasuyuki Matsushita, Yasushi Yagi |
ICCP | 3 |
| 2013 | Multiview Photometric Stereo Using Planar Mesh ParameterizationabstractWe propose a method for accurate 3D shape reconstruction using uncalibrated multiview photometric stereo. A coarse mesh reconstructed using multiview stereo is first parameterized using a planar mesh parameterization technique. Subsequently, multiview photometric stereo is performed in the 2D parameter domain of the mesh, where all geometric and photometric cues from multiple images can be treated uniformly. Unlike traditional methods, there is no need for merging view-dependent surface normal maps. Our key contribution is a new photometric stereo based mesh refinement technique that can efficiently reconstruct meshes with extremely fine geometric details by directly estimating a displacement texture map in the 2D parameter domain. We demonstrate that intricate surface geometry can be reconstructed using several challenging datasets containing surfaces with specular reflections, multiple albedos and complex topologies. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
ICCV | 3 |
| 2013 | Binocular photometric stereo acquisition and reconstruction for 3d talking head applicationsabstractIn order to render a high quality, versatile 3D talking head, a stable, high frame rate AV data acquisition system is con-structed. It can capture 3D position, surface orientation and albedo texture of the talking head video images along with the corresponding speech signals. The system consists of a com-puter controlled LED lighting subsystem; high speed stereo cameras; a microphone; and a computer for synchronous re-cording of multi-stream AV data. The visual image data col-lected is processed through a binocular photometric stereo 3D reconstruction pipeline. The pipeline automatically segments out the face; computes the depth map with binocular stereo; computes the normal map with photometric stereo; generates albedo texture; and finally constructs a high-detailed 3d model with depth and normal cues as constraints. By using the data collected with the built system, we can capture high quality dynamic facial performance, synchronized with the subject’s uttered speech. Index Terms: talking head, binocular photometric stereo, fa-cial performance capture Chaoyang Wang 0001, Yasuyuki Matsushita, Bojun Huang, Magnetro Chen, Frank K. Soong |
INTERSPEECH | 3 |
| 2013 | Depth from Refraction Using a Transparent Medium with Unknown Pose and Refractive IndexabstractIn this paper, we introduce a novel method for depth acquisition based on refraction of light. A scene is captured directly by a camera and by placing a transparent medium between the scene and the camera. A depth map of the scene is then recovered from the displacements of scene points in the images. Unlike other existing depth from refraction methods, our method does not require prior knowledge of the pose and refractive index of the transparent medium, but instead can recover them directly from the input images. By analyzing the displacements of corresponding scene points in the images, we derive closed form solutions for recovering the pose of the transparent medium and develop an iterative method for estimating the refractive index of the medium. Experimental results on both synthetic and real-world data are presented, which demonstrate the effectiveness of the proposed method. Zhihu Chen, Kwan-Yee Kenneth Wong, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 3 |
| 2013 | Guest Editorial: 3D Imaging, Processing and Modelling
Guy Godin, Michael Goesele, Yasuyuki Matsushita, Ryusuke Sagawa, Ruigang Yang |
Int. J. Comput. Vis. | 3 |
| 2013 | Radiometric Calibration by Rank MinimizationabstractWe present a robust radiometric calibration framework that capitalizes on the transform invariant low-rank structure in the various types of observations, such as sensor irradiances recorded from a static scene with different exposure times, or linear structure of irradiance color mixtures around edges. We show that various radiometric calibration problems can be treated in a principled framework that uses a rank minimization approach. This framework provides a principled way of solving radiometric calibration problems in various settings. The proposed approach is evaluated using both simulation and real-world datasets and shows superior performance to previous approaches. Joon-Young Lee, Yasuyuki Matsushita, Boxin Shi, In-So Kweon, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Appearance-Based Gaze Estimation Using Visual SaliencyabstractWe propose a gaze sensing method using visual saliency maps that does not need explicit personal calibration. Our goal is to create a gaze estimator using only the eye images captured from a person watching a video clip. Our method treats the saliency maps of the video frames as the probability distributions of the gaze points. We aggregate the saliency maps based on the similarity in eye images to efficiently identify the gaze points from the saliency maps. We establish a mapping between the eye images to the gaze points by using Gaussian process regression. In addition, we use a feedback loop from the gaze estimator to refine the gaze probability maps to improve the accuracy of the gaze estimation. The experimental results show that the proposed method works well with different people and video clips and achieves a 3.5-degree accuracy, which is sufficient for estimating a user's attention on a display. Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Nonlinear Camera Response Functions and Image Deblurring: Theoretical Analysis and PracticeabstractThis paper investigates the role that nonlinear camera response functions (CRFs) have on image deblurring. We present a comprehensive study to analyze the effects of CRFs on motion deblurring. In particular, we show how nonlinear CRFs can cause a spatially invariant blur to behave as a spatially varying blur. We prove that such nonlinearity can cause large errors around edges when directly applying deconvolution to a motion blurred image without CRF correction. These errors are inevitable even with a known point spread function (PSF) and with state-of-the-art regularization-based deconvolution algorithms. In addition, we show how CRFs can adversely affect PSF estimation algorithms in the case of blind deconvolution. To help counter these effects, we introduce two methods to estimate the CRF directly from one or more blurred images when the PSF is known or unknown. Our experimental results on synthetic and real images validate our analysis and demonstrate the robustness and accuracy of our approaches. Yu-Wing Tai, Sunyeong Kim, Seon Joo Kim, Feng Li 0005, Jie Yang 0002, Jingyi Yu 0001, Yasuyuki Matsushita, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2013 | Graph-based joint clustering of fixations and visual entitiesabstractWe present a method that extracts groups of fixations and image regions for the purpose of gaze analysis and image understanding. Since the attentional relationship between visual entities conveys rich information, automatically determining the relationship provides us a semantic representation of images. We show that, by jointly clustering human gaze and visual entities, it is possible to build meaningful and comprehensive metadata that offer an interpretation about how people see images. To achieve this, we developed a clustering method that uses a joint graph structure between fixation points and over-segmented image regions to ensure a cross-domain smoothness constraint. We show that the proposed clustering method achieves better performance in relating attention to visual entities in comparison with standard clustering techniques. Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001 |
ACM Trans. Appl. Percept. | 2 |
| 2012 | Camera spectral sensitivity estimation from a single image under unknown illumination by using fluorescenceabstractCamera spectral sensitivity plays an important role for various color-based computer vision tasks. Although several methods have been proposed to estimate it, their applicability is severely restricted by the requirement for a known illumination spectrum. In this work, we present a single-image estimation method using fluorescence with no requirement for a known illumination spectrum. Under different illuminations, the spectral distributions of fluorescence emitted from the same material remain unchanged up to a certain scale. Thus, a camera's response to the fluorescence would have the same chromaticity. Making use of this chromaticity invariance, the camera spectral sensitivity can be estimated under an arbitrary illumination whose spectrum is unknown. Through extensive experiments, we proved that our method is accurate under different illuminations. Moreover, we show how to recover the spectra of daylight from the estimated results. Finally, we use the estimated camera spectral sensitivities and daylight spectra to solve color correction problems. Shuai Han 0006, Yasuyuki Matsushita, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
CVPR | 2 |
| 2012 | Robust photometric stereo using sparse regressionabstractThis paper presents a robust photometric stereo method that effectively compensates for various non-Lambertian corruptions such as specularities, shadows, and image noise. We construct a constrained sparse regression problem that enforces both Lambertian, rank-3 structure and sparse, additive corruptions. A solution method is derived using a hierarchical Bayesian approximation to accurately estimate the surface normals while simultaneously separating the non-Lambertian corruptions. Extensive evaluations are performed that show state-of-the-art performance using both synthetic and real-world images. Satoshi Ikehata, David P. Wipf, Yasuyuki Matsushita, Kiyoharu Aizawa |
CVPR | 3 |
| 2012 | Nonlinear camera response functions and image deblurringabstractThis paper investigates the role that nonlinear camera response functions (CRFs) have on image deblurring. In particular, we show how nonlinear CRFs can cause a spatially invariant blur to behave as a spatially varying blur. This can result in noticeable ringing artifacts when deconvolution is applied even with a known point spread function (PSF). In addition, we show how CRFs can adversely affect PSF estimation algorithms in the case of blind deconvolution. To help counter these effects, we introduce two methods to estimate the CRF directly from one or more blurred images when the PSF is known or unknown. While not as accurate as conventional CRF estimation algorithms based on multiple exposures or calibration patterns, our approach is still quite effective in improving deblurring results in situations where the CRF is unknown. Sunyeong Kim, Yu-Wing Tai, Seon Joo Kim, Michael S. Brown, Yasuyuki Matsushita |
CVPR | 5 |
| 2012 | Aligning images in the wildabstractAligning image pairs with significant appearance change is a long standing computer vision challenge. Much of this problem stems from the local patch descriptors' instability to appearance variation. In this paper we suggest this instability is due less to descriptor corruption and more the difficulty in utilizing local information to canonically define the orientation (scale and rotation) at which a patch's descriptor should be computed. We address this issue by jointly estimating correspondence and relative patch orientation, within a hierarchical algorithm that utilizes a smoothly varying parameterization of geometric transformations. By collectively estimating the correspondence and orientation of all the features, we can align and orient features that cannot be stably matched with only local information. At the price of smoothing over motion discontinuities (due to independent motion or parallax), this approach can align image pairs that display significant inter-image appearance variations. Wen-Yan Lin, Yasuyuki Matsushita, Kok-Lim Low |
CVPR | 3 |
| 2012 | A biquadratic reflectance model for radiometric image analysisabstractRadiometric image analysis methods heavily rely on reflectance models. Due to the complexity of real materials, methods based on simple models such as the Lambertian model often suffer from inaccuracy. On the other hand, more advanced models such as the Cook-Torrance model severely complicate the analysis problem. We tackle this dilemma by focusing on the low-frequency component of the reflectance. We propose a compact biquadratic reflectance model to represent the reflectance of a broad class of materials precisely in the low-frequency domain. We validate our model by fitting to both existing parametric models and non-parametric measured data, and show that our model outperforms existing parametric diffuse models. We show applications of reflectometry using general diffuse surfaces and photometric stereo for general isotropic materials. Experimental results show the effectiveness of our biquadratic model and its usefulness in radiometric image analysis. Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi |
CVPR | 3 |
| 2012 | Edge-preserving photometric stereo via depth fusionabstractWe present a sensor fusion scheme that combines active stereo with photometric stereo. Aiming at capturing full-frame depth for dynamic scenes at a minimum of three lighting conditions, we formulate an iterative optimization scheme that (1) adaptively adjusts the contribution from photometric stereo so that discontinuity can be preserved; (2) detects shadow areas by checking the visibility of the estimated point with respect to the light source, instead of using image-based heuristics; and (3) behaves well for ill-conditioned pixels that are under shadow, which are inevitable in almost any scene. Furthermore, we decompose our non-linear cost function into subproblems that can be optimized efficiently using linear techniques. Experiments show significantly improved results over the previous state-of-the-art in sensor fusion. Qing Zhang 0017, Mao Ye 0005, Ruigang Yang, Yasuyuki Matsushita, Bennett Wilburn |
CVPR | 4 |
| 2012 | Shape from Single Scattering for Translucent Objects
Chika Inoshita, Yasuhiro Mukaigawa, Yasuyuki Matsushita, Yasushi Yagi |
ECCV (2) | 3 |
| 2012 | Elevation Angle from Reflectance Monotonicity: Photometric Stereo for General Isotropic Reflectances
Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi |
ECCV (3) | 3 |
| 2012 | Motion Detail Preserving Optical Flow EstimationabstractA common problem of optical flow estimation in the multiscale variational framework is that fine motion structures cannot always be correctly estimated, especially for regions with significant and abrupt displacement variation. A novel extended coarse-to-fine (EC2F) refinement framework is introduced in this paper to address this issue, which reduces the reliance of flow estimates on their initial values propagated from the coarse level and enables recovering many motion details in each scale. The contribution of this paper also includes adaptation of the objective function to handle outliers and development of a new optimization procedure. The effectiveness of our algorithm is demonstrated by Middlebury optical flow benchmarkmarking and by experiments on challenging examples that involve large-displacement motion. Li Xu 0001, Jiaya Jia, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Noise suppression in low-light images through joint denoising and demosaicingabstractWe address the effects of noise in low-light images in this paper. Color images are usually captured by a sensor with a color filter array (CFA). This requires a demosaicing process to generate a full color image. The captured images typically have low signal-to-noise ratio, and the demosaicing step further corrupts the image, which we show to be the leading cause of visually objectionable random noise patterns (splotches). To avoid this problem, we propose a combined framework of denoising and demosaicing, where we use information about the image inferred in the denoising step to perform demosaicing. Our experiments show that such a framework results in sharper low-light images that are devoid of splotches and other noise artifacts. Priyam Chatterjee, Neel Joshi, Sing Bing Kang, Yasuyuki Matsushita |
CVPR | 4 |
| 2011 | High-resolution hyperspectral imaging via matrix factorizationabstractHyperspectral imaging is a promising tool for applications in geosensing, cultural heritage and beyond. However, compared to current RGB cameras, existing hyperspectral cameras are severely limited in spatial resolution. In this paper, we introduce a simple new technique for reconstructing a very high-resolution hyperspectral image from two readily obtained measurements: A lower-resolution hyper-spectral image and a high-resolution RGB image. Our approach is divided into two stages: We first apply an unmixing algorithm to the hyperspectral input, to estimate a basis representing reflectance spectra. We then use this representation in conjunction with the RGB input to produce the desired result. Our approach to unmixing is motivated by the spatial sparsity of the hyperspectral input, and casts the unmixing problem as the search for a factorization of the input into a basis and a set of maximally sparse coefficients. Experiments show that this simple approach performs reasonably well on both simulations and real data examples. Rei Kawakami, Yasuyuki Matsushita, John Wright 0001, Moshe Ben-Ezra, Yu-Wing Tai, Katsushi Ikeuchi |
CVPR | 2 |
| 2011 | Radiometric calibration by transform invariant low-rank structureabstractWe present a robust radiometric calibration method that capitalizes on the transform invariant low-rank structure of sensor irradiances recorded from a static scene with different exposure times. We formulate the radiometric calibration problem as a rank minimization problem. Unlike previous approaches, our method naturally avoids over-fitting problem; therefore, it is robust against biased distribution of the input data, which is common in practice. When the exposure times are completely unknown, the proposed method can robustly estimate the response function up to an exponential ambiguity. The method is evaluated using both simulation and real-world datasets and shows a superior performance than previous approaches. Joon-Young Lee, Boxin Shi, Yasuyuki Matsushita, In-So Kweon, Katsushi Ikeuchi |
CVPR | 3 |
| 2011 | Smoothly varying affine stitchingabstractTraditional image stitching using parametric transforms such as homography, only produces perceptually correct composites for planar scenes or parallax free camera motion between source frames. This limits mosaicing to source images taken from the same physical location. In this paper, we introduce a smoothly varying affine stitching field which is flexible enough to handle parallax while retaining the good extrapolation and occlusion handling properties of parametric transforms. Our algorithm which jointly estimates both the stitching field and correspondence, permits the stitching of general motion source images, provided the scenes do not contain abrupt protrusions. Wen-Yan Lin, Yasuyuki Matsushita, Tian-Tsong Ng, Loong Fah Cheong |
CVPR | 3 |
| 2011 | High-quality shape from multi-view stereo and shading under general illuminationabstractMulti-view stereo methods reconstruct 3D geometry from images well for sufficiently textured scenes, but often fail to recover high-frequency surface detail, particularly for smoothly shaded surfaces. On the other hand, shape-from-shading methods can recover fine detail from shading variations. Unfortunately, it is non-trivial to apply shape-from-shading alone to multi-view data, and most shading-based estimation methods only succeed under very restricted or controlled illumination. We present a new algorithm that combines multi-view stereo and shading-based refinement for high-quality reconstruction of 3D geometry models from images taken under constant but otherwise arbitrary illumination. We have tested our algorithm on several scenes that were captured under several general and unknown lighting conditions, and we show that our final reconstructions rival laser range scans. Chenglei Wu, Bennett Wilburn, Yasuyuki Matsushita, Christian Theobalt |
CVPR | 3 |
| 2011 | Camera calibration with lens distortion from low-rank texturesabstractWe present a simple, accurate, and flexible method to calibrate intrinsic parameters of a camera together with (possibly significant) lens distortion. This new method can work under a wide range of practical scenarios: using multiple images of a known pattern, multiple images of an unknown pattern, single or multiple image(s) of multiple patterns, etc. Moreover, this new method does not rely on extracting any low-level features such as corners or edges. It can tolerate considerably large lens distortion, noise, error, illumination and viewpoint change, and still obtain accurate estimation of the camera parameters. The new method leverages on the recent breakthroughs in powerful high-dimensional convex optimization tools, especially those for matrix rank minimization and sparse signal recovery. We will show how the camera calibration problem can be formulated as an important extension to principal component pursuit, and solved by similar techniques. We characterize to exactly what extent the parameters can be recovered in case of ambiguity. We verify the efficacy and accuracy of the proposed algorithm with extensive experiments on real images. Zhengdong Zhang 0001, Yasuyuki Matsushita, Yi Ma 0001 |
CVPR | 2 |
| 2011 | Self-calibrating depth from refractionabstractIn this paper, we introduce a novel method for depth acquisition based on refraction of light. A scene is captured twice by a fixed perspective camera, with the first image captured directly by the camera and the second by placing a transparent medium between the scene and the camera. A depth map of the scene is then recovered from the displacements of scene points in the images. Unlike other existing depth from refraction methods, our method does not require the knowledge of the pose and refractive index of the transparent medium, but can recover them directly from the input images. We hence call our method self-calibrating depth from refraction. Experimental results on both synthetic and real-world data are presented, which demonstrate the effectiveness of the proposed method. Zhihu Chen, Kwan-Yee Kenneth Wong, Yasuyuki Matsushita, Miaomiao Liu 0001 |
ICCV | 3 |
| 2010 | Hemispherical Confocal Imaging Using Turtleback Reflector
Yasuhiro Mukaigawa, Seiichi Tagawa, Ramesh Raskar, Yasuyuki Matsushita, Yasushi Yagi |
ACCV (1) | 5 |
| 2010 | Robust Photometric Stereo via Low-Rank Matrix Completion and Recovery
Lun Wu, Arvind Ganesh, Boxin Shi, Yasuyuki Matsushita, Yongtian Wang, Yi Ma 0001 |
ACCV (3) | 4 |
| 2010 | Consensus photometric stereoabstractThis paper describes a photometric stereo method that works with a wide range of surface reflectances. Unlike previous approaches that assume simple parametric models such as Lambertian reflectance, the only assumption that we make is that the reflectance has three properties; monotonicity, visibility, and isotropy with respect to the cosine of light direction and surface orientation. In fact, these properties are observed in many non-Lambertian diffuse reflectances. We also show that the monotonicity and isotropy properties hold specular lobes with respect to the cosine of the surface orientation and the bisector between the light direction and view direction. Each of these three properties independently gives a possible solution space of the surface orientation. By taking the intersection of the solution spaces, our method determines the surface orientation in a consensus manner. Our method naturally avoids the need for radiometrically calibrating cameras because the radiometric response function preserves these three properties. The effectiveness of the proposed method is demonstrated using various simulated and real-world scenes that contain a variety of diffuse and specular surfaces. Tomoaki Higo, Yasuyuki Matsushita, Katsushi Ikeuchi |
CVPR | 2 |
| 2010 | Self-calibrating photometric stereoabstractWe present a self-calibrating photometric stereo method. From a set of images taken from a fixed viewpoint under different and unknown lighting conditions, our method automatically determines a radiometric response function and resolves the generalized bas-relief ambiguity for estimating accurate surface normals and albedos. We show that color and intensity profiles, which are obtained from registered pixels across images, serve as effective cues for addressing these two calibration problems. As a result, we develop a complete auto-calibration method for photometric stereo. The proposed method is useful in many practical scenarios where calibrations are difficult. Experimental results validate the accuracy of the proposed method using various real-world scenes. Boxin Shi, Yasuyuki Matsushita, Chao Xu 0006, Ping Tan 0002 |
CVPR | 2 |
| 2010 | Calibration-free gaze sensing using saliency mapsabstractWe propose a calibration-free gaze sensing method using visual saliency maps. Our goal is to construct a gaze estimator only using eye images captured from a person watching a video clip. The key is treating saliency maps of the video frames as probability distributions of gaze points. To efficiently identify gaze points from saliency maps, we aggregate saliency maps based on the similarity of eye appearances. We establish mapping between eye images to gaze points by Gaussian process regression. The experimental result shows that the proposed method works well with different people and video clips and achieves 6 degrees of accuracy, which is useful for estimating a person's attention on monitors. Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001 |
CVPR | 2 |
| 2010 | Estimating demosaicing algorithms using image noise varianceabstractWe propose a method for estimating demosaicing algorithms from image noise variance. We show that the noise variance in interpolated pixels becomes smaller than that of directly observed pixels without interpolation. Our method capitalizes on the spatial variation of image noise variance in demosaiced images to estimate the color filter array patterns and demosaicing algorithms. We verify the effectiveness of the proposed method using various images demosaiced with different demosaicing algorithms extensively. Jun Takamatsu, Yasuyuki Matsushita, Tsukasa Ogasawara, Katsushi Ikeuchi |
CVPR | 2 |
| 2010 | Motion detail preserving optical flow estimationabstractWe discuss the cause of a severe optical flow estimation problem that fine motion structures cannot always be correctly reconstructed in the commonly employed multi-scale variational framework. Our major finding is that significant and abrupt displacement transition wrecks small-scale motion structures in the coarse-to-fine refinement. A novel optical flow estimation method is proposed in this paper to address this issue, which reduces the reliance of the flow estimates on their initial values propagated from the coarser level and enables recovering many motion details in each scale. The contribution of this paper also includes adaption of the objective function and development of a new optimization procedure. The effectiveness of our method is borne out by experiments for both large- and small-displacement optical flow estimation. Li Xu 0001, Jiaya Jia, Yasuyuki Matsushita |
CVPR | 3 |
| 2010 | Shape from Second-Bounce of Light Transport
Tian-Tsong Ng, Yasuyuki Matsushita |
ECCV (2) | 3 |
| 2009 | Interactive Shadow Removal from a Single Image Using Hierarchical Graph Cut
Daisuke Miyazaki, Yasuyuki Matsushita, Katsushi Ikeuchi |
ACCV (1) | 2 |
| 2009 | A hand-held photometric stereo camera for 3-D modelingabstractThis paper presents a simple yet practical 3-D modeling method for recovering surface shape and reflectance from a set of images. We attach a point light source to a hand-held camera to add a photometric constraint to the multi-view stereo problem. Using the photometric constraint, we simultaneously solve for shape, surface normal, and reflectance. Unlike prior approaches, we formulate the problem using realistic assumptions of a near light source, non-Lambertian surfaces, perspective camera model, and the presence of ambient lighting. The effectiveness of the proposed method is verified using simulated and real-world scenes. Tomoaki Higo, Yasuyuki Matsushita, Neel Joshi, Katsushi Ikeuchi |
ICCV | 2 |
| 2009 | Image retargeting using importance diffusionabstractThis paper presents a simple and effective image retargeting method that preserves visually important parts while reducing unwanted distortions of an image. Our approach is based on a novel importance diffusion scheme, which propagates importance of removed pixels to their neighbors for preserving visual contexts and avoiding over-shrinkage of unimportant parts. Importance diffusion enables even a simple row/column removal method, which removes the least important rows/columns repeatedly, to produce visually pleasant results. It also provides control over the trade-off between uniform and non-uniform sampling for the row/column removal and seam carving methods. Experimental result demonstrates that importance diffusion successfully improves the retargeting results of row/column removal and seam carving. Sunghyun Cho, Hanul Choi, Yasuyuki Matsushita, Seungyong Lee 0001 |
ICIP | 3 |
| 2009 | An improved belief propagation method for dynamic collage
Yingzhen Yang, Qunsheng Peng 0001, Yasuyuki Matsushita |
Vis. Comput. | 5 |
| 2008 | Estimating camera response functions using probabilistic intensity similarityabstractWe propose a method for estimating camera response functions using a probabilistic intensity similarity measure. The similarity measure represents the likelihood of two intensity observations corresponding to the same scene radiance in the presence of noise. We show that the response function and the intensity similarity measure are strongly related. Our method requires several input images of a static scene taken from the same viewing position with fixed camera parameters. Noise causes pixel values at the same pixel coordinate to vary in these images, even though they measure the same scene radiance. We use these fluctuations to estimate the response function by maximizing the intensity similarity function for all pixels. Unlike prior noise-based estimation methods, our method requires only a small number of images, so it works with digital cameras as well as video cameras. Moreover, our method does not rely on any special image processing or statistical prior models. Real-world experiments using different cameras demonstrate the effectiveness of the technique. Jun Takamatsu, Yasuyuki Matsushita, Katsushi Ikeuchi |
CVPR | 2 |
| 2008 | Radiometric calibration using temporal irradiance mixturesabstractWe propose a new method for sampling camera response functions: temporally mixing two uncalibrated irradiances within a single camera exposure. Calibration methods rely on some known relationship between irradiance at the camera image plane and measured pixel intensities. Prior approaches use a color checker chart with known reflectances, registered images with different exposure ratios, or even the irradiance distribution along edges in images. We show that temporally blending irradiances allows us to densely sample the camera response function with known relative irradiances. Our first method computes the camera response curve using temporal mixtures of two pixel intensities on an uncalibrated computer display. The second approach makes use of temporal irradiance mixtures caused by motion blur. Both methods require only one input image, although more images can be used for improved robustness to noise or to cover more of the response curve. We show that our methods compute accurate response functions for a variety of cameras. Bennett Wilburn, Yasuyuki Matsushita |
CVPR | 3 |
| 2008 | An Incremental Learning Method for Unconstrained Gaze Estimation
Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001, Hideki Koike |
ECCV (3) | 2 |
| 2008 | Estimating Radiometric Response Functions from Image Noise Variance
Jun Takamatsu, Yasuyuki Matsushita, Katsushi Ikeuchi |
ECCV (4) | 2 |
| 2008 | Statistical Analysis of Global Motion Chains
Jenny Yuen, Yasuyuki Matsushita |
ECCV (2) | 2 |
| 2007 | A Probabilistic Intensity Similarity Measure based on Noise DistributionsabstractWe derive a probabilistic similarity measure between two observed image intensities that is based on the noise properties of the camera. In many vision algorithms, the effect of camera noise is either neglected or reduced in a preprocessing stage. However, noise reduction cannot be performed with high accuracy due to lack of knowledge about the true intensity signal. Our similarity metric specifically represents the likelihood that two intensity observations correspond to the same unknown noise-free scene radiance. By directly accounting for noise in the evaluation of similarity, the proposed measure makes noise reduction unnecessary and enhances many vision algorithms that involve matching of image intensities. Real-world experiments demonstrate the effectiveness of the proposed similarity measure in comparison to the standard L2norm. Yasuyuki Matsushita, Stephen Lin 0001 |
CVPR | 1 |
| 2007 | Radiometric Calibration from Noise DistributionsabstractA method is proposed for estimating radiometric response functions from noise observations. From the statistical properties of noise sources, the noise distribution for each scene radiance value is shown to be symmetric for a radiometrically calibrated camera. However, due to the non-linearity of camera response functions, the observed noise distributions become skewed in an uncalibrated camera. In this paper, we capitalize on these asymmetric profiles of measured noise distributions to estimate radiometric response functions. Unlike prior approaches, the proposed method is not sensitive to noise level, and is therefore particularly useful when the noise level is high. Also, the proposed method does not require registered input images taken with different exposures; only statistical noise distributions at multiple intensity levels are used. Real-world experiments demonstrate the effectiveness of the proposed approach in comparison to standard calibration techniques. Yasuyuki Matsushita, Stephen Lin 0001 |
CVPR | 1 |
| 2007 | Removing Non-Uniform Motion Blur from ImagesabstractWe propose a method for removing non-uniform motion blur from multiple blurry images. Traditional methods focus on estimating a single motion blur kernel for the entire image. In contrast, we aim to restore images blurred by unknown, spatially varying motion blur kernels caused by different relative motions between the camera and the scene. Our algorithm simultaneously estimates multiple motions, motion blur kernels, and the associated image segments. We formulate the problem as a regularized energy function and solve it using an alternating optimization technique. Real- world experiments demonstrate the effectiveness of the proposed method. Sunghyun Cho, Yasuyuki Matsushita, Seungyong Lee 0001 |
ICCV | 2 |
| 2007 | Illumination Brush: Interactive Design of All-Frequency LightingabstractWe present an appearance-based user interface for artists to efficiently design customized image-based lighting environments. 1 Our approach avoids typical iterations of parameter editing, rendering, and confirmation by providing a set of intuitive user interfaces for directly specifying the desired appearance of the model in the scene. Then the system automatically creates the lighting environment by solving the inverse shading problem. To obtain a realistic image, all-frequency lighting is used with a spherical radial basis function (SRBF) representation. Rendering is performed using precomputed radiance transfer (PRT) to achieve a responsive speed. User experiments demonstrated the effectiveness of the proposed system compared to a previous approach. Makoto Okabe, Yasuyuki Matsushita, Li Shen 0003, Takeo Igarashi |
PG | 2 |
| 2006 | Space-Time Video MontageabstractConventional video summarization methods focus predominantly on summarizing videos along the time axis, such as building a movie trailer: The resulting video trailer tends to retain much empty space in the background of the video frames while discarding much informative video content due to size limit. In this paper we propose a novel spacetime video summarization method which we call space-time video montage. The method simultaneously analyzes both the spatial and temporal injbrmation distribution in a video sequence, and extracts the visually informative space-time portions of the input videos. The informative video porlions are represented in volumetric layers. The layers are then packrd together in a smull ouzput video volume such that the total amount of visual information in the video volume is maximized. To achieve the packing process, we develop a new algorithm based upon the first-fit and Graph cut optimization techniques. Since our method is uble to cut off spatially und temporally less informative portions, it is uble to generate much more compact yet highly informative output videos. The effecliveness of our method is validated by extensive experiments over a wide variety of videos. Hong-Wen Kang, Yasuyuki Matsushita, Xiaoou Tang, Xue-Quan Chen |
CVPR (2) | 2 |
| 2006 | Video Completion by Motion Field TransferabstractExisting methods for video completion typically rely on periodic color transitions, layer extraction, or temporally local motion. However, periodicity may be imperceptible or absent, layer extraction is difficult, and temporally local motion cannot handle large holes. This paper presents a new approach for video completion using motion field transfer to avoid such problems. Unlike prior methods, we fill in missing video parts by sampling spatio-temporal patches of local motion instead of directly sampling color. Once the local motion field has been computed within the missing parts of the video, color can then be propagated to produce a seamless hole-free video. We have validated our method on many videos spanning a variety of scenes. We can also use the same approach to perform frame interpolation using motion fields from different videos. Takaaki Shiratori, Yasuyuki Matsushita, Xiaoou Tang, Sing Bing Kang |
CVPR (1) | 2 |
| 2006 | An Intensity Similarity Measure in Low-Light Conditions
François Alter, Yasuyuki Matsushita, Xiaoou Tang |
ECCV (4) | 2 |
| 2006 | Full-Frame Video Stabilization with Motion InpaintingabstractVideo stabilization is an important video enhancement technology which aims at removing annoying shaky motion from videos. We propose a practical and robust approach of video stabilization that produces full-frame stabilized videos with good visual quality. While most previous methods end up with producing smaller size stabilized videos, our completion method can produce full-frame videos by naturally filling in missing image parts by locally aligning image data of neighboring frames. To achieve this, motion inpainting is proposed to enforce spatial and temporal consistency of the completion in both static and dynamic image areas. In addition, image quality in the stabilized video is enhanced with a new practical deblurring algorithm. Instead of estimating point spread functions, our method transfers and interpolates sharper image pixels of neighboring frames to increase the sharpness of the frame. The proposed video completion and deblurring methods enabled us to develop a complete video stabilizer which can naturally keep the original image quality in the stabilized videos. The effectiveness of our method is confirmed by extensive experiments over a wide variety of videos. Yasuyuki Matsushita, Eyal Ofek, Weina Ge, Xiaoou Tang, Harry Shum |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | Dynamic stills and clip trailers
Yaron Caspi, Anat Axelrod, Yasuyuki Matsushita, Alon Gamliel |
Vis. Comput. | 3 |
| 2005 | Full-Frame Video StabilizationabstractVideo stabilization is an important video enhancement technology which aims at removing annoying shaky motion from videos. We propose a practical and robust approach of video stabilization that produces full-frame stabilized videos with good visual quality. While most previous methods end up with producing low resolution stabilized videos, our completion method can produce full-frame videos by naturally filling in missing image parts by locally aligning image data of neighboring frames. To achieve this, motion inpainting is proposed to enforce spatial and temporal consistency of the completion in both static and dynamic image areas. In addition, image quality in the stabilized video is enhanced with a new practical deblurring algorithm. Instead of estimating point spread functions, our method transfers and interpolates sharper image pixels of neighbouring frames to increase the sharpness of the frame. The proposed video completion and deblurring methods enabled us to develop a complete video stabilizer which can naturally keep the original image quality in the stabilized videos. The effectiveness of our method is confirmed by extensive experiments over a wide variety of videos. Yasuyuki Matsushita, Eyal Ofek, Xiaoou Tang, Harry Shum |
CVPR (1) | 1 |
| 2005 | Interactive Shape from ShadingabstractShape from shading (SfS) has always been difficult for real applications due to its intrinsic ill-posedness. In this paper, we propose an interactive SfS method which efficiently uses human knowledge in order to resolve ambiguity. We propose a global solution of continuous surfaces with a few constraints of surface normals that are interactively imposed to regularize the problem. A surface is divided into local patches, and each local solution is estimated with a fast marching SfS. It is shown that the boundaries of local solutions constitute a weighted Voronoi diagram, which allows for the formation of a global solution from the local ones. Finally, we optimize this global estimation by minimizing an energy functional based on shading and smoothness priors. Reconstruction results from both synthetic and real images demonstrate the usability of the new approach for various modeling applications. Yasuyuki Matsushita, Long Quan, Harry Shum |
CVPR (1) | 2 |
| 2005 | A Theory of Inverse Light TransportabstractIn this paper we consider the problem of computing and removing interreflections in photographs of real scenes. Towards this end, we introduce the problem of inverse light transport - given a photograph of an unknown scene, decompose it into a sum of n-bounce images, where each image records the contribution of light that bounces exactly n times before reaching the camera. We prove the existence of a set of interreflection cancelation operators that enable computing each n-bounce image by multiplying the photograph by a matrix. This matrix is derived from a set of "impulse images" obtained by probing the scene with a narrow beam of light. The operators work under unknown and arbitrary illumination, and exist for scenes that have arbitrary spatially-varying BRDFs. We derive a closed-form expression for these operators in the Lambertian case and present experiments with textured and untextured Lambertian scenes that confirm our theory's predictions. Steven M. Seitz, Yasuyuki Matsushita, Kiriakos N. Kutulakos |
ICCV | 2 |
| 2005 | Decorating Surfaces with Bidirectional Texture FunctionsabstractWe present a system for decorating arbitrary surfaces with bidirectional texture functions (BTF). Our system generates BTFs in two steps. First, we automatically synthesize a BTF over the target surface from a given BTF sample. Then, we let the user interactively paint BTF patches onto the surface such that the painted patches seamlessly integrate with the background patterns. Our system is based on a patch-based texture synthesis approach known as quilting. We present a graphcut algorithm for BTF synthesis on surfaces and the algorithm works well for a wide variety of BTF samples, including those which present problems for existing algorithms. We also describe a graphcut texture painting algorithm for creating new surface imperfections (e.g., dirt, cracks, scratches) from existing imperfections found in input BTF samples. Using these algorithms, we can decorate surfaces with real-world textures that have spatially-variant reflectance, fine-scale geometry details, and surfaces imperfections. A particularly attractive feature of BTF painting is that it allows us to capture imperfections of real materials and paint them onto geometry models. We demonstrate the effectiveness of our system with examples. Kun Zhou 0001, Lifeng Wang 0001, Yasuyuki Matsushita, Jiaoying Shi, Baining Guo, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2004 | Estimating Intrinsic Images from Image Sequences with Biased Illumination
Yasuyuki Matsushita, Stephen Lin 0001, Sing Bing Kang, Harry Shum |
ECCV (2) | 1 |
| 2004 | Illumination Normalization with Time-Dependent Intrinsic Images for Video SurveillanceabstractVariation in illumination conditions caused by weather, time of day, etc., makes the task difficult when building video surveillance systems of real world scenes. Especially, cast shadows produce troublesome effects, typically for object tracking from a fixed viewpoint, since it yields appearance variations of objects depending on whether they are inside or outside the shadow. In this paper, we handle such appearance variations by removing shadows in the image sequence. This can be considered as a preprocessing stage which leads to robust video surveillance. To achieve this, we propose a framework based on the idea of intrinsic images. Unlike previous methods of deriving intrinsic images, we derive time-varying reflectance images and corresponding illumination images from a sequence of images instead of assuming a single reflectance image. Using obtained illumination images, we normalize the input image sequence in terms of incident lighting distribution to eliminate shadowing effects. We also propose an illumination normalization scheme which can potentially run in real time, utilizing the illumination eigenspace, which captures the illumination variation due to weather, time of day, etc., and a shadow interpolation method based on shadow hulls. This paper describes the theory of the framework with simulation results and shows its effectiveness with object tracking results on real scene data sets. Yasuyuki Matsushita, Ko Nishino, Katsushi Ikeuchi, Masao Sakauchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Illumination Normalization with Time-dependent Intrinsic Images for Video SurveillanceabstractCast shadows produce troublesome effects for video surveillance systems, typically for object tracking from a fixed viewpoint, since it yields appearance variations of objects depending on whether they are inside or outside the shadow. To robustly eliminate these shadows from image sequences as a preprocessing stage for robust video surveillance, we propose a framework based on the idea of intrinsic images. Unlike previous methods for deriving intrinsic images, we derive time-varying reflectance images and corresponding illumination images from a sequence of images. Using obtained illumination images, we normalize the input image sequence in terms of incident lighting distribution to eliminate shadow effects. We also propose an illumination normalization scheme, which can potentially run in real time, utilizing the illumination eigenspace, which captures the illumination variation due to weather, time of day etc., and a shadow interpolation method based on shadow hulls. This paper describes the theory of the framework with simulation results, and shows its effectiveness with object tracking results on real scene data sets for traffic monitoring. Yasuyuki Matsushita, Ko Nishino, Katsushi Ikeuchi, Masao Sakauchi |
CVPR (1) | 1 |
| 2002 | Lighting Interpolation by Shadow Morphing Using Intrinsic LumigraphsabstractDensely-sampled image representations such as the light field or lumigraph have been effective in enabling photorealistic image synthesis. Unfortunately, lighting interpolation with such representations has not been shown to be possible without the use of accurate 3D geometry and surface reflectance properties. In this paper we propose an approach to image-based lighting interpolation that is based on estimates of geometry and shading from relatively few images. We decompose captured light fields at different lighting conditions into intrinsic images (reflectance and illumination images), and estimate view-dependent scene geometries using multi-view stereo. We call the resulting representation an intrinsic lumigraph. In the same way that the lumigraph uses geometry to permit more accurate view interpolation, the intrinsic lumigraph uses both geometry and intrinsic images to allow high-quality interpolation at different views and lighting conditions. Joint use of geometry and intrinsic images is effective in the computation of shadow masks for shadow prediction at new lighting conditions. We illustrate our approach with images of real scenes. Yasuyuki Matsushita, Sing Bing Kang, Stephen Lin 0001, Harry Shum, Xin Tong 0001 |
PG | 1 |
| 2000 | Occlusion Robust Tracking Utilizing Spatio-Temporal Markov Random Field ModelabstractIt is very important to achieve reliable vehicle tracking in ITS application such as accident detection. The most difficult problem associated with vehicle tracking is the occlusion effect among vehicles. In order to resolve this problem, we applied the dedicated algorithm which we defined as spatio-temporal Markov random field model to traffic images at an intersection. The spatio-temporal MRF considers texture correlations between consecutive images as well as the correlation among neighbors within a image. As a result, we were able to track vehicles at the intersection robustly against occlusions. Vehicles appear in various kinds of shapes and they move in random manners at the intersection. Although occlusions occur in such complicated manners, the algorithm given was able to segment and track such occluded vehicles at a high success rate of 93-96%. The algorithm requires only gray scale images and does not assume any physical models of vehicles. Shunsuke Kamijo, Yasuyuki Matsushita, Katsushi Ikeuchi, Masao Sakauchi |
ICPR | 2 |
| 2000 | Traffic monitoring and accident detection at intersectionsabstractWe have developed an algorithm, referred to as spatio-temporal Markov random field, for traffic images at intersections. This algorithm models a tracking problem by determining the state of each pixel in an image and its transit, and how such states transit along both the x-y image axes as well as the time axes. Our algorithm is sufficiently robust to segment and track occluded vehicles at a high success rate of 93%-96%. This success has led to the development of an extendable robust event recognition system based on the hidden Markov model (HMM). The system learns various event behavior patterns of each vehicle in the HMM chains and then, using the output from the tracking system, identifies current event chains. The current system can recognize bumping, passing, and jamming. However, by including other event patterns in the training set, the system can be extended to recognize those other events, e.g., illegal U-turns or reckless driving. We have implemented this system, evaluated it using the tracking results, and demonstrated its effectiveness. Shunsuke Kamijo, Yasuyuki Matsushita, Katsushi Ikeuchi, Masao Sakauchi |
IEEE Trans. Intell. Transp. Syst. | 2 |