VLDB 2026 Research / reviewers in the wild / expert
Hiroaki Santo
dblp:199/6850
· DBLP profile ↗
23ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0003-2891-5993ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent SpaceabstractThis paper introduces a method and application for automatically detecting behavioral interactions between grazing cattle from a single image, which is essential for smart livestock management in the cattle industry, such as for detecting estrus. Although interaction detection for humans has been actively studied, a non-trivial challenge lies in cattle interaction detection, specifically the lack of a comprehensive behavioral dataset that includes interactions, as the interactions of grazing cattle are rare events. We, therefore, propose CattleAct, a data-efficient method for interaction detection by decomposing interactions into the combinations of actions by individual cattle. Specifically, we first learn an action latent space from a large-scale cattle action dataset. Then, we embed rare interactions via the fine-tuning of the pre-trained latent space using contrastive learning, thereby constructing a unified latent space of actions and interactions. On top of the proposed method, we develop a practical working system integrating video and GPS inputs. Experiments on a commercial-scale pasture demonstrate the accurate interaction detection achieved by our method compared to the baselines. Our implementation is available at https://github.com/rakawanegan/CattleAct. Ren Nakagawa, Yang Yang 0124, Risa Shinoda, Hiroaki Santo, Kenji Oyama, Fumio Okura, Takenao Ohkawa |
WACV | 4 |
| 2026 | Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image AttentionabstractFoundation segmentation models achieve reasonable leaf instance extraction from top-view crop images without training (i.e., zero-shot). However, segmenting entire plant individuals with each consisting of multiple overlapping leaves remains challenging. This problem is referred to as a hierarchical segmentation task, typically requiring annotated training datasets that are often species-specific and require significant human labor. To address this, we introduce ZeroPlantSeg, a zero-shot segmentation for rosette-shaped plant individuals from top-view images. We integrate a foundation segmentation model, extracting leaf instances, and a vision-language model, reasoning about plants’ structures to extract plant individuals without additional training. Evaluations on datasets with multiple plant species, growth stages, and shooting environments demonstrate that our method surpasses existing zero-shot methods and achieves better cross-domain performance than supervised methods. Implementations are available at https://github.com/JunhaoXing/ZeroPlantSeg. Junhao Xing, Ryohei Miyakawa, Yang Yang 0124, Xinpeng Liu 0007, Risa Shinoda, Hiroaki Santo, Yosuke Toda, Fumio Okura |
WACV | 6 |
| 2026 | PlantPose: Universal Plant Skeleton Estimation via Tree-constrained Graph GenerationabstractAbstract Accurate estimation of plant skeletal structures ( e.g. , branching structures) from images is essential for smart agriculture and plant science. Unlike human skeletons with fixed topology, plant skeleton estimation presents a unique challenge, i.e. , estimating arbitrary tree graphs from images. To address this problem, we introduce PlantPose , a universal plant skeleton estimator via tree-constrained graph generation. PlantPose combines learning-based graph generation with traditional graph algorithms to enforce tree constraints during the training loop. To enhance the model’s generalization capability, we curate a large and diverse dataset comprising real-world and synthetic plant images, along with simplified representations ( e.g. , sketches and abstract drawings). This dataset enables the generalized model to adapt to diverse input styles and categories of plant images while preserving topological consistency. Our approach demonstrates robust and accurate plant skeleton estimation across multiple domains, including previously unseen out-of-domain scenarios. Further analyses highlight the method’s strengths and limitations in handling complex, heterogeneous data distributions. All implementations and datasets are available at https://github.com/huntorochi/PlantPose/ . Xinpeng Liu 0007, Hiroaki Santo, Yosuke Toda, Fumio Okura |
Int. J. Comput. Vis. | 2 |
| 2026 | DP-SfM: Dual-Pixel Structure-From-Motion Without Scale AmbiguityabstractMulti-view 3D reconstruction, namely, structure-from-motion followed by multi-view stereo, is a fundamental component of 3D computer vision. In general, multi-view 3D reconstruction suffers from an unknown scale ambiguity unless a reference object of known size is present in the scene. In this article, we show that multi-view images captured using a dual-pixel (DP) sensor can automatically resolve the scale ambiguity, without requiring a reference object or prior calibration. Specifically, the defocus blur observed in DP images provides sufficient information to determine the absolute scale when paired with depth maps (up to scale) recovered from multi-view 3D reconstruction. Based on this observation, we develop a simple yet effective linear method to estimate the absolute scale, followed by the intensity-based optimization stage that aligns the left and right DP images by shifting them back toward each other using cross-view blur kernels. Experiments demonstrate the effectiveness of the proposed approach across diverse scenes captured with different cameras and lenses. Lilika Makabe, Kohei Ashida, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Instance-wise distribution control of text-to-image diffusion modelsabstractText-to-image diffusion models are increasingly used to generate synthetic datasets for downstream vision tasks. However, they often inherit biases from large-scale training data, which can result in unbalanced attribute distributions in the generated images. While prior efforts have attempted to mitigate these biases, most focus on single-object images and struggle to control attributes across object instances in multi-instance generations. To address this limitation, we propose an instance-wise control of the attribute distribution by fine-tuning diffusion models with guidance from a pre-trained object detector and an attribute classifier. Our approach aligns the attribute distribution over object instances in generated images with a user-defined distribution, which enables precise control over attribute proportions at the instance level. Experiments across various objects and attributes demonstrate that our method generates high-quality, multi-instance images that match the specified distribution, supporting the scalable creation of distribution-aware synthetic datasets for in-the-wild vision tasks. Weng Ian Chan, Hiroaki Santo, Yasuyuki Matsushita, Fumio Okura |
Pattern Recognit. | 2 |
| 2025 | Spectral Sensitivity Estimation with an Uncalibrated Diffraction GratingabstractThis paper introduces a practical and accurate calibration method for camera spectral sensitivity using a diffraction grating. Accurate calibration of camera spectral sensitivity is crucial for various computer vision tasks, including color correction, illumination estimation, and material analysis. Unlike existing approaches that require specialized narrow-band filters or reference targets with known spectral reflectances, our method only requires an uncalibrated diffraction grating sheet, readily available off-the-shelf. By capturing images of the direct illumination and its diffracted pattern through the grating sheet, our method estimates both the camera spectral sensitivity and the diffraction grating parameters in a closed-form manner. Experiments on synthetic and real-world data demonstrate that our method outperforms conventional reference target-based methods, underscoring its effectiveness and practicality. Lilika Makabe, Hiroaki Santo, Fumio Okura, Michael S. Brown, Yasuyuki Matsushita |
ICCV | 2 |
| 2025 | NeuraLeaf: Neural Parametric Leaf Models with Shape and Deformation Disentanglement
Yang Yang 0124, Zhendong Mao 0001, Hiroaki Santo, Yasuyuki Matsushita, Fumio Okura |
ICCV | 3 |
| 2025 | TreeFormer: Single-View Plant Skeleton Estimation via Tree-Constrained Graph Generation
Xinpeng Liu 0007, Hiroaki Santo, Yosuke Toda, Fumio Okura |
WACV | 2 |
| 2024 | NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized ImagesabstractWe present NeRSP, a Neural 3D reconstruction technique for Reflective surfaces with Sparse Polarized images. Reflective surface reconstruction is extremely challenging as specular reflections are view-dependent and thus violate the multiview consistency for multiview stereo. On the other hand, sparse image inputs, as a practical capture setting, commonly cause incomplete or distorted results due to the lack of correspondence matching. This paper jointly han-dles the challenges from sparse inputs and reflective surfaces by leveraging polarized images. We derive photomet-ric and geometric cues from the polarimetric image formation model and multiview azimuth consistency, which jointly optimize the surface geometry modeled via implicit neural representation. Based on the experiments on our synthetic and real datasets, we achieve the state-of-the-art surface reconstruction results with only 6 views as input. Yufei Han 0002, Heng Guo 0003, Koki Fukai, Hiroaki Santo, Boxin Shi, Fumio Okura, Zhanyu Ma |
CVPR | 4 |
| 2024 | MVCPS-NeuS: Multi-View Constrained Photometric Stereo for Neural Surface ReconstructionabstractMulti-view photometric stereo (MVPS) recovers a high-fidelity 3D shape of a scene by benefiting from both multi-view stereo and photometric stereo. While photometric stereo boosts detailed shape reconstruction, it necessitates recording images under various light conditions for each viewpoint. In particular, calibrating the light directions for each view significantly increases the cost of acquiring images. To make MVPS more accessible, we introduce a practical and easy-to-implement setup, multi-view constrained photometric stereo (MVCPS), where the light directions are unknown but constrained to move together with the camera. Unlike con-ventional multi-view uncalibrated photometric stereo, our constrained setting reduces the ambiguities of surface normal estimates from per-view linear ambiguities to a single and global linear one, thereby simplifying the disambiguation process. The proposed method integrates the ambiguous surface normal into neural surface reconstruction (NeuS) to simultaneously resolve the global ambiguity and estimate the detailed 3D shape. Experiments demonstrate that our method estimates accurate shapes under sparse viewpoints using only a few multi-view constrained light sources. Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
CVPR | 1 |
| 2024 | Resolving Scale Ambiguity in Multi-view 3D Reconstruction Using Dual-Pixel Sensors
Kohei Ashida, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ECCV (50) | 2 |
| 2024 | Synthesizing Time-Varying BRDFs via Latent Space
Takuto Narumoto, Hiroaki Santo, Fumio Okura |
ECCV (71) | 2 |
| 2023 | Multi-View Azimuth Stereo via Tangent Space ConsistencyabstractWe present a method for 3D reconstruction only using calibrated multi-view surface azimuth maps. Our method, multi-view azimuth stereo, is effective for textureless or specular surfaces, which are difficult for conventional multi-view stereo methods. We introduce the concept of tangent space consistency: Multi-view azimuth observations of a surface point should be lifted to the same tangent space. Leveraging this consistency, we recover the shape by optimizing a neural implicit surface representation. Our method harnesses the robust azimuth estimation capabilities of photometric stereo methods or polarization imaging while bypassing potentially complex zenith angle estimation. Experiments using azimuth maps from various sources validate the accurate shape recovery with our method, even without zenith angles. Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
CVPR | 2 |
| 2023 | Learning to Synthesize Photorealistic Dual-pixel Images from RGBD framesabstractAs a special sensor that implicitly provides ordinal depth information, dual-pixel (DP) appears to be beneficial for various tasks such as defocus deblurring and monocular depth estimation. Recent advances in data-driven dual-pixel (DP) research are bottlenecked by the difficulties in reaching large-scale DP datasets, and a photorealistic image synthesis approach appears to be a credible solution. To benchmark the accuracy of various existing DP image simulators and facilitate data-driven DP image synthesis, this work presents a real-world DP dataset consisting of approximately 5000 high-quality pairs of sharp images, DP defocus blur images, detailed imaging parameters, and accurate depth maps. Based on this large-scale dataset, we also propose a holistic data-driven framework to synthesize photorealistic DP images, where a neural network replaces conventional handcrafted imaging models. Experiments show that our neural DP simulator can generate more photorealistic DP images than existing state-of-the-art methods and effectively benefit data-driven DP-related tasks. Our code and dataset are released at https://github.com/SILI1994/Dual-Pixel-Simulator. Feiran Li, Heng Guo 0003, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ICCP | 3 |
| 2023 | Near-light Photometric Stereo with Symmetric LightsabstractThis paper describes a linear solution method for near-light photometric stereo by exploiting symmetric light source arrangements. Unlike conventional non-convex optimization approaches, by arranging multiple sets of symmetric nearby light source pairs, our method derives a closed-form solution for surface normal and depth without requiring initialization. In addition, our method works as long as the light sources are symmetrically distributed about an arbitrary point even when the entire spatial offset is uncalibrated. Experiments showcase the accuracy of shape recovery accuracy of our method, achieving comparable results to the state-of-the-art calibrated near-light photometric stereo method while significantly reducing requirements of careful depth initialization and light calibration. Lilika Makabe, Heng Guo 0003, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ICCP | 3 |
| 2022 | Bilateral Normal Integration
Hiroaki Santo, Boxin Shi, Fumio Okura, Yasuyuki Matsushita |
ECCV (1) | 2 |
| 2022 | Shape-coded ArUco: Fiducial Marker for Bridging 2D and 3D ModalitiesabstractWe introduce a fiducial marker for the registration of two-dimensional (2D) images and untextured three-dimensional (3D) shapes that are recorded by commodity laser scanners. Specifically, we design a 3D-version of the ArUco marker that retains exactly the same appearance as its 2D counterpart from any viewpoint above the marker but contains shape information. The shape-coded ArUco can naturally work with off-the-shelf ArUco marker detectors in the 2D image domain. For the 3D domain, we develop a method for detecting the marker in an untextured 3D point cloud. Experiments demonstrate accurate 2D-3D registration using our shape-coded ArUco markers in comparison to baseline methods. Lilika Makabe, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
WACV | 2 |
| 2022 | Symmetric-light Photometric StereoabstractThis paper presents symmetric-light photometric stereo for surface normal estimation, in which directional lights are distributed symmetrically with respect to the optic center. Unlike previous studies of ring-light settings that required the information of ring radius, we show that even without the knowledge of the exact light source locations or their distances from the optic center, the symmetric configuration provides us sufficient information for recovering unique surface normals without ambiguity. Specifically, under the symmetric lights, measurements of a pair of scene points having distinct surface normals but the same albedo yield a system of constrained quadratic equations about the surface normal, which has a unique solution. Experiments demonstrate that the proposed method alleviates the need for geometric light source calibration while maintaining the accuracy of calibrated photometric stereo. Kazuma Minami, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
WACV | 2 |
| 2022 | Deep Photometric Stereo Networks for Determining Surface Normal and ReflectancesabstractThis article presents a photometric stereo method based on deep learning. One of the major difficulties in photometric stereo is designing an appropriate reflectance model that is both capable of representing real-world reflectances and computationally tractable for deriving surface normal. Unlike previous photometric stereo methods that rely on a simplified parametric image formation model, such as the Lambert's model, the proposed method aims at establishing a flexible mapping between complex reflectance observations and surface normal using a deep neural network. In addition, the proposed method predicts the reflectance, which allows us to understand surface materials and to render the scene under arbitrary lighting conditions. As a result, we propose a deep photometric stereo network (DPSN) that takes reflectance observations under varying light directions and infers the surface normal and reflectance in a per-pixel manner. To make the DPSN applicable to real-world scenes, a dataset of measured BRDFs (MERL BRDF dataset) has been used for training the network. Evaluation using simulation and real-world scenes shows the effectiveness of the proposed approach in estimating both surface normal and reflectances. Hiroaki Santo, Masaki Samejima, Yusuke Sugano, Boxin Shi, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Deep Near-Light Photometric Stereo for Spatially Varying Reflectances
Hiroaki Santo, Michael Waechter, Yasuyuki Matsushita |
ECCV (8) | 1 |
| 2020 | Light Structure from Pin Motion: Geometric Point Light Source CalibrationabstractAbstract We present a method for geometric point light source calibration. Unlike prior works that use Lambertian spheres, mirror spheres, or mirror planes, we use a calibration target consisting of a plane and small shadow casters at unknown positions above the plane. We show that shadow observations from a moving calibration target under a fixed light follow the principles of pinhole camera geometry and epipolar geometry, allowing joint recovery of the light position and 3D shadow caster positions, equivalent to how conventional structure from motion jointly recovers camera parameters and 3D feature positions from observed 2D features. Moreover, we devised a unified light model that works with nearby point lights as well as distant light in one common framework. Our evaluation shows that our method yields light estimates that are stable and more accurate than existing techniques while having a much simpler setup and requiring less manual labor. Hiroaki Santo, Michael Waechter, Wen-Yan Lin, Yusuke Sugano, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 1 |
| 2018 | Light Structure from Pin Motion: Simple and Accurate Point Light Calibration for Physics-Based Modeling
Hiroaki Santo, Michael Waechter, Masaki Samejima, Yusuke Sugano, Yasuyuki Matsushita |
ECCV (3) | 1 |
| 2017 | Device-free and privacy preserving indoor positioning using infrared retro-reflection imagingabstractIndoor positioning is a core technology for indoor daily life applications such as elderly care and home automation. This study presents a device-free and privacy preserving indoor positioning method using infrared (IR) cameras and retroreflectors. The proposed method uses IR cameras equipped with IR LEDs to capture retroreflections from markers attached to walls in the environment, and detects a person who passes between the camera and a marker by observing the occlusion. Because our method employs occlusion of markers, it can track a person without attaching tags to the person. Also, our camera device permits us to filter out visible light and thus the appearance of the person is not recorded. Our evaluation in real environments showed that our method achieved an average positioning error of about 0.3 meters. Hiroaki Santo, Takuya Maekawa, Yasuyuki Matsushita |
PerCom | 1 |