EDBT 2026 Demo / reviewers in the wild / expert
Fumio Okura
dblp:18/9717
· DBLP profile ↗
41ranked-venue papers
8as first author
30since 2021 · last 2026
0000-0001-7595-1300ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised 3D Human Pose Estimation via Conditional Multi-view Ancestral Sampling
Ryohei Goto, Takuya Fujihashi, Shunsuke Saruwatari, Fumio Okura |
FG | 4 |
| 2026 | Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent SpaceabstractThis paper introduces a method and application for automatically detecting behavioral interactions between grazing cattle from a single image, which is essential for smart livestock management in the cattle industry, such as for detecting estrus. Although interaction detection for humans has been actively studied, a non-trivial challenge lies in cattle interaction detection, specifically the lack of a comprehensive behavioral dataset that includes interactions, as the interactions of grazing cattle are rare events. We, therefore, propose CattleAct, a data-efficient method for interaction detection by decomposing interactions into the combinations of actions by individual cattle. Specifically, we first learn an action latent space from a large-scale cattle action dataset. Then, we embed rare interactions via the fine-tuning of the pre-trained latent space using contrastive learning, thereby constructing a unified latent space of actions and interactions. On top of the proposed method, we develop a practical working system integrating video and GPS inputs. Experiments on a commercial-scale pasture demonstrate the accurate interaction detection achieved by our method compared to the baselines. Our implementation is available at https://github.com/rakawanegan/CattleAct. Ren Nakagawa, Yang Yang 0124, Risa Shinoda, Hiroaki Santo, Kenji Oyama, Fumio Okura, Takenao Ohkawa |
WACV | 6 |
| 2026 | Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image AttentionabstractFoundation segmentation models achieve reasonable leaf instance extraction from top-view crop images without training (i.e., zero-shot). However, segmenting entire plant individuals with each consisting of multiple overlapping leaves remains challenging. This problem is referred to as a hierarchical segmentation task, typically requiring annotated training datasets that are often species-specific and require significant human labor. To address this, we introduce ZeroPlantSeg, a zero-shot segmentation for rosette-shaped plant individuals from top-view images. We integrate a foundation segmentation model, extracting leaf instances, and a vision-language model, reasoning about plants’ structures to extract plant individuals without additional training. Evaluations on datasets with multiple plant species, growth stages, and shooting environments demonstrate that our method surpasses existing zero-shot methods and achieves better cross-domain performance than supervised methods. Implementations are available at https://github.com/JunhaoXing/ZeroPlantSeg. Junhao Xing, Ryohei Miyakawa, Yang Yang 0124, Xinpeng Liu 0007, Risa Shinoda, Hiroaki Santo, Yosuke Toda, Fumio Okura |
WACV | 8 |
| 2026 | PlantPose: Universal Plant Skeleton Estimation via Tree-constrained Graph GenerationabstractAbstract Accurate estimation of plant skeletal structures ( e.g. , branching structures) from images is essential for smart agriculture and plant science. Unlike human skeletons with fixed topology, plant skeleton estimation presents a unique challenge, i.e. , estimating arbitrary tree graphs from images. To address this problem, we introduce PlantPose , a universal plant skeleton estimator via tree-constrained graph generation. PlantPose combines learning-based graph generation with traditional graph algorithms to enforce tree constraints during the training loop. To enhance the model’s generalization capability, we curate a large and diverse dataset comprising real-world and synthetic plant images, along with simplified representations ( e.g. , sketches and abstract drawings). This dataset enables the generalized model to adapt to diverse input styles and categories of plant images while preserving topological consistency. Our approach demonstrates robust and accurate plant skeleton estimation across multiple domains, including previously unseen out-of-domain scenarios. Further analyses highlight the method’s strengths and limitations in handling complex, heterogeneous data distributions. All implementations and datasets are available at https://github.com/huntorochi/PlantPose/ . Xinpeng Liu 0007, Hiroaki Santo, Yosuke Toda, Fumio Okura |
Int. J. Comput. Vis. | 4 |
| 2026 | DP-SfM: Dual-Pixel Structure-From-Motion Without Scale AmbiguityabstractMulti-view 3D reconstruction, namely, structure-from-motion followed by multi-view stereo, is a fundamental component of 3D computer vision. In general, multi-view 3D reconstruction suffers from an unknown scale ambiguity unless a reference object of known size is present in the scene. In this article, we show that multi-view images captured using a dual-pixel (DP) sensor can automatically resolve the scale ambiguity, without requiring a reference object or prior calibration. Specifically, the defocus blur observed in DP images provides sufficient information to determine the absolute scale when paired with depth maps (up to scale) recovered from multi-view 3D reconstruction. Based on this observation, we develop a simple yet effective linear method to estimate the absolute scale, followed by the intensity-based optimization stage that aligns the left and right DP images by shifting them back toward each other using cross-view blur kernels. Experiments demonstrate the effectiveness of the proposed approach across diverse scenes captured with different cameras and lenses. Lilika Makabe, Kohei Ashida, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Instance-wise distribution control of text-to-image diffusion modelsabstractText-to-image diffusion models are increasingly used to generate synthetic datasets for downstream vision tasks. However, they often inherit biases from large-scale training data, which can result in unbalanced attribute distributions in the generated images. While prior efforts have attempted to mitigate these biases, most focus on single-object images and struggle to control attributes across object instances in multi-instance generations. To address this limitation, we propose an instance-wise control of the attribute distribution by fine-tuning diffusion models with guidance from a pre-trained object detector and an attribute classifier. Our approach aligns the attribute distribution over object instances in generated images with a user-defined distribution, which enables precise control over attribute proportions at the instance level. Experiments across various objects and attributes demonstrate that our method generates high-quality, multi-instance images that match the specified distribution, supporting the scalable creation of distribution-aware synthetic datasets for in-the-wild vision tasks. Weng Ian Chan, Hiroaki Santo, Yasuyuki Matsushita, Fumio Okura |
Pattern Recognit. | 4 |
| 2025 | HoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian SplattingabstractNovel view synthesis has demonstrated impressive progress recently, with 3D Gaussian splatting (3DGS) offering efficient training time and photorealistic real-time rendering. However, reliance on Cartesian coordinates limits 3DGS’s performance on distant objects, which is important for reconstructing unbounded outdoor environments. We found that, despite its ultimate simplicity, using homogeneous coordinates, a concept on the projective geometry, for the 3DGS pipeline remarkably improves the rendering accuracies of distant objects. We therefore propose Homogeneous Gaussian Splatting (HoGS) incorporating homogeneous coordinates into the 3DGS framework, providing a unified representation for enhancing near and distant objects. HoGS effectively manages both expansive spatial positions and scales particularly in outdoor unbounded environments by adopting projective geometry principles. Experiments show that HoGS significantly enhances accuracy in reconstructing distant objects while maintaining high-quality rendering of nearby objects, along with fast training speed and real-time rendering capability. Our implementations are available on our project page https://kh129.github.io/hogs/. Xinpeng Liu 0007, Zeyi Huang, Fumio Okura, Yasuyuki Matsushita |
CVPR | 3 |
| 2025 | Spectral Sensitivity Estimation with an Uncalibrated Diffraction GratingabstractThis paper introduces a practical and accurate calibration method for camera spectral sensitivity using a diffraction grating. Accurate calibration of camera spectral sensitivity is crucial for various computer vision tasks, including color correction, illumination estimation, and material analysis. Unlike existing approaches that require specialized narrow-band filters or reference targets with known spectral reflectances, our method only requires an uncalibrated diffraction grating sheet, readily available off-the-shelf. By capturing images of the direct illumination and its diffracted pattern through the grating sheet, our method estimates both the camera spectral sensitivity and the diffraction grating parameters in a closed-form manner. Experiments on synthetic and real-world data demonstrate that our method outperforms conventional reference target-based methods, underscoring its effectiveness and practicality. Lilika Makabe, Hiroaki Santo, Fumio Okura, Michael S. Brown, Yasuyuki Matsushita |
ICCV | 3 |
| 2025 | NeuraLeaf: Neural Parametric Leaf Models with Shape and Deformation Disentanglement
Yang Yang 0124, Zhendong Mao 0001, Hiroaki Santo, Yasuyuki Matsushita, Fumio Okura |
ICCV | 5 |
| 2025 | TreeFormer: Single-View Plant Skeleton Estimation via Tree-Constrained Graph Generation
Xinpeng Liu 0007, Hiroaki Santo, Yosuke Toda, Fumio Okura |
WACV | 4 |
| 2025 | Predicting Future Cognitive Decline From Long-Term Observations of Dual-Task Performance DataabstractEarly stage detection of cognitive decline is crucial for effective prevention and treatment of dementia. However, current approaches based on MRI or biomarkers are expensive and impractical, making them unsuitable for early-stage detection from daily measurements. A suitable option is the dual-task paradigm, which involves simultaneously performing two tasks (typically a physical task combined with a cognitive task). This approach has proven effective in assessing daily cognitive status. The underlying principle is that dual-task performance reflects the maximum cognitive load that can be handled by participants, which in turn reflects their current cognitive function. However, a one-time dual-task test cannot predict future changes in cognitive function. In this study, we present the first attempt at leveraging long-term observations of dual-task performance data. Our results show that changes in dual-task performance over time are associated with future cognitive changes. Our approach extracts temporal features from six months of dual-task performance data, and predicts future cognitive decline over the next two years using a machine learning model. Our experimental results yielded an accuracy comparable to that returned by MRI scans, thus demonstrating that the proposed approach can achieve early detection of future cognitive decline from routine dual-task measurements. Shuqiong Wu, Tomoya Noguchi, Fumio Okura, Yasushi Yagi |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized ImagesabstractWe present NeRSP, a Neural 3D reconstruction technique for Reflective surfaces with Sparse Polarized images. Reflective surface reconstruction is extremely challenging as specular reflections are view-dependent and thus violate the multiview consistency for multiview stereo. On the other hand, sparse image inputs, as a practical capture setting, commonly cause incomplete or distorted results due to the lack of correspondence matching. This paper jointly han-dles the challenges from sparse inputs and reflective surfaces by leveraging polarized images. We derive photomet-ric and geometric cues from the polarimetric image formation model and multiview azimuth consistency, which jointly optimize the surface geometry modeled via implicit neural representation. Based on the experiments on our synthetic and real datasets, we achieve the state-of-the-art surface reconstruction results with only 6 views as input. Yufei Han 0002, Heng Guo 0003, Koki Fukai, Hiroaki Santo, Boxin Shi, Fumio Okura, Zhanyu Ma |
CVPR | 6 |
| 2024 | MVCPS-NeuS: Multi-View Constrained Photometric Stereo for Neural Surface ReconstructionabstractMulti-view photometric stereo (MVPS) recovers a high-fidelity 3D shape of a scene by benefiting from both multi-view stereo and photometric stereo. While photometric stereo boosts detailed shape reconstruction, it necessitates recording images under various light conditions for each viewpoint. In particular, calibrating the light directions for each view significantly increases the cost of acquiring images. To make MVPS more accessible, we introduce a practical and easy-to-implement setup, multi-view constrained photometric stereo (MVCPS), where the light directions are unknown but constrained to move together with the camera. Unlike con-ventional multi-view uncalibrated photometric stereo, our constrained setting reduces the ambiguities of surface normal estimates from per-view linear ambiguities to a single and global linear one, thereby simplifying the disambiguation process. The proposed method integrates the ambiguous surface normal into neural surface reconstruction (NeuS) to simultaneously resolve the global ambiguity and estimate the detailed 3D shape. Experiments demonstrate that our method estimates accurate shapes under sparse viewpoints using only a few multi-view constrained light sources. Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
CVPR | 2 |
| 2024 | Resolving Scale Ambiguity in Multi-view 3D Reconstruction Using Dual-Pixel Sensors
Kohei Ashida, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ECCV (50) | 3 |
| 2024 | Synthesizing Time-Varying BRDFs via Latent Space
Takuto Narumoto, Hiroaki Santo, Fumio Okura |
ECCV (71) | 3 |
| 2023 | Multi-View Azimuth Stereo via Tangent Space ConsistencyabstractWe present a method for 3D reconstruction only using calibrated multi-view surface azimuth maps. Our method, multi-view azimuth stereo, is effective for textureless or specular surfaces, which are difficult for conventional multi-view stereo methods. We introduce the concept of tangent space consistency: Multi-view azimuth observations of a surface point should be lifted to the same tangent space. Leveraging this consistency, we recover the shape by optimizing a neural implicit surface representation. Our method harnesses the robust azimuth estimation capabilities of photometric stereo methods or polarization imaging while bypassing potentially complex zenith angle estimation. Experiments using azimuth maps from various sources validate the accurate shape recovery with our method, even without zenith angles. Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
CVPR | 3 |
| 2023 | Learning to Synthesize Photorealistic Dual-pixel Images from RGBD framesabstractAs a special sensor that implicitly provides ordinal depth information, dual-pixel (DP) appears to be beneficial for various tasks such as defocus deblurring and monocular depth estimation. Recent advances in data-driven dual-pixel (DP) research are bottlenecked by the difficulties in reaching large-scale DP datasets, and a photorealistic image synthesis approach appears to be a credible solution. To benchmark the accuracy of various existing DP image simulators and facilitate data-driven DP image synthesis, this work presents a real-world DP dataset consisting of approximately 5000 high-quality pairs of sharp images, DP defocus blur images, detailed imaging parameters, and accurate depth maps. Based on this large-scale dataset, we also propose a holistic data-driven framework to synthesize photorealistic DP images, where a neural network replaces conventional handcrafted imaging models. Experiments show that our neural DP simulator can generate more photorealistic DP images than existing state-of-the-art methods and effectively benefit data-driven DP-related tasks. Our code and dataset are released at https://github.com/SILI1994/Dual-Pixel-Simulator. Feiran Li, Heng Guo 0003, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ICCP | 4 |
| 2023 | Near-light Photometric Stereo with Symmetric LightsabstractThis paper describes a linear solution method for near-light photometric stereo by exploiting symmetric light source arrangements. Unlike conventional non-convex optimization approaches, by arranging multiple sets of symmetric nearby light source pairs, our method derives a closed-form solution for surface normal and depth without requiring initialization. In addition, our method works as long as the light sources are symmetrically distributed about an arbitrary point even when the entire spatial offset is uncalibrated. Experiments showcase the accuracy of shape recovery accuracy of our method, achieving comparable results to the state-of-the-art calibrated near-light photometric stereo method while significantly reducing requirements of careful depth initialization and light calibration. Lilika Makabe, Heng Guo 0003, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
ICCP | 4 |
| 2023 | Shuffled Linear Regression with Outliers in Both Covariates and Responses
Feiran Li, Kent Fujiwara, Fumio Okura, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 3 |
| 2023 | Discrete Search Photometric Stereo for Fast and Accurate Shape EstimationabstractWe consider the problem of estimating surface normals of a scene with spatially varying, general bidirectional reflectance distribution functions (BRDFs) observed by a static camera under varying distant illuminations. Unlike previous approaches that rely on continuous optimization of surface normals, we cast the problem as a discrete search problem over a set of finely discretized surface normals. In this setting, we show that the expensive processes can be precomputed in a scene-independent manner, resulting in accelerated inference. We discuss two variants of our discrete search photometric stereo (DSPS), one working with continuous linear combinations of BRDF bases and the other working with discrete BRDFs sampled from a BRDF space. Experiments show that DSPS has comparable accuracy to state-of-the-art exemplar-based photometric stereo methods while achieving 10-100x acceleration. Kenji Enomoto, Michael Waechter, Fumio Okura, Kiriakos N. Kutulakos, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Bilateral Normal Integration
Hiroaki Santo, Boxin Shi, Fumio Okura, Yasuyuki Matsushita |
ECCV (1) | 4 |
| 2022 | Shape-coded ArUco: Fiducial Marker for Bridging 2D and 3D ModalitiesabstractWe introduce a fiducial marker for the registration of two-dimensional (2D) images and untextured three-dimensional (3D) shapes that are recorded by commodity laser scanners. Specifically, we design a 3D-version of the ArUco marker that retains exactly the same appearance as its 2D counterpart from any viewpoint above the marker but contains shape information. The shape-coded ArUco can naturally work with off-the-shelf ArUco marker detectors in the 2D image domain. For the 3D domain, we develop a method for detecting the marker in an untextured 3D point cloud. Experiments demonstrate accurate 2D-3D registration using our shape-coded ArUco markers in comparison to baseline methods. Lilika Makabe, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
WACV | 3 |
| 2022 | Symmetric-light Photometric StereoabstractThis paper presents symmetric-light photometric stereo for surface normal estimation, in which directional lights are distributed symmetrically with respect to the optic center. Unlike previous studies of ring-light settings that required the information of ring radius, we show that even without the knowledge of the exact light source locations or their distances from the optic center, the symmetric configuration provides us sufficient information for recovering unique surface normals without ambiguity. Specifically, under the symmetric lights, measurements of a pair of scene points having distinct surface normals but the same albedo yield a system of constrained quadratic equations about the surface normal, which has a unique solution. Experiments demonstrate that the proposed method alleviates the need for geometric light source calibration while maintaining the accuracy of calibrated photometric stereo. Kazuma Minami, Hiroaki Santo, Fumio Okura, Yasuyuki Matsushita |
WACV | 3 |
| 2022 | Shape and Albedo Recovery by Your Phone using Stereoscopic Flash and No-Flash PhotographyabstractRecovering shape and albedo for the immense number of existing cultural heritage artifacts is challenging. Accurate 3D reconstruction systems are typically expensive and thus inaccessible to many and cheaper off-the-shelf 3D sensors often generate results of unsatisfactory quality. This paper presents a high-fidelity shape and albedo recovery method that only requires a stereo camera and a flashlight, a typical camera setup equipped in many off-the-shelf smartphones. The stereo camera allows us to infer rough shape from a pair of no-flash images, and a flash image is further captured for shape refinement based on our flash/no-flash image formation model. We verify the effectiveness of our method on real-world artifacts in indoor and outdoor conditions using smartphones with different camera/flashlight configurations. Comparison results demonstrate that our stereoscopic flash and no-flash photography benefits the high-fidelity shape and albedo recovery on a smartphone. Using our method, people can immediately turn their phones into high-fidelity 3D scanners, facilitating the digitization of cultural heritage artifacts. Michael Waechter, Boxin Shi, Fumio Okura, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 6 |
| 2022 | Multispectral Photometric Stereo for Spatially-Varying Spectral Reflectances
Heng Guo 0003, Fumio Okura, Boxin Shi, Takuya Funatomi, Yasuhiro Mukaigawa, Yasuyuki Matsushita |
Int. J. Comput. Vis. | 2 |
| 2021 | Normal Integration via Inverse Plane Fitting With Minimum Point-to-Plane DistanceabstractThis paper presents a surface normal integration method that solves an inverse problem of local plane fitting. Surface reconstruction from normal maps is essential in photometric shape reconstruction. To this end, we formulate normal integration in the camera coordinates and jointly solve for 3D point positions and local plane displacements. Unlike existing methods that consider the vertical distances between 3D points, we minimize the sum of squared point-to-plane distances. Our method can deal with both orthographic or perspective normal maps with arbitrary boundaries. Compared to existing normal integration methods, our method avoids the checkerboard artifact and performs more robustly against natural boundaries, sharp features, and outliers. We further provide a geometric analysis of the source of artifacts that appear in previous methods based on our plane fitting formulation. Experimental results on analytically computed, synthetic, and real-world surfaces show that our method yields accurate and stable reconstruction for both orthographic and perspective normal maps1. Boxin Shi, Fumio Okura, Yasuyuki Matsushita |
CVPR | 3 |
| 2021 | Multispectral Photometric Stereo for Spatially-Varying Spectral Reflectances: A Well Posed Problem?abstractMultispectral photometric stereo (MPS) aims at recovering the surface normal of a scene from a single-shot multi-spectral image, which is known as an ill-posed problem. To make the problem well-posed, existing MPS methods rely on restrictive assumptions, such as shape prior, surfaces having a monochromatic with uniform albedo. This paper alleviates the restrictive assumptions in existing methods. We show that the problem becomes well-posed for a surface with a uniform chromaticity but spatially-varying albedos based on our new formulation. Specifically, if at least three (or two) scene points share the same chromaticity, the proposed method uniquely recovers their surface normals and spectral reflectance with the illumination of more than or equal to four (or five) spectral lights. Besides, our method can be made robust by having many (i.e., 4 or more) spectral bands using robust estimation techniques for conventional photometric stereo. Experiments on both synthetic and real-world scenes demonstrate the effectiveness of our method. Our data and result can be found at https://github.com/GH-HOME/MultispectralPS.git. Heng Guo 0003, Fumio Okura, Boxin Shi, Takuya Funatomi, Yasuhiro Mukaigawa, Yasuyuki Matsushita |
CVPR | 2 |
| 2021 | Generalized Shuffled Linear RegressionabstractWe consider the shuffled linear regression problem where the correspondences between covariates and responses are unknown. While the existing formulation assumes an ideal underlying bijection in which all pieces of data should match, such an assumption barely holds in real-world applications due to either missing data or outliers. Therefore, in this work, we generalize the formulation of shuffled linear regression to a broader range of conditions where only part of the data should correspond. Moreover, we present a remarkably simple yet effective optimization algorithm with guaranteed global convergence. Distinct tasks validate the effectiveness of the proposed method.1 Feiran Li, Kent Fujiwara, Fumio Okura, Yasuyuki Matsushita |
ICCV | 3 |
| 2021 | A Closer Look at Rotation-invariant Deep Point Cloud AnalysisabstractWe consider the deep point cloud analysis tasks where the inputs of the networks are randomly rotated. Recent progress in rotation-invariant point cloud analysis is mainly driven by converting point clouds into their respective canonical poses, and principal component analysis (PCA) is a practical tool to achieve this. Due to the imperfect alignment of PCA, most of the current works are devoted to developing powerful network structures and features to overcome this deficiency, without thoroughly analyzing the PCA-based canonical poses themselves. In this work, we present a detailed study w.r.t. the PCA-based canonical poses of point clouds. Our investigation reveals that the ambiguity problem associated with the PCA-based canonical poses is handled insufficiently in some recent works. To this end, we develop a simple pose selector module for disambiguation, which presents noticeable enhancement (i.e., 5.3% classification accuracy) over state-of-the-art approaches on the challenging real-world dataset.1 Feiran Li, Kent Fujiwara, Fumio Okura, Yasuyuki Matsushita |
ICCV | 3 |
| 2021 | Human Localization Using a Single Camera Towards Social Distance Monitoring During Sports
Ryosuke Hasegawa, Akira Uchiyama, Fumio Okura, Daigo Muramatsu, Issei Ogasawara, Hiromi Takahata, Ken Nakata, Teruo Higashino |
MobiQuitous | 3 |
| 2020 | Descriptor-Free Multi-view Region Matching for Instance-Wise 3D Reconstruction
Takuma Doi, Fumio Okura, Toshiki Nagahara, Yasuyuki Matsushita, Yasushi Yagi |
ACCV (5) | 2 |
| 2018 | Probabilistic Plant Modeling via Multi-View Image-to-Image TranslationabstractThis paper describes a method for inferring three-dimensional (3D) plant branch structures that are hidden under leaves from multi-view observations. Unlike previous geometric approaches that heavily rely on the visibility of the branches or use parametric branching models, our method makes statistical inferences of branch structures in a probabilistic framework. By inferring the probability of branch existence using a Bayesian extension of image-to-image translation applied to each of multi-view images, our method generates a probabilistic plant 3D model, which represents the 3D branching pattern that cannot be directly observed. Experiments demonstrate the usefulness of the proposed approach in generating convincing branch structures in comparison to prior approaches. Takahiro Isokane, Fumio Okura, Ayaka Ide, Yasuyuki Matsushita, Yasushi Yagi |
CVPR | 2 |
| 2017 | Realtime Novel View Synthesis with Eigen-Texture Regression
Yuta Nakashima, Fumio Okura, Norihiko Kawai, Ryosuke Kimura, Hiroshi Kawasaki, Katsushi Ikeuchi, Ambrosio Blanco |
BMVC | 2 |
| 2015 | Unifying Color and Texture Transfer for Predictive Appearance ManipulationabstractAbstract Recent color transfer methods use local information to learn the transformation from a source to an exemplar image, and then transfer this appearance change to a target image. These solutions achieve very successful results for general mood changes, e.g., changing the appearance of an image from “sunny” to “overcast”. However, such methods have a hard time creating new image content, such as leaves on a bare tree. Texture transfer, on the other hand, can synthesize such content but tends to destroy image structure. We propose the first algorithm that unifies color and texture transfer, outperforming both by leveraging their respective strengths. A key novelty in our approach resides in teasing apart appearance changes that can be modeled simply as changes in color versus those that require new image content to be generated. Our method starts with an analysis phase which evaluates the success of color transfer by comparing the exemplar with the source. This analysis then drives a selective, iterative texture transfer algorithm that simultaneously predicts the success of color transfer on the target and synthesizes new content where needed. We demonstrate our unified algorithm by transferring large temporal changes between photographs, such as change of season – e.g., leaves on bare trees or piles of snow on a street – and flooding. Fumio Okura, Kenneth Vanhoey, Adrien Bousseau, Alexei A. Efros, George Drettakis |
Comput. Graph. Forum | 1 |
| 2014 | Indirect augmented reality considering real-world illumination changeabstractIndirect augmented reality (IAR) utilizes pre-captured omnidirectional images and offline superimposition of virtual objects for achieving high-quality geometric and photometric registration. Meanwhile, IAR causes inconsistency between the real world and the pre-captured image. This paper describes the first-ever study focusing on the temporal inconsistency issue in IAR. We propose a novel IAR system which reflects real-world illumination change by selecting an appropriate image from a set of images pre-captured under various illumination. Results of a public experiment show that the proposed system can improve the realism in IAR. Fumio Okura, Takayuki Akaguma, Tomokazu Sato, Naokazu Yokoya |
ISMAR | 1 |
| 2013 | Teleoperation of mobile robots by generating augmented free-viewpoint imagesabstractThis paper proposes a teleoperation interface by which an operator can control a robot from freely configured viewpoints using realistic images of the physical world. The viewpoints generated by the proposed interface provide human operators with intuitive control using a head-mounted display and head tracker, and assist them to grasp the environment surrounding the robot. A state-of-the-art free-viewpoint image generation technique is employed to generate the scene presented to the operator. In addition, an augmented reality technique is used to superimpose a 3D model of the robot onto the generated scenes. Through evaluations under virtual and physical environments, we confirmed that the proposed interface improves the accuracy of teleoperation. Fumio Okura, Yuko Ueda, Tomokazu Sato, Naokazu Yokoya |
IROS | 1 |
| 2013 | Interactive exploration of augmented aerial scenes with free-viewpoint image generation from pre-rendered imagesabstractThis study proposes a framework to photorealistically synthesize virtual objects and virtualized real-world. We combine the offline rendering of virtual objects and the free-viewpoint image generation to take advantage of the higher quality of offline rendering without the computational cost of online computer graphics (CG) rendering; i.e., it incurs only the cost of the online computation for the free-viewpoint image generation. In addition, the generation of structured viewpoints (e.g., at every grid point) reduces the computational costs required to online process. Fumio Okura, Masayuki Kanbara, Naokazu Yokoya |
ISMAR | 1 |
| 2012 | Full Spherical High Dynamic Range Imaging from the SkyabstractThis paper describes a method for acquiring full spherical high dynamic range (HDR) images with no missing areas by using two omni directional cameras mounted on the top and bottom of an unmanned airship. The full spherical HDR images are generated by combining multiple omni directional images that are captured with different shutter speeds. The images generated are intended for uses in telepresence, augmented telepresence, and image-based lighting. Fumio Okura, Masayuki Kanbara, Naokazu Yokoya |
ICME | 1 |
| 2012 | Spacetime freeview generation using image-based rendering, relighting, and augmented telepresenceabstractThis paper proposes an freeview generation technique providing the users to change their viewpoints beyond time and space. The study consists of three technical elements: image-based rendering, relighting, and augmented telepresence. Before now, we have developed two systems relating this study: an augmented telepresence system and a full spherical HDR aerial imaging system. Fumio Okura |
ACM Multimedia | 1 |
| 2012 | Fly-through heijo palace site: historical tourism system using augmented telepresenceabstractWe have developed an augmented telepresence system which enables virtual tourism beyond time and space. Augmented telepresence provides a user with both the view of a remote location and related information using augmented reality techniques. This study deals with the geometric and photometric registration problems to generate movie-quality augmented omnidirectional videos automatically. The user can look around the scene from the sky above Heijo palace Site which is an ancient capital in Nara, Japan in the technical demonstration. Fumio Okura, Masayuki Kanbara, Naokazu Yokoya |
ACM Multimedia | 1 |
| 2010 | Augmented telepresence using autopilot airship and omni-directional cameraabstractThis study is concerned with a large-scale telepresence system based on remote control of mobile robot or aerial vehicle. The proposed system provides a user with not only view of remote site but also related information by AR technique. Such systems are referred to as augmented telepresence in this paper. Aerial imagery can capture a wider area at once than image capturing from the ground. However, it is difficult for a user to change position and direction of viewpoint freely because of the difficulty in remote control and limitation of hardware. To overcome these problems, the proposed system uses an autopilot airship to support changing user's viewpoint and employs an omni-directional camera for changing viewing direction easily. This paper describes hardware configuration for aerial imagery, an approach for overlaying virtual objects, and automatic control of the airship, as well as experimental results using a prototype system. Fumio Okura, Masayuki Kanbara, Naokazu Yokoya |
ISMAR | 1 |