VLDB 2026 Research / reviewers in the wild / expert
Hideo Saito 0001
dblp:12/6217
· DBLP profile ↗
137ranked-venue papers
12as first author
26since 2021 · last 2025
0000-0002-2421-9862ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 118 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 33 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 33 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | User-in-the-Loop View Sampling with Error Peaking VisualizationabstractAugmented reality (AR) provides ways to visualize missing view samples for novel view synthesis. Existing approaches present 3D annotations for new view samples and task users with taking images by aligning the AR display. This data collection task is known to be mentally demanding and limits capture areas to pre-defined small areas due to the ideal but restrictive underlying sampling theory. To free users from 3D annotations and limited scene exploration, we propose using locally reconstructed light fields and visualizing errors to be removed by inserting new views. Our results show that the error-peaking visualization is less invasive, reduces disappointment in final results, and is satisfactory with fewer view samples in our mobile view synthesis system. We also show that our approach can contribute to recent radiance field reconstruction for larger scenes, such as 3D Gaussian splatting. Ayaka Yasunaga, Hideo Saito 0001, Shohei Mori |
ICIP | 2 |
| 2025 | IntelliCap: Intelligent Guidance for Consistent View SamplingabstractNovel view synthesis from images, for example, with 3D Gaussian splatting, has made great progress. Rendering fidelity and speed are now ready even for demanding virtual reality applications. However, the problem of assisting humans in collecting the input images for these rendering algorithms has received much less attention. High-quality view synthesis requires uniform and dense view sampling. Unfortunately, these requirements are not easily addressed by human camera operators, who are in a hurry, impatient, or lack understanding of the scene structure and the photographic process. Existing approaches to guide humans during image acquisition concentrate on single objects or neglect view-dependent material characteristics. We propose a novel situated visualization technique for scanning at multiple scales. During the scanning of a scene, our method identifies important objects that need extended image coverage to properly represent view-dependent appearance. To this end, we leverage semantic segmentation and category identification, ranked by a vision-language model. Spherical proxies are generated around highly ranked objects to guide the user during scanning. Our results show superior performance in real scenes compared to conventional view sampling strategies. Ayaka Yasunaga, Hideo Saito 0001, Dieter Schmalstieg, Shohei Mori |
ISMAR | 2 |
| 2025 | Occlusion-Free 4D Gaussians for Open Surgery Videos Using Multi-camera Shadowless Lamps
Yuna Kato, Shohei Mori, Hideo Saito 0001, Yoshifumi Takatsume, Hiroki Kajita, Mariko Isogawa |
MICCAI (10) | 3 |
| 2025 | 8th ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'25)abstractThe 8th ACM International Workshop on Multimedia Content Analysis in Sports is held in Dublin, Ireland on October 28th, 2025. It is co-located with ACM Multimedia 2025. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing the multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation as well as understanding, statistical analysis, and evaluation in amateur and professional sports. There is a lack of research communities focusing on the fusion of multiple modalities. Thus, this workshop series on multimedia content analysis in sports aims to contribute to the closure of this research gap by bringing together the breadth and depth of these diverse approaches to stimulate each other with new ideas and foster research progress. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 3 |
| 2025 | Towards Predicting Any Human Trajectory In ContextabstractPredicting accurate future trajectories of pedestrians is essential for autonomous systems but remains a challenging task due to the need for adaptability in different environments and domains. A common approach involves collecting scenario-specific data and performing fine-tuning via backpropagation. However, the need to fine-tune for each new scenario is often impractical for deployment on edge devices. To address this challenge, we introduce TrajICL, an In-Context Learning (ICL) framework for pedestrian trajectory prediction that enables adaptation without fine-tuning on the scenario-specific data at inference time without requiring weight updates. We propose a spatio-temporal similarity-based example selection (STES) method that selects relevant examples from previously observed trajectories within the same scene by identifying similar motion patterns at corresponding locations. To further refine this selection, we introduce prediction-guided example selection (PG-ES), which selects examples based on both the past trajectory and the predicted future trajectory, rather than relying solely on the past trajectory. This approach allows the model to account for long-term dynamics when selecting examples. Finally, instead of relying on small real-world datasets with limited scenario diversity, we train our model on a large-scale synthetic dataset to enhance its prediction ability by leveraging in-context examples. Extensive experiments demonstrate that TrajICL achieves remarkable adaptation across both in-domain and cross-domain scenarios, outperforming even fine-tuned approaches across multiple public benchmarks. Ryo Fujii, Hideo Saito 0001, Ryo Hachiuma |
NeurIPS | 2 |
| 2025 | CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
Ryo Fujii, Ryo Hachiuma, Hideo Saito 0001 |
WACV | 3 |
| 2025 | Dense Depth from Event Focal StackabstractWe propose a method for dense depth estimation from an event stream generated when sweeping the focal plane of the driving lens attached to an event camera. In this method, a depth map is inferred from an “event focal stack” composed of the event stream using a convolutional neural network trained with synthesized event focal stacks. The synthesized event stream is created from a focal stack generated by Blender for any arbitrary 3D scene. This allows for training on scenes with diverse structures. Additionally, we explored methods to eliminate the domain gap between real event streams and synthetic event streams. Our method demonstrates superior performance over a depth-from-defocus method in the image domain on synthetic and real datasets. Kenta Horikawa, Mariko Isogawa, Hideo Saito 0001, Shohei Mori |
WACV | 3 |
| 2025 | Autoencoder-based unsupervised one-class learning for abnormal activity detection in egocentric videosabstractAbstract In recent years, abnormal human activity detection has become an important research topic. However, most existing methods focus on detecting abnormal activities of pedestrians in surveillance videos; even those methods using egocentric videos deal with the activities of pedestrians around the camera wearer. In this paper, the authors present an unsupervised auto‐encoder‐based network trained by one‐class learning that inputs RGB image sequences recorded by egocentric cameras to detect abnormal activities of the camera wearers themselves. To improve the performance of network, the authors introduce a ‘re‐encoding’ architecture and a regularisation loss function term, minimising the KL divergence between the distributions of features extracted by the first and second encoders. Unlike the common use of KL divergence loss to obtain a feature distribution close to an already‐known distribution, the aim is to encourage the features extracted by the second encoder to have a close distribution to those extracted from the first encoder. The authors evaluate the proposed method on the Epic‐Kitchens‐55 dataset and conduct an ablation study to analyse the functions of different components. Experimental results demonstrate that the method outperforms the comparison methods in all cases and demonstrate the effectiveness of the proposed re‐encoding architecture and the regularisation term. Haowen Hu, Ryo Hachiuma, Hideo Saito 0001 |
IET Comput. Vis. | 3 |
| 2025 | EventPointMesh: Human Mesh Recovery Solely From Event Point CloudsabstractHow much can we infer about human shape using an event camera that only detects the pixel position where the luminance changed and its timestamp? This neuromorphic vision technology captures changes in pixel values at ultra-high speeds, regardless of the variations in environmental lighting brightness. Existing methods for human mesh recovery (HMR) from event data need to utilize intensity images captured with a generic frame-based camera, rendering them vulnerable to low-light conditions, energy/memory constraints, and privacy issues. In contrast, we explore the potential of solely utilizing event data to alleviate these issues and ascertain whether it offers adequate cues for HMR, as illustrated in Fig. 1. This is a quite challenging task due to the substantially limited information ensuing from the absence of intensity images. To this end, we propose EventPointMesh, a framework which treats event data as a three-dimensional (3D) spatio-temporal point cloud for reconstructing the human mesh. By employing a coarse-to-fine pose feature extraction strategy, we extract both global features and local features. The local features are derived by processing the spatio-temporally dispersed event points into groups associated with individual body segments. This combination of global and local features allows the framework to achieve a more accurate HMR, capturing subtle differences in human movements. Experiments demonstrate that our method with only sparse event data outperforms baseline methods. Ryosuke Hori, Mariko Isogawa, Dan Mikami, Hideo Saito 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
Masashi Hatano, Ryo Hachiuma, Ryo Fujii, Hideo Saito 0001 |
ECCV (33) | 4 |
| 2024 | Weakly Semi-Supervised Tool Detection in Minimally Invasive Surgery VideosabstractSurgical tool detection is essential for analyzing and evaluating minimally invasive surgery videos. Current approaches are mostly based on supervised methods that require large, fully instance-level labels (i.e., bounding boxes). However, large image datasets with instance-level labels are often limited because of the burden of annotation. Thus, surgical tool detection is important when providing image-level labels instead of instance-level labels since image-level annotations are considerably more time-efficient than instance-level annotations. In this work, we propose to strike a balance between the extremely costly annotation burden and detection performance. We further propose a co-occurrence loss, which considers a characteristic that some tool pairs often co-occur together in an image to leverage image-level labels. Encapsulating the knowledge of co-occurrence using the co-occurrence loss helps to overcome the difficulty in classification that originates from the fact that some tools have similar shapes and textures. Extensive experiments conducted on the Endovis2018 dataset in various data settings show the effectiveness of our method. Ryo Fujii, Ryo Hachiuma, Hideo Saito 0001 |
ICASSP | 3 |
| 2024 | E2GS: Event Enhanced Gaussian SplattingabstractEvent cameras, known for their high dynamic range, absence of motion blur, and low energy usage, have recently found a wide range of applications thanks to these attributes. In the past few years, the field of event-based 3D reconstruction saw remarkable progress, with the Neural Radiance Field (NeRF) based approach demonstrating photorealistic view synthesis results. However, the volume rendering paradigm of NeRF necessitates extensive training and rendering times. In this paper, we introduce Event Enhanced Gaussian Splatting (E2GS), a novel method that incorporates event data into Gaussian Splatting, which has recently made significant advances in the field of novel view synthesis. Our E2GS effectively utilizes both blurry images and event data, significantly improving image deblurring and producing high-quality novel view synthesis. Our comprehensive experiments on both synthetic and real-world datasets demonstrate our E2GS can generate visually appealing renderings while offering faster training and rendering speed (140 FPS). Our code is available at https://github.com/deguchihiroyuki/E2GS. Hiroyuki Deguchi 0003, Mana Masuda, Takuya Nakabayashi, Hideo Saito 0001 |
ICIP | 4 |
| 2024 | Free-Viewpoint Visual Inspection via 3D Gaussian Splatting for Direct Template MatchingabstractMachine vision systems play a pivotal role in streamlining manufacturing processes, notably in quality control through automatic in-line visual inspections. A common practice for inspecting parts, components, and final products is to use a master part benchmark for quality comparison. However, challenges arise when objects enter inspection points in unintended orientations. This misalignment potentially leads to erroneous decisions by automated systems, resulting in additional checkpoints or wastage affecting the production rate. To tackle this issue, we propose a visual inspection pipeline that leverages recent machine learning-based approaches to compare the inspection target and a master part virtually oriented to the same perspective. Specifically, we suggest combining 3D Gaussian Splatting and DUSt3R as a practical solution. Our approach demonstrates its efficacy in real-world scenarios through testing on three mock parts and a real industrial component. Kenta Ito, Shiori Ueda, Shohei Mori, Junichi Sugano, Hideyuki Adachi, Hideo Saito 0001 |
IECON | 6 |
| 2024 | Visuo-Tactile Zero-Shot Object Recognition with Vision-Language ModelabstractTactile perception is vital, especially when distinguishing visually similar objects. We propose an approach to incorporate tactile data into a Vision-Language Model (VLM) for visuo-tactile zero-shot object recognition. Our approach leverages the zero-shot capability of VLMs to infer tactile properties from the names of tactilely similar objects. The proposed method translates tactile data into a textual description solely by annotating object names for each tactile sequence during training, making it adaptable to various contexts with low training costs. The proposed method was evaluated on the FoodReplica and Cube datasets, demonstrating its effectiveness in recognizing objects that are difficult to distinguish by vision alone. Shiori Ueda, Atsushi Hashimoto 0001, Masashi Hamaya, Kazutoshi Tanaka, Hideo Saito 0001 |
IROS | 5 |
| 2024 | EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos
Ryo Fujii, Masashi Hatano, Hideo Saito 0001, Hiroki Kajita |
MICCAI (6) | 3 |
| 2023 | Learning Food Picking without Food: Fracture Anticipation by Breaking Reusable Fragile ObjectsabstractFood picking is trivial for humans but not for robots, as foods are fragile. Presetting foods' physical properties does not help robots much due to the objects' inter- and intra-category diversity. A recent study proved that learning-based fracture anticipation with tactile sensors could overcome this problem; however, the method trains the model for each food to deal with intra-category differences, and tuning robots for each food leads to an undesirable amount of food consumption. This study proposes a novel framework for learning food-picking tasks without consuming foods. The key idea is to leverage the object-breaking experiences of several reusable fragile objects instead of consuming real foods while making the picking ability object-invariant with domain generalization (DG). In real-robot experiments, we trained a model with reusable objects (toy blocks, ping-pong balls, and jellies), selected based on the three common fracture types (crack, rupture, and crush). We then tested the model with four real food objects (tofu, bananas, potato chips, and tomatoes). The results showed that the proposed combination of reusable objects' breaking experiences and DG is effective for the food-picking task. Rinto Yagawa, Reina Ishikawa, Masashi Hamaya, Kazutoshi Tanaka, Atsushi Hashimoto 0001, Hideo Saito 0001 |
ICRA | 6 |
| 2023 | High-Quality Virtual Single-Viewpoint Surgical Video: Geometric Autocalibration of Multiple Cameras in Surgical Lights
Yuna Kato, Mariko Isogawa, Shohei Mori, Hideo Saito 0001, Hiroki Kajita, Yoshifumi Takatsume |
MICCAI (9) | 4 |
| 2023 | MMSports '23: 6th International Workshop on Multimedia Content Analysis in SportsabstractThe sixth ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'23) is part of the ACM International Conference on Multimedia 2023 (ACM Multimedia 2023). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing multimedia/multimodal data in sports, sports broadcasts, sports games and sports medicine. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation and understanding, for statistical analysis and evaluation, and for sensor fusion during workouts as well as competitions. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Hideo Saito 0001, Thomas B. Moeslund, Rainer Lienhart |
ACM Multimedia | 1 |
| 2023 | Multi-Layer Scene Representation from Composed Focal StacksabstractMulti-layer images are a powerful scene representation for high-performance rendering in virtual/augmented reality (VR/AR). The major approach to generate such images is to use a deep neural network trained to encode colors and alpha values of depth certainty on each layer using registered multi-view images. A typical network is aimed at using a limited number of nearest views. Therefore, local noises in input images from a user-navigated camera deteriorate the final rendering quality and interfere with coherency over view transitions. We propose to use a focal stack composed of multi-view inputs to diminish such noises. We also provide theoretical analysis for ideal focal stacks to generate multi-layer images. Our results demonstrate the advantages of using focal stacks in coherent rendering, memory footprint, and AR-supported data capturing. We also show three applications of imaging for VR. Reina Ishikawa, Hideo Saito 0001, Denis Kalkofen, Shohei Mori |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Shared Transformer Encoder with Mask-Based 3d Model Estimation for Container Mass EstimationabstractFor human-safe robot control in human-to-robot handover, the physical properties of containers and fillings should be accurately estimated. In this paper, we propose a Transformer encoder that shares the same architecture and parameters for filling level and type estimation. We also propose a mask-based geometric algorithm to estimate 3D models of containers for the estimation of their capacity and dimensions. We further use these estimations to estimate their mass in a Convolutional Neural Network model. Experiments show that our Transformer model produced encouraging results in both estimations. While challenges remain in our mask-based algorithm and Convolutional Neural Network model, their results revealed several ways for improvement. Tomoya Matsubara, Seitaro Otsuki, Yuiga Wada, Haruka Matsuo, Takumi Komatsu, Yui Iioka, Komei Sugiura, Hideo Saito 0001 |
ICASSP | 8 |
| 2022 | Neural Implicit Event Generator for Motion TrackingabstractWe present a novel framework of motion tracking from event data using implicit expression. Our framework uses pre-trained event generation MLP called the implicit event generator (IEG) and carries out motion tracking by updating its state (position and velocity) based on the difference between the observed event and generated event from the current state estimation. The difference is computed implicitly by the IEG. Unlike the conventional explicit approach, which requires dense computation to evaluate the difference, our implicit approach realizes the update of the efficient state directly from sparse event data. Our sparse algorithm is especially suitable for mobile robotics applications in which computational resources and battery life are limited. To verify the effectiveness of our method on real-world data, we applied it to the AR marker tracking application. We have confirmed that our framework works well in real-world environments in the presence of noise and background clutter. Mana Masuda, Yusuke Sekikawa, Ryo Fujii, Hideo Saito 0001 |
ICRA | 4 |
| 2022 | MMSports'22: 5th International ACM Workshop on Multimedia Content Analysis in SportsabstractThe fifth ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'22) is part of the ACM International Conference on Multimedia 2022 (ACM Multimedia 2022). After two years of pure virtual MMSports workshops due to COVID-19, MMSports'22 is held on-site again. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing multimedia/multimodal data in sports, sports broadcasts, sports games and sports medicine. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation and understanding, for statistical analysis and evaluation, and for sensor fusion during workouts as well as competitions. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Related Workshop Proceedings are available in the ACM DL at: https://dl.acm.org/doi/proceedings/10.1145/3552437. Hideo Saito 0001, Thomas B. Moeslund, Rainer Lienhart |
ACM Multimedia | 1 |
| 2022 | Special issue on "drones as enablers of novel services: operational and technology challenges"
Ana M. Bernardos, Juan A. Besada, Jesús García 0001, Hideo Saito 0001, Patrizia Marti |
Pers. Ubiquitous Comput. | 4 |
| 2021 | Silhouette-Based Synthetic Data Generation For 3D Human Pose Estimation With A Single Wrist-Mounted 360° CameraabstractIn this paper, we propose a framework for 3D human pose estimation with a single 360° camera mounted on the user’s wrist. Perceiving a 3D human pose with such a simple setting has remarkable potential for various applications (e.g., daily-living activity monitoring, motion analysis for sports enhancement). However, no existing work has tackled this task due to the difficulty of estimating a human pose from a single camera image in which only a part of the human body is captured and the lack of training data. Therefore, we propose an effective method for translating wrist-mounted 360° camera images into 3D human poses. We also propose silhouette-based synthetic data generation dedicated to this task, which enables us to bridge the domain gap between real-world data and synthetic data. We achieved higher estimation accuracy quantitatively and qualitatively compared with other baseline methods. Ryosuke Hori, Ryo Hachiuma, Hideo Saito 0001, Mariko Isogawa, Dan Mikami |
ICIP | 3 |
| 2021 | Toward Unsupervised 3d Point Cloud Anomaly Detection Using Variational AutoencoderabstractIn this paper, we present an end-to-end unsupervised anomaly detection framework for 3D point clouds. To the best of our knowledge, this is the first work to tackle the anomaly detection task on a general object represented by a 3D point cloud. We propose a deep variational autoencoder based unsupervised anomaly detection network adapted to the 3D point cloud and an anomaly score specifically for 3D point clouds. To verify the effectiveness of the model, we conducted extensive experiments on ShapeNet dataset. Through quantitative and qualitative evaluation, we demonstrate that the proposed method outperforms the baseline method. Mana Masuda, Ryo Hachiuma, Ryo Fujii, Hideo Saito 0001, Yusuke Sekikawa |
ICIP | 4 |
| 2021 | MMSports'21: 4th International Workshop on Multimedia Content Analysis in Sports
Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 3 |
| 2020 | Single-modal Incremental Terrain Clustering from Self-Supervised Audio-Visual Feature LearningabstractThe key to an accurate understanding of terrain is to extract the informative features from the multi-modal data obtained from different devices. Sensors, such as RGB cameras, depth sensors, vibration sensors, and microphones, are used as the multi-modal data. Many studies have explored ways to use them, especially in the robotics field. Some papers have successfully introduced single-modal or multi-modal methods. However, in practice, robots can be faced with extreme conditions; microphones do not work well in crowded scenes, and an RGB camera cannot capture terrains well in the dark. In this paper, we present a novel framework using the multimodal variational autoencoder and the Gaussian mixture model clustering algorithm on image data and audio data for terrain type clustering. Our method enables the terrain type clustering even if one of the modalities (either image or audio) is missing at the test-time. We evaluated the clustering accuracy with a conventional multi-modal terrain type clustering method and we conducted ablation studies to show the effectiveness of our approach. Reina Ishikawa, Ryo Hachiuma, Akiyoshi Kurobe, Hideo Saito 0001 |
ICPR | 4 |
| 2020 | Deep Selection: A Fully Supervised Camera Selection Network for Surgery Recordings
Ryo Hachiuma, Tomohiro Shimizu, Hideo Saito 0001, Hiroki Kajita, Yoshifumi Takatsume |
MICCAI (3) | 3 |
| 2020 | MMSports'20: 3rd International Workshop on Multimedia Content Analysis in SportsabstractThe third ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'20) is part of the ACM International Conference on Multimedia 2020 (ACM Multimedia 2020). Exceptionally, due to the corona pandemic, the workshop is held virtually. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding and visualizing the multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 3 |
| 2020 | Video-Annotated Augmented Reality Assembly TutorialsabstractWe present a system for generating and visualizing interactive 3D Augmented Reality tutorials based on 2D video input, which allows viewpoint control at runtime. Inspired by assembly planning, we analyze the input video using a 3D CAD model of the object to determine an assembly graph that encodes blocking relationships between parts. Using an assembly graph enables us to detect assembly steps that are otherwise difficult to extract from the video, and generally improves object detection and tracking by providing prior knowledge about movable parts. To avoid information loss, we combine the 3D animation with relevant parts of the 2D video so that we can show detailed manipulations and tool usage that cannot be easily extracted from the video. To further support user orientation, we visually align the 3D animation with the real-world object by using texture information from the input video. We developed a presentation system that uses commonly available hardware to make our results accessible for home use and demonstrate the effectiveness of our approach by comparing it to traditional video tutorials. Masahiro Yamaguchi 0002, Shohei Mori, Peter Mohr, Markus Tatzgern, Ana Stanescu 0003, Hideo Saito 0001, Denis Kalkofen |
UIST | 6 |
| 2020 | InpaintFusion: Incremental RGB-D Inpainting for 3D ScenesabstractState-of-the-art methods for diminished reality propagate pixel information from a keyframe to subsequent frames for real-time inpainting. However, these approaches produce artifacts, if the scene geometry is not sufficiently planar. In this article, we present InpaintFusion, a new real-time method that extends inpainting to non-planar scenes by considering both color and depth information in the inpainting process. We use an RGB-D sensor for simultaneous localization and mapping, in order to both track the camera and obtain a surfel map in addition to RGB images. We use the RGB-D information in a cost function for both the color and the geometric appearance to derive a global optimization for simultaneous inpainting of color and depth. The inpainted depth is merged in a global map by depth fusion. For the final rendering, we project the map model into image space, where we can use it for effects such as relighting and stereo rendering of otherwise hidden structures. We demonstrate the capabilities of our method by comparing it to inpainting results with methods using planar geometric proxies. Shohei Mori, Okan Erat, Wolfgang Broll, Hideo Saito 0001, Dieter Schmalstieg, Denis Kalkofen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Mobile Photometric Stereo with Keypoint-Based SLAM for Dense 3D ReconstructionabstractThe standard photometric stereo is a technique to densely reconstruct objects' surfaces using light variation under the assumption of a static camera with a moving light source. In this work, we use photometric stereo to reconstruct dense 3D scenes while moving the camera and the light altogether. In such non-static case, camera poses as well as correspondences between pixels of each frame to apply photometric stereo are required. ORB-SLAM is a technique that can be used to estimate camera poses. To retrieve correspondences, our idea is to start from a sparse 3D mesh obtained with ORB SLAM and then densify the mesh by a plane sweep method using a multi-view photometric consistency. By combining ORB-SLAM and photometric stereo, it is possible to reconstruct dense 3D scenes with a off-the-shelf smartphone and its embedded torchlight. Note that SLAM systems usually struggle with textureless object, which is effectively compensated by the photometric stereo in our method. Experiments are conducted to show that our proposed method gives better results than SLAM alone or COLMAP, especially for partially textureless surfaces. Remy Maxence, Hideaki Uchiyama, Hiroshi Kawasaki, Diego Thomas, Vincent Nozick, Hideo Saito 0001 |
3DV | 6 |
| 2019 | DetectFusion: Detecting and Segmenting Both Known and Unknown Dynamic Objects in Real-time SLAM
Ryo Hachiuma, Christian Pirchheim, Dieter Schmalstieg, Hideo Saito 0001 |
BMVC | 4 |
| 2019 | EventNet: Asynchronous Recursive Event ProcessingabstractEvent cameras are bio-inspired vision sensors that mimic retinas to asynchronously report per-pixel intensity changes rather than outputting an actual intensity image at regular intervals. This new paradigm of image sensor offers significant potential advantages; namely, sparse and non-redundant data representation. Unfortunately, however, most of the existing artificial neural network architectures, such as a CNN, require dense synchronous input data, and therefore, cannot make use of the sparseness of the data. We propose EventNet, a neural network designed for real-time processing of asynchronous event streams in a recursive and event-wise manner. EventNet models dependence of the output on tens of thousands of causal events recursively using a novel temporal coding scheme. As a result, at inference time, our network operates in an event-wise manner that is realized with very few sum-of-the-product operations---look-up table and temporal feature aggregation---which enables processing of 1 mega or more events per second on standard CPU. In experiments using real data, we demonstrated the real-time performance and robustness of our framework. Yusuke Sekikawa, Kosuke Hara, Hideo Saito 0001 |
CVPR | 3 |
| 2019 | Incremental Class Discovery for Semantic Segmentation With RGBD Sensing
Yoshikatsu Nakajima, Byeongkeun Kang, Hideo Saito 0001, Kris Makoto Kitani |
ICCV | 3 |
| 2019 | MMSports'19: 2nd ACM International Workshop on Multimedia Content Analysis in SportsabstractThe second ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'19) is held in Nice, France on October 25th, 2019 co-located with the ACM International Conference on Multimedia 2019 (ACM Multimedia 2019). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding and visualizing the multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 3 |
| 2019 | Removing fences from sweep motion videos using global 3D reconstruction and fence-aware light field renderingabstractDiminishing the appearance of a fence in an image is a challenging research area due to the characteristics of fences (thinness, lack of texture, etc.) and the need for occluded background restoration. In this paper, we describe a fence removal method for an image sequence captured by a user making a sweep motion, in which occluded background is potentially observed. To make use of geometric and appearance information such as consecutive images, we use two well-known approaches: structure from motion and light field rendering. Results using real image sequences show that our method can stably segment fences and preserve background details for various fence and background combinations. A new video without the fence, with frame coherence, can be successfully provided. Chanya Lueangwattana, Shohei Mori, Hideo Saito 0001 |
Comput. Vis. Media | 3 |
| 2018 | Constant Velocity 3D ConvolutionabstractWe propose a novel three-dimensional (3D)-convolution method, cv3dconv, for detecting spatiotemporal features from videos. It reduces the number of sum-of-products of 3D convolution by thousands of times by assuming the constant moving velocity of the camera. We observed that a specific class of video sequences, such as those captured by an in-vehicle camera, can be well approximated with piece-wise linear movements of 2D features in the temporal dimension. Our principal finding is that the 3D kernel, represented by the constant-velocity, can be decomposed into a convolution of a 2D kernel representing the shapes and a 3D kernel representing the velocity. We derived the efficient recursive algorithm for this class of 3D convolution which is exceptionally suited for sparse data, and this parameterized decomposed representation imposes a structured regularization along the temporal direction. We experimentally verified the validity of our approximation using a controlled dataset, and we also showed the effectiveness of cv3dconv for the visual odometry estimation task using real event camera data captured in urban road scene. Yusuke Sekikawa, Kohta Ishikawa, Kosuke Hara, Yuichi Yoshida, Koichiro Suzuki, Ikuro Sato, Hideo Saito 0001 |
3DV | 7 |
| 2018 | Fast and Accurate Semantic Mapping through Geometric-based Incremental SegmentationabstractWe propose an efficient and scalable method for incrementally building a dense, semantically annotated 3D map in real-time. The proposed method assigns class probabilities to each region, not each element (e.g., surfel and voxel), of the 3D map which is built up through a robust SLAM framework and incrementally segmented with a geometric-based segmentation method. Differently from all other approaches, our method has a capability of running at over 30Hz while performing all processing components, including SLAM, segmentation, 2D recognition, and updating class probabilities of each segmentation label at every incoming frame, thanks to the high efficiency that characterizes the computationally intensive stages of our framework. By utilizing a specifically designed CNN to improve the frame-wise segmentation result, we can also achieve high accuracy. We validate our method on the NYUv2 dataset by comparing with the state of the art in terms of accuracy and computational efficiency, and by means of an analysis in terms of time and space complexity. Yoshikatsu Nakajima, Keisuke Tateno, Federico Tombari, Hideo Saito 0001 |
IROS | 4 |
| 2018 | 1st ACM International Workshop on Multimedia Content Analysis in SportsabstractThe first ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'18) is held in Seoul, South Korea on October 26th, 2018 and is co-located with the ACM International Conference on Multimedia 2018 (ACM Multimedia 2018). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining and content analysis of multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this first workshop of a serious workshops on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 3 |
| 2018 | FusionMLS: Highly dynamic 3D reconstruction with consumer-grade RGB-D camerasabstractMulti-view dynamic three-dimensional reconstruction has typically required the use of custom shutter-synchronized camera rigs in order to capture scenes containing rapid movements or complex topology changes. In this paper, we demonstrate that multiple unsynchronized low-cost RGB-D cameras can be used for the same purpose. To alleviate issues caused by unsynchronized shutters, we propose a novel depth frame interpolation technique that allows synchronized data capture from highly dynamic 3D scenes. To manage the resulting huge number of input depth images, we also introduce an efficient moving least squares-based volumetric reconstruction method that generates triangle meshes of the scene. Our approach does not store the reconstruction volume in memory, making it memory-efficient and scalable to large scenes. Our implementation is completely GPU based and works in real time. The results shown herein, obtained with real data, demonstrate the effectiveness of our proposed method and its advantages compared to state-of-the-art approaches. Siim Meerits, Diego Thomas, Vincent Nozick, Hideo Saito 0001 |
Comput. Vis. Media | 4 |
| 2018 | Extrinsic Camera Calibration Without Visible Corresponding Points Using Omnidirectional CamerasabstractThis paper proposes a novel algorithm that calibrates multiple cameras scattered across a broad area. The key idea of the proposed method is “using the position of an omnidirectional camera as a reference point.” The common approach to calibrating multiple cameras assumes that the cameras capture at least some common points. This means calibration becomes quite difficult if there are no shared points in each camera's field of view (FOV). The proposed method uses the position of an omnidirectional camera to determine point correspondence. The position of an omnidirectional camera relative to the calibrated camera is estimated by the theory of epipolar geometry, even if the omnidirectional camera is placed outside the camera's FOV. This property makes our method applicable to multiple cameras scattered across a broad area. Qualitative and quantitative evaluations using synthesized and real data, e.g., a sports field, demonstrate the advantages of the proposed method. Shogo Miyata, Hideo Saito 0001, Kosuke Takahashi, Dan Mikami, Mariko Isogawa, Akira Kojima |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Message from the ISMAR 2016 Science and Technology Program Chairs and Guest EditorsabstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Wolfgang Broll, Hideo Saito 0001, J. Edward Swan II |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Eating and Drinking Recognition via Integrated Information of Head Directions and Joint Positions in a Group
Naoto Ienaga, Yuko Ozasa, Hideo Saito 0001 |
ICPRAM | 3 |
| 2017 | Diminished hand: A diminished reality-based work area visualizationabstractLive instructors perspective videos are useful to present intuitive visual instructions for trainees in medical and industrial settings. In such videos, the instructors hands often hide the work area. In this demo, we present a diminished hand for visualizing the work area hidden by hands by capturing the work area with multiple cameras. To achieve the diminished reality, we use a light field rendering technique, in which light rays avoid passing through penalty points set in the unstructured light fields reconstructed from the multiple viewpoint images. Shohei Mori, Momoko Maezawa, Naoto Ienaga, Hideo Saito 0001 |
VR | 4 |
| 2017 | Robust camera pose estimation by viewpoint classification using deep learningabstractCamera pose estimation with respect to target scenes is an important technology for superimposing virtual information in augmented reality (AR). However, it is difficult to estimate the camera pose for all possible view angles because feature descriptors such as SIFT are not completely invariant from every perspective. We propose a novel method of robust camera pose estimation using multiple feature descriptor databases generated for each partitioned viewpoint, in which the feature descriptor of each keypoint is almost invariant. Our method estimates the viewpoint class for each input image using deep learning based on a set of training images prepared for each viewpoint class. We give two ways to prepare these images for deep learning and generating databases. In the first method, images are generated using a projection matrix to ensure robust learning in a range of environments with changing backgrounds. The second method uses real images to learn a given environment around a planar pattern. Our evaluation results confirm that our approach increases the number of correct matches and the accuracy of camera pose estimation compared to the conventional method. Yoshikatsu Nakajima, Hideo Saito 0001 |
Comput. Vis. Media | 2 |
| 2016 | Synthesis of a stroboscopic image from a hand-held camera sequence for a sports analysisabstractThis paper presents a method for synthesizing a stroboscopic image of a moving sports player from a hand-held camera sequence. This method has three steps: synthesis of background image, synthesis of stroboscopic image, and removal of player’s shadow. In synthesis of background image step, all input frames masked a bounding box of the player are stitched together to generate a background image. The player is extracted by an HOG-based people detector. In synthesis of stroboscopic image step, the background image, the input frame, and a mask of the player synthesize a stroboscopic image. In removal of shadow step, we remove the player’s shadow which negatively affects an analysis by using mean-shift. In our previous work, synthesis of background image has been timeconsuming. In this paper, by using the bounding box of the player detected by HOG and by subtracting the images for synthesizing a mask, computational speed and accuracy can be improved. These have contributed greatly to the improvement from the previous method. These are main improvements and novelty points from our previous method. In experiments, we confirmed the effectiveness of the proposed method, measured the player’s speed and stride length, and made a footprint image. The image sequence was captured under a simple condition that no other people were in the background and the person controlling the video camera was standing still, such like a motion parallax was not occurred. In addition, we applied the synthesis method to various scenes to confirm its versatility. Kunihiro Hasegawa, Hideo Saito 0001 |
Comput. Vis. Media | 2 |
| 2015 | An omnidirectional vision system for bus safety surveillanceabstractSudden brake seems to be unavoidable in driving practice to prevent from accidence. The sudden brake may often cause falling risks to on-board passengers in public transport and even threaten the lives of the senior. Thus, this paper presents an omnidirectional vision system to assist bus drivers in ensuring the safety for passengers. To this end, we configure an auxiliary camera in the boarding gate to recognize the age of passengers and an omnidirectional camera on the ceiling of the bus to detect and track the passengers. Their trajectories and positions are mapped into a bus layout, together with their age information transferred from the auxiliary camera. This bus layout is shown on a small display in front of bus drivers. Instead of looking to the mirror to guess and gather information of passengers, the driver can collect such necessary information for safe driving. We set up experiments on a real bus and demonstrate the system by preliminary results. Dao Huu Hung, Hideo Saito 0001, Keiichi Yamamoto, Hiromitsu Sato |
AVSS | 2 |
| 2015 | Combination Photometric Stereo Using Compactness of Albedo and Surface Normal in the Presence of Shadows and Specular Reflection
Naoto Ienaga, Hideo Saito 0001, Kouichi Tezuka, Yasumasa Iwamura, Masayoshi Shimizu |
CAIP (1) | 2 |
| 2015 | Motion estimation for non-overlapping cameras by improvement of feature points matching based on urban 3D structureabstractWe propose a method of ego-motion estimation for a self-driving vehicle using multiple cameras. By finding corresponding points between the multi-camera images, we aim to enhance the accuracy of the ego-motion estimation. However since the viewing directions are very different from one camera to the other, a conventional algorithm such as SURF cannot detect a sufficient number of correspondences. We propose a novel matching algorithm by warping feature patches detected in different cameras based on urban 3D structure. We assume that detected features exist on the surface of buildings or roads and the patch around the feature is planar. Based on this assumption, we can warp the patches so that the feature descriptors are similar for the corresponding feature points. We apply Bundle Adjustment to the found correspondences to optimizes the odometry. The result shows higher estimation accuracy when compared to other matching method. Atsushi Kawasaki, Hideo Saito 0001, Kosuke Hara |
ICIP | 2 |
| 2015 | Plane Fitting and Depth Variance Based Upsampling for Noisy Depth Map from 3D-ToF Cameras in Real-time
Kazuki Matsumoto, François de Sorbier, Hideo Saito 0001 |
ICPRAM (2) | 3 |
| 2015 | Retrieving Lights Positions Using Plane Segmentation with Diffuse Illumination Reinforced with Specular ComponentabstractWe present a novel method to retrieve multiple positions of point lights in real indoor scenes based on a 3D reconstruction. This method takes advantage of illumination over planes detected using a segmentation of the reconstructed mesh of the scene. We can also provide an estimation without suffering from the presence of specular highlights but rather use this component to refine the final estimation. This allows consistent relighting throughout the entire scene for aumented reality purposes. Paul-Emile Buteau, Hideo Saito 0001 |
ISMAR | 2 |
| 2015 | Remote Welding Robot Manipulation Using Multi-view ImagesabstractThis paper proposes a remote welding robot manipulation system by using multi-view images. After an operator specifies two-dimensional path on images, the system transforms it into three-dimensional path and displays the movement of the robot by overlaying graphics with images. The accuracy of our system is sufficient to weld objects when combining with a sensor in the robot. The system allows the non-expert operator to weld objects remotely and intuitively, without the need to create a 3D model of a processed object beforehand. Yuichi Hiroi, Kei Obata, Katsuhiro Suzuki, Naoto Ienaga, Maki Sugimoto, Hideo Saito 0001, Tadashi Takamaru |
ISMAR | 6 |
| 2015 | Rubix: Dynamic Spatial Augmented Reality by Extraction of Plane Regions with a RGB-D CameraabstractDynamic spatial augmented reality requires accurate real-time 3D pose information of the physical objects that are to be projected onto. Previous depth-based methods for tracking objects required strong features to enable recognition; making it difficult to estimate an accurate 6DOF pose for physical objects with a small set of recognizable features (such as a non-textured cube). We propose a more accurate method with fewer limitations for the pose estimation of a tangible object that has known planar faces and using depth data from an RGB-D camera only. In this paper, the physical object's shape is limited to cubes of different sizes. We apply this new tracking method to achieve dynamic projections onto these cubes. In our method, 3D points from an RGB-D camera are divided into a cluster of planar regions, and the point cloud inside each face of the object is fitted to an already-known geometric model of a cube. With the 6DOF pose of the physical object, SAR generated imagery is then projected correctly onto the physical object. The 6DOF tracking is designed to support tangible interactions with the physical object. We implemented example interactive applications with one or multiple cubes to show the capability of our method. Masayuki Sano, Kazuki Matsumoto, Bruce H. Thomas, Hideo Saito 0001 |
ISMAR | 4 |
| 2014 | Accurate Camera Pose Estimation for KinectFusion Based on Line Segment Matching by LEHFabstractKinect Fusion is able to build a 3D reconstruction in real time and provide a 3D model. Kinect Fusion uses Iterative Closest Point (ICP) algorithm for point cloud alignment from the each camera frame and estimates each camera pose. However, ICP algorithm has its limits and the camera poses lack in accuracy. We propose an alignment method which is not only based on point cloud but also line segments. This method significantly improve the camera pose accuracy obtained from Kinect Fusion and creates better 3D model. In this method, we use line segment matching by Line-based Eight-directional Histogram Feature(LEHF). We also propose an improved version of LEHF for this alignment method. The basic idea is to get a set of 2D-3D line segment correspondences between 2D line segments on camera images and 3D line segments of 3D line segment based models, to solve the PnL problem and to recompute the camera pose. The experimental result that the camera pose estimated by our method is more accurate than the original one obtained from Kinect Fusion. Yusuke Nakayama, Toshihiro Honda, Hideo Saito 0001, Masayoshi Shimizu, Nobuyasu Yamaguchi |
ICPR | 3 |
| 2014 | RGB-D-T camera system for AR display of temperature changeabstractThe anomalies of power equipment can be founded using temperature changes compared to its normal state. In this paper we present a system for visualizing temperature changes in a scene using a thermal 3D model. Our approach is based on two precomputed 3D models of the target scene achieved with a RGB-D camera coupled with the thermal camera. The first model contains the RGB information, while the second one contains the thermal information. For comparing the status of the temperature between the model and the current time, we accurately estimate the pose of the camera by finding keypoint correspondences between the current view and the RGB 3D model. Knowing the pose of the camera, we are then able to compare the thermal 3D model with the current status of the temperature from any viewpoint. Kazuki Matsumoto, Wataru Nakagawa, François de Sorbier, Maki Sugimoto, Hideo Saito 0001, Shuji Senda, Takashi Shibata 0001, Akihiko Iketani |
ISMAR | 5 |
| 2014 | Tablet system for visual, overlay of 3D virtual object onto real environmentabstractWe propose a novel system for visual overlay of 3D virtual object onto real environment observed by tablet PC with camera. This system allows us to visually simulate the layout of virtual 3D objects such as furniture in the real environment captured by the tablet PC. For estimating the pose and position of the tablet PC in the 3D structure of the target environment, we propose and implement the 2 procedures using the captured image and using the motion sensor in tablet PC. Those performances are presented in the demonstration. Hiroyuki Yoshida, Takuya Okamoto, Hideo Saito 0001 |
ISMAR | 3 |
| 2014 | An AR edutainment system supporting bone anatomy learningabstractWe present a medical Augmented Reality (AR) edutainment system for bone anatomy learning. This learning environment, called AR bone puzzle, is a metaphor for bone anatomy learning with AR visualization and intuitive interaction. AR bone puzzle uses its user's body as a puzzle frame and computer generated virtual bones as puzzle pieces. Users learn bone anatomy by assembling the virtual bone pieces on their body. Key features of this system are 3D AR visualization and intuitive gesture based user interaction. Philipp Stefan, Patrick Wucherer, Yuji Oyamada, Alexander Schoch, Motoko Kanegae, Naoki Shimizu, Tatsuya Kodera, Sebastien Cahier, Matthias Weigl, Maki Sugimoto, Pascal Fallavollita, Hideo Saito 0001, Nassir Navab |
VR | 13 |
| 2013 | Virtual Slicer: Development of Interactive Visualizer for Tomographic Medical Images Based on Position and Orientation of Handheld DeviceabstractThis paper proposes an interface that helps understanding the correspondence between the patient and medical images. In our proposed method, we have developed an interactive visualizer for tomographic images based on the relative position and orientation of the handheld device and the patient. Sho Shimamura, Motoko Kanegae, Yuji Uema, Masahiko Inami, Tetsu Hayashida, Hideo Saito 0001, Maki Sugimoto |
CW | 6 |
| 2013 | Region-based tracking using sequences of relevance measuresabstractWe present the preliminary results of our proposal: a region-based detection and tracking method of arbitrary shapes. The method is designed to be robust against orientation and scale changes and also occlusions. In this work, we study the effectiveness of sequence of shape descriptors for matching purpose. We detect and track surfaces by matching the sequences of descriptor so called relevance measures with their correspondences in the database. First, we extract stable shapes as the detection target using Maximally Stable Extreme Region (MSER) method. The keypoints on the stable shapes are then extracted by simplifying the outline of the stable regions. The relevance measures that are composed by three keypoints are then computed and the sequences of them are composed as descriptors. During runtime, the sequences of relevance measures are extracted from the captured image and are matched with those in the database. When a particular region is matched with one in the database, the orientation of the region is then estimated and virtual annotations can be superimposed. We apply this approach in an interactive task support system that helps users for creating paper craft objects. Sandy Martedi, Bruce H. Thomas, Hideo Saito 0001 |
ISMAR | 3 |
| 2012 | Fast Line Description for Line-based SLAM
Keisuke Hirose, Hideo Saito 0001 |
BMVC | 2 |
| 2012 | Illumination estimation from shadow and incomplete object shape captured by an RGB-D camera
Takuya Ikeda, Yuji Oyamada, Maki Sugimoto, Hideo Saito 0001 |
ICPR | 4 |
| 2012 | Calibration-free projector-camera system for spatial augmented reality on planar surfaces
Takayuki Nakamura, François de Sorbier, Sandy Martedi, Hideo Saito 0001 |
ICPR | 4 |
| 2012 | Workshop 3: IEEE ISMAR 2012 workshop on tracking methods and applications (TMA)abstractThe focus of this workshop is on presenting, discussing and demonstrating recent tracking methods and applications that work well in practice and that show some superiority over state-of-the-art methods. Rather than focusing on pure novelty, this workshop encourages presentations that concentrate on complete systems and integrated approaches. The TMA workshop looks at pose tracking from an end-to-end point of view. Daniel Wagner 0003, Jonathan Ventura, Gerhard Reitmayr, Hideo Saito 0001, Selim Benhimane |
ISMAR | 4 |
| 2011 | Template-based methods for sentence generation and speech synthesisabstractHere we propose a sentence-generation method using templates that can be applied to create a speech database. This method requires the recording of a relatively small sentence set, and the resultant speech database can generate comparatively natural sounding synthesized speech. Applying this method to the Japan Broadcasting Corporation (NHK) weather report radio program reduced the size of the required sentence set to just a fraction of that required by comparable methods. We also propose a speech-synthesis method using templates. In an evaluation test, 66% of the speech samples synthesized by the proposed method using templates were preferred to those produced by the conventional concatenative speech synthesis method. Hiroyuki Segi, Reiko Takou, Nobumasa Seiyama, Tohru Takagi, Hideo Saito 0001, Shinji Ozawa |
ICASSP | 5 |
| 2011 | Random dot markersabstractThis paper presents a novel approach for detecting and tracking markers with randomly scattered dots for augmented reality applications. Compared with traditional markers with square pattern, our random dot markers have several significant advantages for flexible marker design, robustness against occlusion and user interaction. The retrieval and tracking of these markers are based on geometric feature based keypoint matching and tracking. We experimentally demonstrate that the discriminative ability of forty random dots per marker is applicable for retrieving up to one thousand markers. Hideaki Uchiyama, Hideo Saito 0001 |
VR | 2 |
| 2011 | Random dot markersabstractWe introduce a novel type of markers with randomly scattered dots for augmented reality applications. Compared with traditional square markers, our markers have several significant advantages for flexible marker design, robustness against occlusion and user interaction. Our markers do not need to have a black frame, and their shape is not limited to square because the retrieval and tracking of the markers are based on geometric feature based keypoint matching. In our demonstration, we show real-time simultaneous retrieval and tracking of the markers on a laptop. Hideaki Uchiyama, Hideo Saito 0001 |
VR | 2 |
| 2011 | Guest Editors' Introduction: Special Section on the IEEE International Symposium on Mixed and Augmented Reality (ISMAR)abstractThe two papers in this special section are extended versions of papers originally presented at the International Symposium on Mixed and Augmented Reality (ISMAR) 2007. These two papers won awards at the symposium. Gudrun Klinker, Tobias Höllerer, Hideo Saito 0001, Oliver Bimber |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2010 | Cepstral analysis based blind deconvolution for motion blurabstractCamera shake during exposure blurs the captured image. Despite several decades of studies, image deconvolution to restore a blurred image still remains an issue, particularly in blind deconvolution cases in which the actual shape of the blur is unknown. Approaches based on cepstral analysis succeeded in restoring images degraded by a uniform blur caused by a camera moving straight in a single direction. In this paper, we propose to estimate, from a single blurred image, the point spread function (PSF) caused by a normal camera undergoing a 2D curved motion, and to restore the image. To extend the traditional cepstral analysis, we derive assumptions about the PSF effects in the cepstrum domain. In a first phase, we estimate several PSF candidates from the cepstrum of a blurred image and restore the image with a fast deconvolution algorithm. In a second phase, we select the best PSF candidate by evaluating the restored images. Finally, a slower but more accurate deconvolution algorithm recovers the latent image with the chosen PSF. We validate the proposed method with synthetic and real experiments. Haruka Asai, Yuji Oyamada, Julien Pilet, Hideo Saito 0001 |
ICIP | 4 |
| 2010 | An Augmented Reality Setup with an Omnidirectional Camera Based on Multiple Object DetectionabstractWe propose a novel augmented reality (AR) setup with an omni directional camera on a table top display. The table acts as a mirror on which real playing cards appear augmented with virtual elements. The omni directional camera captures and recognizes its surrounding based on a feature based image retrieval approach which achieves fast and scalable registration. It allows our system to superimpose virtual visual effects to the omni directional camera image. In our AR card game, users sit around a table top display and show a card to the other players. The system recognizes it and augments it with virtual elements in the omni directional image acting as a mirror. While playing the game, the users can interact with each other directly and through the display. Our setup is a new, simple, and natural approach to augmented reality. It opens new doors to traditional card games. Tomoki Hayashi, Hideaki Uchiyama, Julien Pilet, Hideo Saito 0001 |
ICPR | 4 |
| 2010 | Video Retrieval Based on Tracked Features QuantizationabstractIn this paper, we present an image retrieval method based on feature tracking. Feature tracks are summarized into a compact discreet value and used for video indexing purpose. As opposed to existing space-time features, we do not make any assumption on the motion visible on the indexed videos. As a result, given an example query, our system is able to retrieve related videos from a large database. We evaluated our system with the copy detection benchmark MUSCLE-VCD-2007. We also ran retrieval experiment on hours of TV broadcast. Hiroaki Kubo, Julien Pilet, Hideo Saito 0001, Shin'ichi Satoh 0001 |
ICPR | 3 |
| 2010 | 3D Human Body Modeling Using Range DataabstractFor the 3D modeling of walking humans the determination of body pose and extraction of body parts, from the sensed 3D range data, are challenging image processing problems. Real body data may have holes because of self-occlusions and grazing angle views. Most of the existing modeling methods rely on direct fitting a 3D model into the data without considering the fact that the parts in an image are indeed the human body parts. In this paper, we present a method for 3D human body modeling using range data that attempts to overcome these problems. In our approach the entire human body is first decomposed into major body parts by a parts-based image segmentation method, and then a kinematics model is fitted to the segmented body parts in an optimized manner. The fitted model is adjusted by the iterative closest point (ICP) algorithm to resolve the gaps in the body data. Experimental results and comparisons demonstrate the effectiveness of our approach. Koichiro Yamauchi 0002, Bir Bhanu, Hideo Saito 0001 |
ICPR | 3 |
| 2010 | Task support system by displaying instructional video onto AR workspaceabstractThis paper presents an instructional support system based on augmented reality (AR). This system helps a user to work intuitively by overlaying visual information in the same way of a navigation system. In usual AR systems, the contents to be overlaid onto real space are created with 3D Computer Graphics. In most cases, such contents are newly created according to applications. However, there are many 2D videos that show how to take apart or build electric appliances and PCs, how to cook, etc. Therefore, our system employs such existing 2D videos as instructional videos. By transforming an instructional video to display, according to the user's view, and by overlaying the video onto the user's view space, the proposed system intuitively provides the user with visual guidance. In order to avoid the problem that the display of the instructional video and the user's view may be visually confused, we add various visual effects to the instructional video, such as transparency and enhancement of contours. By dividing the instructional video into sections according to the operations to be carried out in order to complete a certain task, we ensure that the user can interactively move to the next step in the instructional video after a certain operation is completed. Therefore, the user can carry on with the task at his/her own pace. In the usability test, users evaluated the use of the instructional video in our system through two tasks: a task involving building blocks and an origami task. As a result, we found that a user's visibility improves when the instructional video is transformed to display according to his/her view. Further, for the evaluation of visual effects, we can classify these effects according to the task and obtain the guideline for the use of our system as an instructional support system for performing various other tasks. Michihiko Goto, Yuko Uematsu, Hideo Saito 0001, Shuji Senda, Akihiko Iketani |
ISMAR | 3 |
| 2010 | Foldable augmented mapsabstractThis paper presents folded surface detection and tracking for augmented maps. For the detection, plane detection is iteratively applied to 2D correspondences between an input image and a reference plane because the folded surface is composed of multiple planes. In order to compute the exact folding line from the detected planes, the intersection line of the planes is computed from their positional relationship. After the detection is done, each plane is individually tracked by frame-by-frame descriptor update. For a natural augmentation on the folded surface, we overlay virtual geographic data on each detected plane. The user can interact with the geographic data by finger pointing because the finger tip of the user is also detected during the tracking. As scenario of use, some interactions on the folded surface are introduced. Experimental results show the accuracy and performance of folded surface detection for evaluating the effectiveness of our approach. Sandy Martedi, Hideaki Uchiyama, Guillermo Enriquez, Hideo Saito 0001, Tsutomu Miyashita, Takenori Hara |
ISMAR | 4 |
| 2010 | Foldable augmented mapsabstractThis demonstration presents folded surface detection and tracking for augmented maps. We model the folded surface as multiple planes. To detect a folded surface, plane detection is iteratively applied to 2D correspondences between an input image and a reference plane. In order to compute the exact folding line from the detected planes, the intersection line of the planes is computed from their positional relationship. After the detection is done, each plane is individually tracked by frame-by-frame descriptor update. For a natural augmentation on the folded surface, we overlay virtual geographic data on each detected plane. Sandy Martedi, Hideaki Uchiyama, Guillermo Enriquez, Hideo Saito 0001, Tsutomu Miyashita, Takenori Hara |
ISMAR | 4 |
| 2010 | Bilateral depth-discontinuity filter for novel view synthesisabstractIn this paper, a new filtering technique addresses the disocclusions problem issued from the depth image based rendering (DIBR) technique within 3DTV framework. An inherent problem with DIBR is to fill in the newly exposed areas (holes) caused by the image warping process. In opposition with multiview video (MVV) systems, such as free viewpoint television (FTV), where multiple reference views are used for recovering the disocclusions, we consider in this paper a 3DTV system based on a video-plus-depth sequence which provides only one reference view of the scene. To overcome this issue, disocclusion removal can be achieved by pre-processing the depth video and/or post-processing the warped image through hole-filling techniques. Specifically, we propose in this paper a pre-processing of the depth video based on a bilateral filtering according to the strength of the depth discontinuity. Experimental results are shown to illustrate the efficiency of the proposed method compared to the traditional methods. Ismaël Daribo, Hideo Saito 0001 |
MMSP | 2 |
| 2010 | Generation of see-through baseball movie from multi-camera viewsabstractThis paper presents a method of generating new view point movie for the baseball game. One of the most interesting view point on the baseball game is looking from behind the catcher. If only one camera is placed behind the catcher, however, the view is occluded by the umpire and catcher. In this paper, we propose a method for generating a see-through movie which is captured from behind the catcher by recovering the pitcher's appearance with multiple cameras, so that we can virtually remove the obstacles (catcher and umpire) from the movie. Our method consists of three processes; recovering the pitcher's appearance by Homography, detecting obstacles by Graph Cut, projecting the ball's trajectory. For demonstrating the effectiveness of our method, in the experiment, we generate a see-through movie by applying our method to the multiple camera movies which are taken in the real baseball stadium. In the see-through movie, the pitcher can be appeared through the catcher and umpire. Takanori Hashimoto, Yuko Uematsu, Hideo Saito 0001 |
MMSP | 3 |
| 2010 | Clickable augmented documentsabstractThis paper presents an Augmented Reality (AR) system for physical text documents that enable users to click a document. In the system, we track the relative pose between a camera and a document to overlay some virtual contents on the document continuously. In addition, we compute the trajectory of a fingertip based on skin color detection for clicking interaction. By merging a document tracking and an interaction technique, we have developed a novel tangible document system. As an application, we develop an AR dictionary system that overlays the meaning and explanation of words by clicking on a document. In the experiment part, we present the accuracy of the clicking interaction and the robustness of our document tracking method against the occlusion. Sandy Martedi, Hideaki Uchiyama, Hideo Saito 0001 |
MMSP | 3 |
| 2010 | Depth camera based system for auto-stereoscopic displaysabstractStereoscopic displays are becoming very popular since more and more contents are now available. As an extension, auto-stereoscopic screens allow several users to watch stereoscopic images without wearing any glasses. For the moment, synthetized content are the easiest solutions to provide, in realtime, all the multiple input images required by such kind of technology. However, live videos are a very important issue in some fields like augmented reality applications, but remain difficult to be applied on auto-stereoscopic displays. In this paper, we present a system based on a depth camera and a color camera that are combined to produce the multiple input images in realtime. The result of this approach can be easily used with any kind of auto-stereoscopic screen. François de Sorbier, Yuko Uematsu, Hideo Saito 0001 |
MMSP | 3 |
| 2010 | Influence of wavelet-based depth coding in multiview video systemsabstractMultiview video representation based on depth data, such as multiview video-plus-depth (MVD), is emerging 3D video communication services raising in the meantime the problem of coding and transmitting depth video in addition to classical texture video. Depth video is considered as a key side information in novel view synthesis within multiview video systems, such as three-dimensional television (3 DTV) or free viewpoint television (FTV), wherein the influence of depth compression on the novel synthesized view is still a contentious issue. In this paper, we propose to discuss and investigate the impact of wavelet-based compression of the depth video on the quality of the view synthesis. Experimental results show that significant gains can be obtained by improving depth edge preservation through shorter wavelet-based filtering on depth edges. Ismaël Daribo, Hideo Saito 0001 |
PCS | 2 |
| 2010 | Virtually augmenting hundreds of real pictures: An approach based on learning, retrieval, and trackingabstractTracking is a major issue of virtual and augmented reality applications. Single object tracking on monocular video streams is fairly well understood. However, when it comes to multiple objects, existing methods lack scalability and can recognize only a limited number of objects. Thanks to recent progress in feature matching, state-of-the-art image retrieval techniques can deal with millions of images. However, these methods do not focus on real-time video processing and can not track retrieved objects. In this paper, we present a method that combines the speed and accuracy of tracking with the scalability of image retrieval. At the heart of our approach is a bi-layer clustering process that allows our system to index and retrieve objects based on tracks of features, thereby effectively summarizing the information available on multiple video frames. As a result, our system is able to track in real-time multiple objects, recognized with low delay from a database of more than 300 entries. Julien Pilet, Hideo Saito 0001 |
VR | 2 |
| 2010 | Real-time video-based rendering from uncalibrated cameras using plane-sweep algorithm
Songkran Jarusirisawad, Vincent Nozick, Hideo Saito 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2010 | Guest Editors' Introduction: Special Section on The International Symposium on Mixed and Augmented Reality (ISMAR)abstractThe three papers in this special section are extended versions of papers presented at the International Symposium on Mixed and Augmented Reality (ISMAR). Mark A. Livingston, Ronald T. Azuma, Oliver Bimber, Hideo Saito 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2009 | Dynamics Analysis of Facial Expressions for Person Identification
Hidenori Tanaka, Hideo Saito 0001 |
CAIP | 2 |
| 2009 | Augmenting text document by on-line learning of local arrangement of keypointsabstractWe propose a technique for text document tracking over a large range of viewpoints. Since the popular SIFT or SURF descriptors typically fail on such documents, our method considers instead local arrangement of keypoints. We extends locally likely arrangement hashing (LLAH), which is limited to fronto-parallel images: We handle a large range of viewpoints by learning the behavior of keypoint patterns when the camera viewpoint changes. Our method starts tracking a document from a nearly frontal view. Then, it undergoes motion, and new configurations of keypoints appear. The database is incrementally updated to reflect these new observations, allowing the system to detect the document under the new viewpoint. We demonstrate the performance and robustness of our method by comparing it with the original LLAH. Hideaki Uchiyama, Hideo Saito 0001 |
ISMAR | 2 |
| 2009 | Rotated Image Based Photomosaic Using Combination of Principal Component Hashing
Hideaki Uchiyama, Hideo Saito 0001 |
PSIVT | 2 |
| 2009 | Multiple planes based registration using 3D Projective Space for Augmented Reality
Yuko Uematsu, Hideo Saito 0001 |
Image Vis. Comput. | 2 |
| 2009 | 3DTV view generation using uncalibrated pure rotating and zooming cameras
Songkran Jarusirisawad, Hideo Saito 0001 |
Signal Process. Image Commun. | 2 |
| 2008 | Defocus Blur Correcting Projector-Camera System
Yuji Oyamada, Hideo Saito 0001 |
ACIVS | 2 |
| 2008 | Live video object tracking and segmentation using graph cutsabstractGraph cuts have proven to be powerful tools in image segmentation. Previous graph cut research has proposed methods for cutting across large graphs constructed from multiple layered video frames, resulting in an object being tracked across multiple frames. However, this research focuses on cutting graphs constructed from a prerecorded video sequence. In live video scenarios, frames cannot be layered to construct 3D volumes, since the contents of the subsequent frames are unknown. Instead, new graphs must be created and cut for each frame on demand. Resource limitations make this unfeasible on high-resolution videos. In addition, object tracking requires a method for incorporating the previous frame's object position and shape into the current graph. We propose a method for tracking and segmenting objects in live video that utilizes regional graph cuts and object pixel probability maps. The regionalization of the cuts around the tracked object will increase the speed of the tracker, and the object pixel probability maps will enable more flexible tracking. Zachary A. Garrett, Hideo Saito 0001 |
ICIP | 2 |
| 2008 | Calibration of a structured light system by observing planar object from unknown viewpointsabstractA calibration method for a structured light system by observing a planar object from unknown viewpoints is proposed. A structured light system captures a 3D shape by a camera that observes a light stripe on an object illuminated by a projector. The 3D shape, obtained from the system defined by a pinhole model for the projection of a light stripe, is solved using the equation of a plane model for the projector. The coefficients of each light stripe¿s equation are estimated using the 4×3 image-to-camera transformation matrix that is expressed by camera parameters. Experimental results demonstrate a high degree of accuracy when following the proposed approach. Koichiro Yamauchi 0002, Hideo Saito 0001, Yukio Sato |
ICPR | 2 |
| 2008 | PCA Based 3D Shape Reconstruction of Human Foot Using Multiple Viewpoint Cameras
Edmée Amstutz, Tomoaki Teshima, Makoto Kimura, Masaaki Mochimaru, Hideo Saito 0001 |
ICVS | 5 |
| 2007 | Real-Time Free Viewpoint from Multiple Moving Cameras
Vincent Nozick, Hideo Saito 0001 |
ACIVS | 2 |
| 2007 | Focal Pre-Correction of Projected Image for Deblurring Screen ImageabstractWe propose a method for reducing out-of-focus blur caused by projector projection. In this method, we estimate the Point-Spread-Function (PSF) of the out-of-focus blur in the image projected onto the screen by comparing the screen image captured by a camera with the original image projected by the projector. According to the estimated PSF, the projected image is pre-corrected, so that the screen image can be deblurred. Experimental results show that our method can reduce out-of-focus projection blur. Yuji Oyamada, Hideo Saito 0001 |
CVPR | 2 |
| 2007 | Vision-Based Guitarist Fingering Tracking Using a Bayesian Classifier and Particle Filters
Chutisant Kerdvibulvech, Hideo Saito 0001 |
PSIVT | 2 |
| 2007 | Online Multiple View Computation for Autostereoscopic Display
Vincent Nozick, Hideo Saito 0001 |
PSIVT | 2 |
| 2007 | 3D Reconstruction of a Human Body from Multiple Viewpoints
Koichiro Yamauchi 0002, Hideto Kameshima, Hideo Saito 0001, Yukio Sato |
PSIVT | 3 |
| 2007 | Virtual Viewpoint Replay for a Soccer Match by View Interpolation From Multiple CamerasabstractThis paper presents a novel method for virtual view synthesis that allows viewers to virtually fly through real soccer scenes, which are captured by multiple cameras in a stadium. The proposed method generates images of arbitrary viewpoints by view interpolation of real camera images near the chosen viewpoints. In this method, cameras do not need to be strongly calibrated since projective geometry between cameras is employed for the interpolation. For avoiding the complex and unreliable process of 3-D recovery, object scenes are segmented into several regions according to the geometric property of the scene. Dense correspondence between real views, which is necessary for intermediate view generation, is automatically obtained by applying projective geometry to each region. By superimposing intermediate images for all regions, virtual views for the entire soccer scene are generated. The efforts for camera calibration are reduced and correspondence matching requires no manual operation; hence, the proposed method can be easily applied to dynamic events in a large space. An application for fly-through observations of soccer match replays is introduced along with the algorithm of view synthesis and experimental results. This is a new approach for providing arbitrary views of an entire dynamic event. Naho Inamoto, Hideo Saito 0001 |
IEEE Trans. Multim. | 2 |
| 2006 | Texture overlay for virtual clothing based on PCA of silhouettesabstractIn this paper, we propose a method for overlaying an arbitrary texture image onto a surface of a plain T-shirt worn by a user. For overlaying arbitrary textures onto the surface of the T-shirt, we need to know the deformation of the surface. For estimating the deformation of the surface from the input images, we use a two-phase process: learning and searching. In the learning phase, the system learns the relationship between the deformation of the surface and the silhouette of the T-shirt region in the image. A database of a number of training images in which a person wearing a T-shirt with markers moves through a variety of positions is used for this learning. Using the database, the system can learn the relationship between the shape of the silhouette and the surface deformation that is provided by the 2D positions of the markers on the surface of the T-shirt. In the searching phase, the silhouette of the user's T-shirt is extracted from the input image, and then, a search for a similar silhouette in the database is conducted in the subspace of the silhouette, which is computed using a PCA of the database. By using the proposed method for estimating the deformation of the surface of the T-shirt, we perform experiments for overlaying virtual clothing. Jun Ehara, Hideo Saito 0001 |
ISMAR | 2 |
| 2006 | Support system for guitar playing using augmented reality displayabstractLearning to play the guitar is difficult. We proposed a system that assists people learning to play the guitar using augmented reality. This system shows a learner how to correctly hold the strings by overlaying a virtual hand model and lines onto a real guitar. The player learning to play the guitar can easily understand the required position by overlapping their hand on a visual guide. An important issue for this system to address is the accurate registration between the visual guide and the guitar, therefore we need to track the pose and the position of the guitar. We also proposed a method to track the guitar with a visual marker and natural features of the guitar. Since we used marker information and edge information as natural features, the system could continually track the guitar. Accordingly, our system can constantly display visual guides at the required position to enable a player to learn to play the guitar in a natural manner. Yoichi Motokawa, Hideo Saito 0001 |
ISMAR | 2 |
| 2006 | Position Estimation of Solid Balls from Handy Camera for Pool Supporting System
Hideaki Uchiyama, Hideo Saito 0001 |
PSIVT | 2 |
| 2005 | Free-viewpoint image synthesis from multiple-view images taken with uncalibrated moving camerasabstractIn this paper, we propose methods for free-viewpoint image synthesis from multiple-view images taken with uncalibrated cameras. In our method, two viewpoints are selected as basis images for defining a projective 3D coordinate of the object scene in which the scene structure is recovered from the input multiple-view images. The 3D coordinate is defined by the image coordinates of the selected two basis images according to the epipolar geometry between those two images, which is represented by a fundamental matrix. The multiple-view images are also related to the projective 3D coordinate by their epipolar geometry to the basis images. Based on such a framework, we do not need to strongly calibrate the cameras, so we can recover 3D structure of the scene without effort for the camera calibration. In addition to that, we can also render free-viewpoint images from hand-held moving cameras. We will demonstrate the effectiveness of the proposed method by showing free-viewpoint images from multi-view images taken with hand-held moving cameras. Yosuke Ito, Hideo Saito 0001 |
ICIP (3) | 2 |
| 2005 | Image-based augmentation of virtual object for handy camera video sequence using arbitrary multiple planesabstractWe propose a novel vision-based registration method for augmented reality with merging of arbitrary multiple planes. In our approach, we require neither artificial markers nor sensors, and estimate the camera rotation and translation by an uncalibrated image sequence in which arbitrary multiple planes in the real world exist. Since the geometrical relationship of those planes is unknown, for merging of them, we assign a 3D coordinate system for each plane independently and construct projective 3D space defined by projective geometry of two reference images. By merging with the projective space, we can use arbitrary multi-planes, and achieve high-accurate registration for every position in the input images. Yuko Uematsu, Hideo Saito 0001 |
ICIP (1) | 2 |
| 2005 | Free viewpoint video synthesis and presentation from multiple sporting videosabstractThis paper introduces two kinds of free viewpoint observation systems for sporting events captured with uncalibrated multiple cameras in a stadium. In the first system (viewpoint on demand system), a user can watch the realistic sporting scenes with the original stadium. In the second system (mixed reality presentation system), a user can virtually watch the scenes overlaid on a desktop stadium model via a video see-through head mounted display (HMD). In the both systems, the user can observe sporting events from his/her favorite viewpoints, where the virtual view images are synthesized and presented by performing view interpolation. As outdoor scenes often have variations in lighting, we develop the systems to handle the changes of lighting condition. If the captured scene contains shadows, we synthesize the virtual view image of the shadows of the background and the foreground independently from real camera images using projective geometry between cameras. The shadows of the foreground objects are then overlaid on the synthesized background of the original stadium or the stadium model in front of the user respectively. The results indicate that the appearance of shadows can produce a realistic mixed reality presentation. Naho Inamoto, Hideo Saito 0001 |
ICME | 2 |
| 2005 | Video Synthesis at Tennis Player Viewpoint from Multiple View VideosabstractIn this paper, we propose a new method for synthesizing player-view from multiple videos in tennis. Our method is divided into two key techniques: virtual-view synthesis and player's view-point estimation. In the former technique, the existing method view interpolation, which is a technique of synthesizing images in an intermediate viewpoint from the real images in two view-points, restricts a viewpoint position between two viewpoints. To avoid this restriction, we propose a new method of the ability to set up a viewpoint position freely. In the latter, viewpoint position is computed from the center of gravity of a player region in captured videos using epipolar geometry. By applying the computed player's viewpoint to the former method, we can synthesize player-view image. Experimental results demonstrate that the proposed method can successfully provide tennis player-view video in tennis. Kenji Kimura, Hideo Saito 0001 |
VR | 2 |
| 2004 | Sports scene analysis and visualization from multiple-view videoabstractWe introduce methods for sports scene analysis and visualization from multiple videos captured with multiple cameras. As the scene analysis, we present a method for tracking multiple soccer players. Tracking is done by integrating the tracking data from all cameras, using the geometrical relationship between cameras called homography. Integrating information from all cameras enables stable tracking on the scene, where the tracking by a single camera often fails in the case of occlusion. We also present a method for free-viewpoint visualization of a soccer game as a technique for visualizing the soccer scene. The free-viewpoint image is synthesized by view interpolation between actual cameras near the virtual viewpoint at each frame. Such free-viewpoint video can also be presented in a system of augmented reality (AR) for immersive visualization of soccer matches Hideo Saito 0001, Naho Inamoto, Sachiko Iwase |
ICME | 1 |
| 2004 | Texture Overlay onto Deformable Surface Using HMD
Mototsugu Emori, Hideo Saito 0001 |
VR | 2 |
| 2004 | Arbitrary viewpoint video synthesis from multiple uncalibrated camerasabstractWe propose a method for arbitrary view synthesis from uncalibrated multiple camera system, targeting large spaces such as soccer stadiums. In Projective Grid Space (PGS), which is a three-dimensional space defined by epipolar geometry between two basis cameras in the camera system, we reconstruct three-dimensional shape models from silhouette images. Using the three-dimensional shape models reconstructed in the PGS, we obtain a dense map of the point correspondence between reference images. The obtained correspondence can synthesize the image of arbitrary view between the reference images. We also propose a method for merging the synthesized images with the virtual background scene in the PGS. We apply the proposed methods to image sequences taken by a multiple camera system, which installed in a large concert hall. The synthesized image sequences of virtual camera have enough quality to demonstrate effectiveness of the proposed method. Satoshi Yaguchi, Hideo Saito 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2003 | 3D Shape and Pose Estimation of Deformable Tapes from Multiple ViewsabstractIn this paper, we propose a method to estimate 3D shape of deformable plastic tapes from multiple camera images. In this method, the tape is modeled as serial connection of multiple rectangular plates, where the size of each plate is previously known and node angles of between plates represent the shape of the object. The node angles of the object are estimated by 2D silhouette shapes taken in the multiple images. The estimation is performed by minimizing the difference of the silhouette shapes between the input images and synthesized images of the model shape. For demonstrating the proposed method, 3D shape of a tape is estimated with two camera images. The accuracy of the estimation is sufficient for making the assembling robot in our plant to handle the tape. Computation time is also sufficiently short for applying the proposed algorithm in the assembling plant. 1 Hitoshi Kubota, Masakazu Ono, Masami Takeshi, Hideo Saito 0001 |
BMVC | 4 |
| 2003 | Factorization Method Using Interpolated Feature Tracking via Projective GeometryabstractThe factorization method by Tomasi and Kanade simultaneously recovers camera motion and object shape from an image sequence. This method is robust because the solution is linear by assuming the orthographic camera model. However, the only feature points that are tracked throughout the image sequence can be reconstructed, it is difficult to recover whole object shape by the factorization method. In this paper, we propose a new method to interpolate feature tracking so that even the loci of unseen feature points can be used as inputs of the factorization for object shape reconstruction. In this method, we employ projective reconstruction to interpolate untracked feature points. All loci of all detected feature points throuout the input image sequence provide correct reconstructed shape of the object via the factorization. The results of reconstruction are evaluated by the experiment using synthetic images and real images. 1 Introdution Hideo Saito 0001, Shigeharu Kamijima |
BMVC | 1 |
| 2003 | Displaying Digital Documents on Real Paper Surface with Arbitrary ShapeabstractIn this paper, we propose a system that displays digital components on real paper surface with arbitrary shape, so that the viewer can feel as if the digital document images are printed on the real paper surface. Such displaying of the digital documents is realized by rendering the document images on the arbitrary shaped surface via a projector. We apply a holography between a source image plane and a projector image plane to render the images on the surface. For adapting the arbitrary shape of the surface, we divide the shaped surface into many small rectangular regions, and generate warp images of each region by calculating this holography of the plane of each divided region. By protecting the warp image on the real surface, the image can be observed as if the image is printed on the surface. Since the system always compute the holography of each divided region, the image can be aligned onto the surface even the surface moves. Shinichiro Hirooka, Hideo Saito 0001 |
ISMAR | 2 |
| 2003 | Immersive Observation of Virtualized Soccer Match at Real Stadium ModelabstractThis paper presents a novel observation system for immersive soccer match taken by multiple video cameras at a real stadium. The user sees the soccer field model in front of his/her eyes from the viewpoint through head-mounted display, while the images of players and a soccer ball are also rendered onto the display. For geometric registration between the soccer field model in the real world and the dynamic soccer scene in the rendered images, the viewpoint position of the user is computed by using only natural feature lines in the HMD camera image. Since it is difficult to strongly calibrate the HMD camera and the multiple cameras that capture the real soccer scene, we employ projective geometry for the registration and the rendering. For demonstrating the efficacy of the proposed system, video images of soccer matches taken at real stadium are rendered onto the HMD camera images of a tabletop stadium model. This is a completely new challenge to apply augmented reality to a dynamic event in a large-space. Naho Inamoto, Hideo Saito 0001 |
ISMAR | 2 |
| 2003 | Fly-through viewpoint video system for multi-view soccer movie using viewpoint interpolation
Naho Inamoto, Hideo Saito 0001 |
VCIP | 2 |
| 2003 | Tracking soccer players based on homography among multiple views
Sachiko Iwase, Hideo Saito 0001 |
VCIP | 2 |
| 2003 | Light field rendering with omni-directional camera
Hiroshi Todoroki, Hideo Saito 0001 |
VCIP | 2 |
| 2003 | Recovery of shape and surface reflectance of specular object from relative rotation of light source
Hideo Saito 0001, Kazuko Omata, Shinji Ozawa |
Image Vis. Comput. | 1 |
| 2003 | Appearance-based virtual view generation from multicamera videos captured in the 3-D roomabstractWe present an appearance-based virtual view generation method that allows viewers to fly through a real dynamic scene. The scene is captured by multiple synchronized cameras. Arbitrary views are generated by interpolating two original camera-views near the given viewpoint. The quality of the generated synthetic view is determined by the precision, consistency and density of correspondences between the two images. All or most of previous work that uses interpolation extracts the correspondences from these two images. However, not only is it difficult to do so reliably (the task requires a good stereo algorithm), but also the two images alone sometimes do not have enough information, due to problems such as occlusion. Instead, we take advantage of the fact that we have many views, from which we can extract much more reliable and comprehensive 3D geometry of the scene as a 3D model. Dense and precise correspondences between the two images, to be used for interpolation, are obtained using this constructed 3D model. Hideo Saito 0001, Shigeyuki Baba, Takeo Kanade |
IEEE Trans. Multim. | 1 |
| 2003 | Special issue on 3-D image analysis and modeling
Hongbin Zha, Hideo Saito 0001, Vittorio Murino, Andrea Fusiello |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2002 | Virtual 3-D interface system via hand motion recognition from two camerasabstractWe propose a new human interface system, the virtual 3D interface system, which allows a user to work in a virtual three-dimensional (3D) space on the computer by way of hand motions and hand poses. This system has two cameras: one takes the top-view of the desktop and the other takes the side-view. With the top-view image, the system recognizes the bending of the five fingers and interprets the command. At the same time, the system extracts the fingertip of the forefinger in both the top-view and the side-view images, and then estimates the 3D position of the fingertip. These procedures are very simple so that we can develop some real-time 3D application systems. Kiyofumi Abe, Hideo Saito 0001, Shinji Ozawa |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2001 | Interpolation of three views based on epipolar geometry
Makoto Kimura, Hideo Saito 0001 |
VCIP | 2 |
| 2000 | 3D Reconstruction of Book Surface Taken from Image Sequence with Handy CameraabstractWe propose a method for reconstruction of the 3D shape of a book surface by integrating shape that is recovered from an image sequence taken with a moving camera. The image sequence is divided into a number of blocks. The factorization method is applied to each block for obtaining a part of the 3D shape of the object. The 3D shapes reconstructed from the block of the sequence are merged into one integrated 3D model. For the merging, position and pose of each 3D shape are adjusted based on the surface normal of the overlapped area. The merged 3D model of all regions of the book surface is shown for demonstrating the effectiveness of the proposed method. Hideo Saito 0001, Shinji Ozawa |
ICPR | 2 |
| 2000 | 3D Shape Recovery of Non-Convex Object from RotationabstractWe propose a method for estimation of the 3D shape of a non-convex object from an image sequence taken with the object rotating under a fixed light source. First, all the territory candidates of the object are prepared in voxel space. The empty voxels are selected by using cross-section estimation using a silhouette, that is the conventional method. However this method cannot reconstruct the non-convex object, because change isn't shown in the silhouette by the dents. Therefore, we remove part of the dents by using the cross-section estimation using occlusion and shading. For the cross-section estimation using occlusion, we apply a technique similar to the shape from silhouette. That is to say, the ray connects the point of view and occlusion to the tangent toward the front object of occlusion and removes the territory surrounded by the corresponding two tangents of the two neighboring views. For the cross-section estimation using shading, we decide the most suitable position of the surface point in the searched ray, which connects a point of view and the surface and on which the surface point is searched. To decide a suitable position, we examine a change in brightness of backward rays, which connect a voxel and each point of view. At each voxel on the searched ray, the most suitable position is where the degree of agreement with the bi-directional reflectance model is the highest. For demonstrating the effectiveness of the proposed method, we show the reconstructed image of non-convex objects, such as balls, or a doll and chair, which are successfully recovered. Taichi Sato, Hideo Saito 0001, Shinji Ozawa |
ICPR | 2 |
| 2000 | 3-D drawing system via hand motion recognition from two camerasabstractThe authors propose a novel human interface system, 3-D Drawing System, which allows a user to draw on virtual 3D space in a computer with a user's hand motions and hand poses. The system has two cameras; one camera takes a top-view of the desktop and another takes a side-view. The input image of the top camera is transformed into the distance-transformed image, which is used for deciding Palm Square that gives the base size, position, and direction of hand. The Palm Square determines Finger Windows which are used for detecting each finger's expansion that is decided by the number of pixels in the Finger Window. The user's command is input to the system by the finger's expansion. At the same time, we extract the fingertip and estimate the 3D position by comparison with the image obtained with the side camera. The procedures are very simple so we can develop the real time 3D drawing system, which allows the user to draw figures in 3D space, and handle 3D virtual objects. Kiyofumi Abe, Hideo Saito 0001, Shinji Ozawa |
SMC | 2 |
| 2000 | Generation of 3D model with super resolved texture from image sequenceabstractProposes a method for the reconstruction of a 3D shape with super-resolved texture of a book surface by integrating the shape that is recovered from an image sequence taken with a moving camera. The image sequence is divided into a number of blocks. A factorization method is applied to each block to obtain part of the 3D shape of the object. The 3D shapes reconstructed from the blocks of the sequence are merged into a single integrated 3D model. During this time, we generate a super-resolved texture using multiple images. For the merging, the position and pose of each 3D shape is adjusted based on the surface normal of the overlapped area. The merged 3D model of all regions of the book surface is shown to demonstrate the effectiveness of the proposed method. Furthermore, we generate a super-resolved texture that is finer than the original image using multiple images. Hideo Saito 0001, Shinji Ozawa |
SMC | 2 |
| 2000 | Super resolving texture mapping from multiple view images for 3D model constructionabstractThere are wide varieties of applications of 3D modeling from real images, which can generate virtual view images by rendering the input images on to a 3D shape model. Because the virtual viewpoint is not generally at the same point, the synthesized image in a virtual view does not have sufficient resolution. To improve this resolution problem, we propose a method for improving the rendered texture resolution. For each vertex of the 3D shape model, the position of the vertex is adjusted to improve the texture of the neighboring triangle patches. For the improvement, the position of the vertex is adjusted by comparing the observation (input image) with the projected image of the model texture. By repeating the visit to every vertex, not only can a super-resolved texture be obtained but also a more accurate 3D patch model. By performing experiments using synthesized multiple-view images of 51 calibrated cameras around the object scene and a real VTR image sequence, it is demonstrated that the proposed method can successfully provide a super-resolved texture image with an improved 3D shape model. Hideo Saito 0001 |
SMC | 1 |
| 2000 | The evaluating system of human skin surface condition by image processingabstractWe propose an automatic evaluation system of human skin surface condition based on subjective evaluation provided by cosmeticians. In the proposed system, image features extracted on the skin image and subjective evaluation by cosmeticians are flexibly connected by using a backpropagation neural network, so that it can automatically estimate human skin surface condition based on subjective evaluation from various skin images. We show some experimental results. Using the trained neural network, human skin surface condition based on subjective evaluation is estimated for unlearned skin images. Then subjective evaluation by this system was compared with that by cosmeticians. Since the proposed system can successfully estimate human skin surface condition like cosmeticians, the effect of the system is demonstrated. Yoshinao Takemae, Hideo Saito 0001, Shinji Ozawa |
SMC | 2 |
| 1999 | Shape Reconstruction in Projective Grid Space from Large Number of ImagesabstractThis paper proposes a new scheme for multi-image projective reconstruction based on a projective grid space. The projective grid space is defined by two basis views and the fundamental matrix relating these views. Given fundamental matrices relating other views to each of the two basis views, this projective grid space can be related to any view. In the projective grid space as a general space that is related to all images, a projective shape can be reconstructed from all the images of weakly calibrated cameras. The projective reconstruction is one way to reduce the effort of the calibration because it does not need Euclid metric information, but rather only correspondences of several points between the images. For demonstrating the effectiveness of the proposed projective grid definition, we modify the voxel coloring algorithm for the projective voxel scheme. The quality of the virtual view images re-synthesized from the projective shape demonstrates the effectiveness of our proposed scheme for projective reconstruction from a large number of images. Hideo Saito 0001, Takeo Kanade |
CVPR | 1 |
| 1999 | 3D Voxel Construction Based on Epipolar Geometry
Makoto Kimura, Hideo Saito 0001, Takeo Kanade |
ICIP (3) | 2 |
| 1999 | Face Pose Estimating System Based on Elgen Space AnalysisabstractIn this paper, we propose a new system for estimating face pose from a facial image. In this system, an input facial image is compared with a database of images of various face pose, then the matched image provides the face pose. The database of images includes not only various face poses but also various illumination conditions, so that the face pose estimation system can be used under various illumination conditions. For collecting such various facial images, they are generated by computer, rather than using real images. The eigenspace method is used for searching a matched image with an input facial image. Since various illumination images are collected in the database of facial images, the extracted principle eigenvectors mostly depend on the face pose. By performing the matching process in the eigenspace, a matched image with the input facial image can be found. The pose of the matched image is closest to the input face. The matching process is also fast because it is performed in small dimensional space spanned by selected eigenvectors only. The proposed pose estimation system can continuously track the face pose of different persons under different light conditions. Hideo Saito 0001, Akihiro Watanabe, Shinji Ozawa |
ICIP (1) | 1 |
| 1998 | Shape Modeling from Multiple View Images Using GAs
Satoshi Kirihara, Hideo Saito 0001 |
ACCV (2) | 2 |
| 1997 | Analytical Solution of Shape from Shading Problem
Seong Ik Cho, Hideo Saito 0001, Shinji Ozawa |
BMVC | 2 |
| 1997 | A Divide-and-conquer Strategy in Shape from Shading ProblemabstractA divide-and-conquer strategy in shape from shading problem under fully perspective conditions is proposed for the information recovery of book surfaces. The whole recovery process is composed of three sequential steps: preprocessing, apparent shape recovery, and ortho-image generation. Pure shade images are extracted in the preprocessing step by introducing a phenomenological model of interreflection and by removing pigment parts from observed images. A recurrence relation is derived from the definition of mean slope in the discrete image. Theoretically, it becomes possible to recover unique shape without iteration using the derived equations in the case of a Lambertian cylinder. However, a feedback shape recovery process is implemented as a practical algorithm in order to overcome self-shadows. Results of simulations and real experiments show the properness and acceptability of the proposed strategy and implemented algorithms. Seong Ik Cho, Hideo Saito 0001, Shinji Ozawa |
CVPR | 2 |
| 1997 | Evaluation of the relationship between emotional concepts and emotional parameters on speechabstractWe propose a linear model of the relationship between the physical changes in speech and perceived emotional concepts. We make use of orthogonal bases in spite of the emotional words and the physical parameters themselves in order to avoid dependence on the method of selecting words and parameters. Furthermore we regard the emotions that listeners perceive from speech as the standard or emotional concepts because the emotions that speakers intended rely on their personality and temporary psychological state. Evaluation for relative information indicates that the proposed linear model is representable for the relationship between the physical quantities and psychological quantities in speech. Tsuyoshi Moriyama, Hideo Saito 0001, Shinji Ozawa |
ICASSP | 2 |
| 1996 | Object modeling from multiple images using genetic algorithmsabstractThis paper describes an application of genetic algorithms (GAs) to modeling of multiple objects from CCD images. Shape modeling is a very important issue for shape recognition for robot vision, representing 3-D shapes in the virtual world, and so on. In this paper, we propose a method for object modeling from multiple view images using genetic algorithms (GAs). In this method, similarity between the model and the image at each view angle is evaluated. The model having the maximum evaluation is found by GAs. In the proposed method, a sharing scheme is used for finding multiple solutions efficiently. Some results of object modeling experiments from synthetic and real multiple view images demonstrate that the proposed method can robustly generate models by using GAs. Hideo Saito 0001, Masayuki Mori |
ICPR | 1 |
| 1995 | Application of genetic algorithms to stereo matching of images
Hideo Saito 0001, Masayuki Mori |
Pattern Recognit. Lett. | 1 |
| 1994 | Halftoning technique using genetic algorithmabstractA new method for the halftoning technique using genetic algorithms (GA) is proposed. For representing gray-tone images on bilevel displays or printers, the gray-tone images need to be transformed into binary images by using halftoning techniques. Visually pleasing halftone images must have both a high spatial resolution and a high gray level resolution. In this method, a halftoning technique is regarded as an optimizing problem that is to search for a visually pleasing distribution of white and black pixels. In the study, spatial and gray level resolutions of halftone images are evaluated, and then the evaluated value is optimized by using GA. For demonstrating an efficacy of the proposed two halftoning techniques, the results obtained by a computer simulation are shown.> Naoki Kobayashi 0008, Hideo Saito 0001 |
ICASSP (5) | 2 |
| 1994 | Estimation of 3-D parametric models from shading image using genetic algorithmsabstractIn this paper, a method for estimating parameters of a 3-D shape from a 2-D shading image using a genetic algorithms (GAs) is proposed. The shape of the object is represented by a superquadrics model, and then the model parameters are coded for application to GAs. The coded string is evaluated according to the similarity of the shading image calculated from the 3-D model shape represented by the parameters to the given 2-D shading image. By applying the GAs to the optimization of the evaluation value, the string having the minimum difference can be found. The parameters are estimated from some shading images of various 3-D shapes by using the proposed method, and the results are presented. Hideo Saito 0001, Nobuhiro Tsunashima |
ICPR (1) | 1 |