Itaru Kitahara

dblp:85/2726 · DBLP profile ↗
← Back
42ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0002-5186-789XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 17 · 4 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2025 Dual-S3D: Hierarchical Dual-Path Selective SSM-CNN for High-Fidelity Implicit Reconstruction
Luoxi Zhang, Pragyan Shrestha, Chun Xie, Itaru Kitahara
ICCV5
2025 MGP-KAD: Multimodal Geometric Priors and Kolmogorov-Arnold Decoder for Single-View 3d Reconstruction in Complex Scenes
abstract
Single-view 3D reconstruction in complex real-world scenes is challenging due to noise, object diversity, and limited dataset availability. To address these challenges, we propose MGP-KAD, a novel multimodal feature fusion framework that integrates RGB and geometric prior to enhance reconstruction accuracy. The geometric prior is generated by sampling and clustering ground-truth object data, producing class-level features that dynamically adjust during training to improve geometric understanding. Additionally, we introduce a hybrid decoder based on Kolmogorov-Arnold Networks (KAN) to overcome the limitations of traditional linear decoders in processing complex multimodal inputs. Extensive experiments on the Pix3D dataset demonstrate that MGP-KAD achieves state-of-the-art (SOTA) performance, significantly improving geometric integrity, smoothness, and detail preservation. Our work provides a robust and effective solution for advancing single-view 3D reconstruction in complex scenes.
Luoxi Zhang, Chun Xie, Itaru Kitahara
ICIP3
2025 SV-DRR: High-Fidelity Novel View X-Ray Synthesis Using Diffusion Model
Chun Xie, Yuichi Yoshii, Itaru Kitahara
MICCAI (4)3
2024 RayEmb: Arbitrary Landmark Detection in X-Ray Images Using Ray Embedding Subspace
Pragyan Shrestha, Chun Xie, Yuichi Yoshii, Itaru Kitahara
ACCV (2)4
2023 A Method for Completing Missing 3D Point Cloud Reconstructed from Aerial Multi-View Images Using Self-Attention Mechanism
abstract
This paper proposes a method to complete the missing 3D point cloud reconstructed from aerial multi-view images by using a deep learning method with self-attention. The advancement of drone technology has made it easier to acquire aerial multi-view images. While it is possible to generate 3D point clouds of the terrain by applying 3D photogrammetric techniques to these images, when capturing multi-view aerial images with a drone, high-altitude vertical shooting is often necessary for privacy protection. For example, some portions of the generated 3D point clouds are lost due to shadowed areas caused by roofs and eaves. To address this issue, this research proposes a method to complete the missing 3D point cloud by using a deep learning. In order to obtain accurate and sufficient amount of training data, 3D CG building models are used for generating sets of missing 3D point cloud data and their corresponding Ground Truth. In the experiment, we applied our method to a 3D point cloud generated from actual captured aerial multi-view images and confirmed that the point cloud with a reasonable shape for the missing parts are successfully completed.
Takenobu Kiyama, Chun Xie, Hidehiko Shishido, Hisatoshi Toriya, Itaru Kitahara
IGARSS5
2023 X-Ray to CT Rigid Registration Using Scene Coordinate Regression
Pragyan Shrestha, Chun Xie, Hidehiko Shishido, Yuichi Yoshii, Itaru Kitahara
MICCAI (10)5
2022 Neural Density-Distance Fields
Itsuki Ueda, Yoshihiro Fukuhara, Hirokatsu Kataoka, Hiroaki Aizawa, Hidehiko Shishido, Itaru Kitahara
ECCV (32)6
2022 Size Does Matter: An Experimental Study of Anxiety in Virtual Reality
abstract
The emotional response of users induced by VR scenarios has become a topic of interest, however, whether changing the size of objects in VR scenes induces different levels of anxiety remains a question to be studied. In this study, we conducted an experiment to initially reveal how the size of a large object in a VR environment affects changes in participants’ (N = 38) anxiety level and heart rate. To holistically quantify the size of large objects in the VR visual field, we used the omnidirectional field of view occupancy (OFVO) criterion for the first time to represent the dimension of the object in the participant’s entire field of view. The results showed that the participants’ heartbeat and anxiety while viewing the large objects were positively and significantly correlated to OFVO. These study reveals that the increase of object size in VR environments is accompanied by a higher degree of user’s anxiety.
Itaru Kitahara, Shinichi Koyama, Qiaoge Li
VRST2
2021 A Method to Correct Perspective Distortion of Ground Area without Camera Parameters
abstract
This study proposes a method for geometric transformation of a ground area in mobile camera images without cameras' internal parameters for registering them to Geographic Information System (GIS) database. If mobile camera images can be registered in GIS, it will be possible to share the latest geographic information by combining it with crowdsourcing in times of disaster. To achieve it, it is necessary to map the mobile camera images to the GIS image database. However, it is difficult to estimate the direct correspondence because most of the mobile camera images are landscape images, and the perspective of the ground area differs significantly. In contrast, GIS image data contain directly downward images. If the cameras' internal parameters and the posture at the shooting are known, it is possible to convert the image to top-view. However, considering the premise of crowdsourcing image collection, these parameters are not necessarily known. Therefore, we develop a method to correct the perspective distortion of the ground area in the image and convert it to a top-eye view even when the camera parameters are unknown.
Hisatoshi Toriya, Ashraf M. Dewan, Itaru Kitahara
IEEE BigData3
2020 Adaptive Image Scaling for Corresponding Points Matching between Images with Differing Spatial Resolutions
abstract
In this study, an image scaling method to improve the accuracy of the image registration between images using different imaging devices is proposed. It is known that conventional keypoint detection, description, and matching methods do not work well between images with different spatial resolutions, such as those captured by drones and satellites. Thus, we propose a method to improve the geometric accuracy of image registration through an adaptive combination of super-resolution and low- resolution images and downscaling to high-resolution images based on the assumption that artificial structures are perceived as relatively simple shapes in top-view images. If the superresolution factor is too high, artifacts are generated, and the accuracy of the corresponding matching points will be decreased. Thus, estimating the highest super-resolution factor while avoiding the artifacts is necessary. Using the super-resolution factor, super-resolution processing of satellite images and downscaling to drone images are simultaneously performed. This is followed by corresponding points matching to achieve high estimation accuracy in the image registration process. Through quantitative evaluation experiments performed using pairs of images with 12 times difference in spatial resolutions, we demonstrated that high-accuracy image registration is possible by applying super-resolution processing at a factor of 4 to 6.
Hisatoshi Toriya, Ashraf M. Dewan, Itaru Kitahara
IEEE BigData3
2019 Super Long Interval Time-Lapse Image Generation for Proactive Preservation of Cultural Heritage Using Crowdsourcing
abstract
To establish advanced analytical methods for preserving cultural heritage, this research proposes a method to generate a time-lapse image with a super-long temporal interval. The key issue is to realize an image collection method using crowdsourcing and a method to improve the matching accuracy between images of cultural heritage buildings captured 50 to 100 years ago and current images. As degradation and damage to the appearance of cultural heritage buildings occurs due to ageing, rebuilding, and renovation, image features of the timed images are changed. This decreases the accuracy of the matching process that uses the appearance of patch-region. In addition, we need to give more consideration to incorrect feature correspondence that is prominent in buildings with considerable symmetry. We aim to solve these difficulties by applying an Autoencoder and a guided matching method. Our method involves utilizing the function of crowdsourcing, which can easily obtain the current image captured at the same position and orientation as the past image. We propose this method to address the inability to obtain the correspondence points between two images when observation times are significantly different.
Hidehiko Shishido, Hansung Kim 0001, Itaru Kitahara
IEEE BigData3
2019 SAR2OPT: Image Alignment Between Multi-Modal Images Using Generative Adversarial Networks
abstract
This work proposes an image-alignment method for multi-modal images (e.g., synthetic aperture radar (SAR) and optical satellite images) using an image-feature-based keypoint-matching algorithm. In applying the matching algorithm to multi-modal images, common features need to be obtained at the corresponding positions. However, the appearances of features among images are different. We solve this issue by translating the appearance of one modal image to the other using generative adversarial networks (GANs). In this work, we attempt to generate optical images from SAR images as a way to extract common features. Through an experiment, we confirm that the proposed method can estimate accurate correspondences between SAR and optical images.
Hisatoshi Toriya, Ashraf M. Dewan, Itaru Kitahara
IGARSS3
2019 Visual Exploratory Activity under Microgravity Conditions in VR: An Exploratory Study during a Parabolic Flight
abstract
This work explores the human visual exploratory activity (VEA) in a microgravity environment compared to one-G. Parabolic flights are the only way to experience microgravity without astronaut training, and the duration of each microgravity segment is less than 20 seconds. Under such special conditions, the test subject visually searches a virtual representation of the International Space Station located in his Field of Regard (FOR). The task was repeated in two different postural positions. Interestingly, the test subject reported a significant reduction of microgravity-related motion sickness while experiencing the VR simulation, in comparison to his previous parabolic flights without VR.
César Daniel Rojas Ferrer, Hidehiko Shishido, Itaru Kitahara, Yoshinari Kameda
VR3
2019 Smooth switching method for asynchronous multiple viewpoint videos using frame interpolation
Hidehiko Shishido, Aoi Harazaki, Yoshinari Kameda, Itaru Kitahara
J. Vis. Commun. Image Represent.4
2018 A Method to Collect Multi-view Images of High Importance Using Disaster Map and Crowdsourcing
abstract
In recent years, research efforts have enabled using the internet in disaster areas, and information technology (IT) is expected to improve our comprehension and evaluation of disasters. In disaster areas, crowdsourcing can be employed to secure human resources and controlled-task distribution. In crowdsourcing, many workers with publicly defined tasks (also called microtasks) are used. A simplified and more efficient microtask framework is required for disaster areas due to the lack of human resources and facilities. In this paper, we describe a method that incorporates information collection from a disaster area into microtasks with machine processing support.
Koyo Kobayashi, Hidehiko Shishido, Yoshinari Kameda, Itaru Kitahara
IEEE BigData4
2018 Time-Lapse Image Generation using Image-Based Modeling by Crowdsourcing
abstract
In recent years, the pillars of the World Heritage Angkor Thom Bayon temple have become a problem of deterioration due to moss breeding. We aim to generate an image to support observation of moss breeding on a pillar. Even under environment that prevent image processing, we can achieve accurate overlay processing by combining corresponding points between images and 3D shapes. In order to generate the timelapse image of the observation target, many accurate images of different capturing timings are necessary. We are going to use a lot of images collected by crowdsourcing for time lapse images. In this research, we use two crowdsourcing models with the "capturing image of the target region" and the "classification of the captured images" as the micro task. Therefore, image acquisition using crowdsourcing and generation of time lapse image are looped. Time lapse image will be more accurate by repeating this flow.
Hidehiko Shishido, Emi Kawasaki, Yutaka Ito, Youhei Kawamura, Toshiya Matsui, Itaru Kitahara
IEEE BigData6
2018 A Calibration Method of Floor Projection System for Learning Aids at School Gym
abstract
This paper proposes a calibration method for a large-scale floor projection system at a school gym. In our system, the projectors are installed on the ceiling of a school gym, and the contents are projected onto the floor. The projection results suffer from both projective distortion and lens distortion. It is hard to use chessboard-based calibration methods or self-calibration methods in our system due to the restrictive deploy environment and practical limitations. Our method is basing on the projector-camera system and "straight lines have to be straight". The experiments show that our method fulfills the on-site requirements of the floor projection system at school gym for learning aid usages.
Chun Xie, Hidehiko Shishido, Mika Oki, Yoshinari Kameda, Kenji Suzuki 0002, Itaru Kitahara
IPAS6
2018 A Calibration Method for Large-Scale Projection Based Floor Display System
abstract
We propose a calibration method for deploying a large-scale projection-based floor display system. In our system, multiple projectors are installed on the ceiling of a large indoor space like a gymnasium to achieve a large projection area on the floor. The projection results suffer from both perspective distortion and lens distortion. In this paper, we use projector-camera systems, in which a camera is mounted on each projector, with the “straight lines have to be straight” methodology, to calibrate our projection system. Different from conventional approaches, our method does not use any calibration board and makes no requirement on the overlapping among the projections and the cameras' fields of view.
Chun Xie, Hidehiko Shishido, Yoshinari Kameda, Kenji Suzuki 0002, Itaru Kitahara
VR5
2017 Method to generate disaster-damage map using 3D photometry and crowd sourcing
abstract
Thanks to the rapid progress of the Internet and mobile devices, information related to disaster areas can be collected through the Internet. To grasp the degree of damage in a disaster situation, the use of crowdsourcing for coordinating the individual efforts (micro tasks) of an enormous number of users (workers) on the Internet has been drawing attention as a means of quickly solving problems. However, the information gathered from the Internet is huge and diverse, so it is difficult to formulate as a crowdsourcing task. This paper proposes a conversion platform for the images of a disaster site photographed by various users as information about the site, integrating the images into a single map using 3D image processing, and providing the map to crowdsourcing as a micro task.
Koyo Kobayashi, Hidehiko Shishido, Yoshinari Kameda, Itaru Kitahara
IEEE BigData4
2017 Proactive preservation of world heritage by crowdsourcing and 3D reconstruction technology
abstract
Since over one million tourists annually visit the Angkor ruins, the effect on the buildings from the vibrations caused by these tourists is a huge problem for maintaining them. Such organisms as bryophytes, which adhere to the surface of the stones of the ruins, is another factor that damages them. Using crowdsourcing and 3D reconstruction technology, we are organizing a proactive preservation project for the Angkor Thom Bayon Temple, which is a world cultural heritage site. We evaluated its damaged parts and visualized the damaged state.
Hidehiko Shishido, Yutaka Ito, Youhei Kawamura, Toshiya Matsui, Atsuyuki Morishima, Itaru Kitahara
IEEE BigData6
2015 Remote Mixed Reality System Supporting Interactions with Virtualized Objects
abstract
Mixed Reality (MR) can merge real and virtual worlds seamlessly. This paper proposes a method to realize smooth collaboration using a remote MR, which makes it possible for geographically distributed users to share the same objects and communicate in real time as if they are at the same place. In this paper, we consider a situation where the users at local and remote sites perform a collaborative work, and real objects to be operated exist only at the local site. It is necessary to share the real objects between the two sites. However, prior studies have shown sharing real objects by duplication is either too costly or unrealistic. Therefore, we propose a method to share the objects by virtualizing the real objects using Computer Vision (CV) and then rendering the virtualized objects using MR. We have proposed a remote collaborative work system to create a smoother user experience for collaborative work with virtualized objects for remote users. Through experiments, we confirmed the effectiveness of our approach.
Itaru Kitahara, Yuichi Ohta
ISMAR2
2014 MR Simulation for Re-wallpapering a Room in a Free-Hand Movie
Masashi Ueda, Itaru Kitahara, Yuichi Ohta
MMM (1)2
2014 A projection-based mixed-reality display for exterior and interior of a building diorama
abstract
This paper proposes an interactive display system that displays both of the exterior and interior construction of a building diorama by using a projection-based Mixed-Reality (MR) technique, which is useful for understanding the complex construction and the spatial relationships between outside and inside. The users can hold and move the diorama model using their hands/body motion, so that they can observe the model from their favorite viewpoint. Our system obtains both of the user's information (the viewpoint and the gesture) and the diorama model's information (the pose) in 3D space by using two RGB-D cameras. The CG image corresponding to the user's viewpoint, gesture and the pose of the diorama is rendered by Dual Rendering algorithm in real time. As the result, the generated CG image is projected onto the diorama to realize MR display. We confirm the effectiveness of our proposed method by developing a pilot system.
Itaru Kitahara, Yoshinari Kameda, Yuichi Ohta
VRST2
2013 6DOF iterative closest point matching considering a priori with maximum a posteriori estimation
abstract
We present a new matching algorithm considering a priori (the prior probability) based on Bayes' theorem. Performance of point cloud registration between target and source clouds is effectively improved by introducing maximum a posteriori (MAP) estimation. The standard Iterative Closest Point (ICP) algorithm for the registration sometimes falls into misalignment due to measurement errors, narrow sensing field of view, or the movement of objects during measurement. Our approach resolves such problems by considering both the likelihood of the measurement and the prior probability of the initial guess for registration in the objective function. We have implemented a new 6DOF Iterative Closest Point matching using MAP estimation, and evaluated the method in real environments comparing with conventional registration methods. The experimental results have shown that our proposed method has wide convergence region and matches point clouds accurately preventing the misalignment problem.
Yoshitaka Hara, Shigeru Bando, Takashi Tsuboucffl, Akira Oshima, Itaru Kitahara, Yoshinari Kameda
IROS5
2013 A Trajectory Estimation Method for Badminton Shuttlecock Utilizing Motion Blur
Hidehiko Shishido, Itaru Kitahara, Yoshinari Kameda, Yuichi Ohta
PSIVT2
2012 Mixed-reality snapshot system using environmental depth sensors
Hiroyoshi Tsuru, Itaru Kitahara, Yuichi Ohta
ICPR2
2010 See-Through Vision: A Visual Augmentation Method for Sensing-Web
Yuichi Ohta, Yoshinari Kameda, Itaru Kitahara, Masayuki Hayashi, Shinya Yamazaki
IPMU (2)3
2010 Real-time soccer player tracking method by utilizing shadow regions
abstract
Our research aims to generate a player's view video stream by using a 3D free-viewpoint video technique. Since player trajectories are necessary to generate the video, we propose a real-time player trajectory estimation method by utilizing the shadow regions from soccer scenes. This paper describes our trial to realize real-time processing. We divide the process into capture and server computers. In addition, we reduced the processing cost with pipeline parallelization and optimization. We apply our proposed method to an actual soccer match held in a stadium and show its effectiveness.
Nozomu Kasuya, Itaru Kitahara, Yoshinari Kameda, Yuichi Ohta
ACM Multimedia2
2009 Lets go out: Research in outdoor mixed and augmented reality
Christian Sandor, Itaru Kitahara, Gerhard Reitmayr, Steven K. Feiner, Yuichi Ohta
ISMAR2
2009 Toward cinematizing our daily lives
Hansung Kim 0001, Ryuuki Sakamoto, Itaru Kitahara, Tomoji Toriyama, Kiyoshi Kogure
Multim. Tools Appl.3
2008 Robust trajectory estimation of soccer players by using two cameras
abstract
This paper proposes a method to estimate the trajectories of soccer players by using two cameras set in a large-scale outdoor space such as a soccer stadium, which is normally not advantageous for image processing. We estimate a player¿s position as the intersection of the primary axes of two body regions, corresponding to the two cameras, and a shadow region on the surface of the field. By projecting the foreground silhouette regions extracted from both cameras¿ images onto the soccer field, each player¿s body region and its shadow region are identified. Furthermore, our method utilizes color information of the players' uniforms to improve the accuracy of object tracking. We have applied our proposed method to a real soccer game held in a soccer stadium and demonstrated its effectiveness.
Nozomu Kasuya, Itaru Kitahara, Yoshinari Kameda, Yuichi Ohta
ICPR2
2008 Generating perceptually-correct shadows for mixed reality
abstract
When human cannot perceive the inconsistency of artificial shadows which are not physically correct, they are acceptable as “perceptually-correct” shadows. This paper focuses on the simplification of light-source models for generating the perceptually-correct artificial shadows. First, we conducted subjective evaluations to obtain knowledge about the human perception of the shadows. Then the knowledge was applied to control the resolution of the light-source map to generate perceptually-correct artificial shadows. Comparative studies among artificial and real shadows justified perceptually correctness. All experiments were done using still images, not videos. Our research becomes a reference to determine the resolution of light-source map in an MR scene.
Gaku Nakano, Itaru Kitahara, Yuichi Ohta
ISMAR2
2007 Robust Foreground Extraction Technique Using Gaussian Family Model and Multiple Thresholds
Hansung Kim 0001, Ryuuki Sakamoto, Itaru Kitahara, Tomoji Toriyama, Kiyoshi Kogure
ACCV (1)3
2007 Viewpoint-Dependent Quality Control on Microfacet Billboarding Model for Sports Video
abstract
We propose a new on-line modeling and rendering method for visualizing players in sports. This method targets the intricate geometry of players in sports such as wrestling. Our system can maintain modeling and rendering quality on producing free-viewpoint 3D video of the players by utilizing the virtual viewpoint information of the viewer. Viewers can move a virtual camera freely over the scene while the system manages the processing cost by changing the size of the voxels and microfacets used for spatial modeling of the players. The system first estimates the rough 3D shape of the players in voxel format, and then assigns a microfacet to each voxel. The texture of the microfacet is obtained from the video cameras surrounding the players. Our preliminary system was evaluated on the wrestling data taken in the Yoyogi National Gymnasium, Tokyo, Japan. We also conducted subjective evaluation of the free-viewpoint video thus produced and found a good balance of rendering quality and high frame-rate performance.
Hitoshi Furuya, Itaru Kitahara, Yoshinari Kameda, Yuichi Ohta
ICME2
2007 Face-to-Face Tabletop Remote Collaboration in Mixed Reality
abstract
This paper proposes a novel remote face-to-face mixed reality (MR) system that enables two people in distant places to share MR space. Challenging issues to realize such an MR system include capturing, sending, and rendering each user's appearance in real time. We developed a method to represent user's upper body and hands on the table as a single deformed-billboard. An MR Othello game is implemented as a test bed of the remote face-to-face MR system. Users can play the tabletop game as if their opponent were sitting across from the table, despite being physically separated. By detecting and sending the status of each real game board to the other site, both users feel that they are sharing tabletop objects.
Shinya Minatani, Itaru Kitahara, Yoshinari Kameda, Yuichi Ohta
ISMAR2
2007 Reliability-based 3D reconstruction in real environment
abstract
We present a practical 3D reconstruction method that guarantees robust visual hull construction in real environments where segmentation errors and occlusion exist. The proposed method consists of foreground extraction and reliability-based shape-from-silhouette, and they are connected by the intra-/inter-silhouette reliabilities. In foreground extraction, all regions are classified into four categories based on their intra-reliabilities. Then the reliability-based shape-from-silhouette technique reconstructs a visual hull by carving a 3D space based on the intra-/inter-silhouette reliabilities. The proposed method provides a reliable visual hull in real environments without much increment of the system complexity compared with conventional systems.
Hansung Kim 0001, Ryuuki Sakamoto, Itaru Kitahara, Tomoji Toriyama, Kiyoshi Kogure
ACM Multimedia3
2007 A Nested Marker for Augmented Reality
abstract
A Nested Marker, a novel visual marker for camera calibration in augmented reality (AR), enables accurate calibration even when the observer is moving very close to or far away from the marker. Our proposed Nested Marker has a recursive layered structure. One marker at an upper layer contains four smaller markers at the lower layer. Smaller markers can also have lower-layer markers nesting inside them. Each marker can be identified by its inside pattern, so the system can select a proper calibration parameter set for the marker. When the observer views the marker close-up, the lowest layer marker will work. When the observer views the marker from a distance, the top-layer marker will work. It is also possible to simultaneously utilize all visible markers in different layers for more stable calibration. Note that Nested Marker can be used in a standard ARToolkit framework. We have also developed an AR system to demonstrate the ability of Nested Marker
Keisuke Tateno, Itaru Kitahara, Yuichi Ohta
VR2
2007 Live 3D Video in Soccer Stadium
Yuichi Ohta, Itaru Kitahara, Yoshinari Kameda, Hiroyuki Ishikawa, Takayoshi Koyama
Int. J. Comput. Vis.2
2006 Photometric inconsistency on a mixed-reality face
abstract
A mixed-reality face (MR face) is a mosaic face with real and virtual facial parts, presented by overlaying a virtual facial part on a real face using mixed-reality techniques. An MR face is an effective means to improve communication in mixed-reality space by restoring the eye expressions lost when wearing HMDs. Photometric registration between the real and virtual parts is important because our eyes are very sensitive, even to small changes in human faces. However, efforts to achieve perfect 'physical' photometric registration on an MR face are not feasible in an ordinary MR space. Therefore, it is essential to clarify the sensitivity of our eyes to the photometric inconsistencies on an MR face, and to concentrate on resolving them. In this paper, we first present the results of a systematic experiment that evaluated our sensitivity to the photometric inconsistencies on an MR face. Then, a technique to resolve the inconsistency and an experimental system to demonstrate the effectiveness of an MR face are described.
Masayuki Takemura, Itaru Kitahara, Yuichi Ohta
ISMAR2
2003 Live Mixed-Reality 3D Video in Soccer Stadium
abstract
This paper proposes a method to realize a 3D video display system that can capture video from multiple cameras, reconstruct 3D models and transmit 3D video data in real time. We represent a target object with a simplified 3D model consisting of a single plane and a 2D texture extracted from multiple cameras. This 3D model is simple enough to be transmitted via a network. We have developed a prototype system that can capture multiple videos, reconstruct 3D models, transmit the models via a network, and display 3D video in real time. A 3D video of a typical soccer scene that includes a dozen players was processed at 26 frames per second.
Takayoshi Koyama, Itaru Kitahara, Yuichi Ohta
ISMAR2
2003 Live 3D video in soccer stadium
abstract
A live 3D video system in a real soccer stadium is presented. Players on the pitch are represented with a simplified 3D model. The model is reconstructed from multiple videos obtained by nine CCD cameras surrounding the pitch. Observers can watch the game from arbitrary viewing position. By using real textures of the players, realistic video presentation is possible. All processes are fully automatic and real time. The system can transmit live 3D video to distant places.
Takayoshi Koyama, Itaru Kitahara, Yuichi Ohta
SIGGRAPH2
2003 Scalable 3D Representation for 3D Video Display in a Large-scale Spac
abstract
The authors introduce their research for realizing a 3D video display system in a very large-scale space such as a soccer stadium, concert hall, etc. They propose a method for describing the shape of a 3D object with a set of planes in order to synthesize a novel view of the object effectively. The most effective layout of the planes can be determined based on the relative locations of an observer's viewing position, multiple cameras, and 3D objects. A method is described for controlling the LOD of the 3D representation by adjusting the orientation, interval, and resolution of planes. The data size of the 3D model and the processing time can be reduced drastically. The effectiveness of the proposed method is demonstrated by experimental results.
Itaru Kitahara, Yuichi Ohta
VR1