Tomokazu Sato

dblp:55/182 · DBLP profile ↗
← Back
43ranked-venue papers
3as first author
0since 2021 · last 2018
0000-0002-7337-8157ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 2 first-authorArtificial intelligence and machine learning · 17 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
8 papers
Virtual and augmented reality · 56% Image and video processing · 26% Multimedia analysis and retrieval · 10%
Artificial intelligence
4 papers
Video understanding and tracking · 48% 3D vision · 26% Representation and self-supervised learning · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 25 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality › mixed reality
diminished reality
0.632016
Diminished Reality Based on Image Inpainting Considering Background Geometry · IEEE Trans. Vis. Comput. Graph. 2016
Diminished reality considering background structures · ISMAR 2013
AR marker hiding based on image inpainting and reflection of illumination changes · ISMAR 2012
Image and video processing › image restoration
image inpainting
0.632016
Diminished Reality Based on Image Inpainting Considering Background Geometry · IEEE Trans. Vis. Comput. Graph. 2016
Diminished reality considering background structures · ISMAR 2013
AR marker hiding based on image inpainting and reflection of illumination changes · ISMAR 2012
Virtual and augmented reality
augmented reality
0.522017
Augmented Reality Marker Hiding with Texture Deformation · IEEE Trans. Vis. Comput. Graph. 2017
Indirect augmented reality considering real-world illumination change · ISMAR 2014
Virtual and augmented reality › augmented reality › augmented reality rendering
marker hiding
0.422017
Augmented Reality Marker Hiding with Texture Deformation · IEEE Trans. Vis. Comput. Graph. 2017
AR marker hiding based on image inpainting and reflection of illumination changes · ISMAR 2012
Computer vision › Video understanding and tracking
action recognition
0.312018
Summarization of User-Generated Sports Video by Using Deep Action Recognition Features · IEEE Trans. Multim. 2018
Multimedia analysis and retrieval
video summarization
0.312018
Summarization of User-Generated Sports Video by Using Deep Action Recognition Features · IEEE Trans. Multim. 2018
Image and video processing
motion estimation
0.312017
Augmented Reality Marker Hiding with Texture Deformation · IEEE Trans. Vis. Comput. Graph. 2017
Virtual and augmented reality › visual realism
illumination consistency
0.222014
Indirect augmented reality considering real-world illumination change · ISMAR 2014
AR marker hiding based on image inpainting and reflection of illumination changes · ISMAR 2012
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.212013
What is the Most EfficientWay to Select Nearest Neighbor Candidates for Fast Approximate Nearest Neighbor Search? · ICCV 2013
Information retrieval
candidate selection
0.212013
What is the Most EfficientWay to Select Nearest Neighbor Candidates for Fast Approximate Nearest Neighbor Search? · ICCV 2013
Information retrieval › similarity search
nearest neighbor search
0.212013
What is the Most EfficientWay to Select Nearest Neighbor Candidates for Fast Approximate Nearest Neighbor Search? · ICCV 2013
Rendering
view-dependent rendering
0.212013
Augmented reality image generation with virtualized real objects using view-dependent texture and geometry · ISMAR 2013
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
deep feature extraction
0.112018
Summarization of User-Generated Sports Video by Using Deep Action Recognition Features · IEEE Trans. Multim. 2018
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.112016
Diminished Reality Based on Image Inpainting Considering Background Geometry · IEEE Trans. Vis. Comput. Graph. 2016
Computational photography and imaging
omnidirectional imaging
0.112014
Indirect augmented reality considering real-world illumination change · ISMAR 2014
Computational photography and imaging › image-based modeling
3d reconstruction from images
0.012013
Augmented reality image generation with virtualized real objects using view-dependent texture and geometry · ISMAR 2013
Virtual and augmented reality
tracking and registration
0.012013
Diminished reality considering background structures · ISMAR 2013
Computer vision › 3D vision
camera calibration
0.012004
Extrinsic Camera Parameter Recovery from Multiple Image Sequences Captured by an Omni-Directional Multi-camera System · ECCV (2) 2004
Computer vision › 3D vision › multi-view geometry › multi-view vision
multi-camera systems
0.012004
Extrinsic Camera Parameter Recovery from Multiple Image Sequences Captured by an Omni-Directional Multi-camera System · ECCV (2) 2004
Virtual and augmented reality › augmented reality
augmented reality rendering
0.012012
AR marker hiding based on image inpainting and reflection of illumination changes · ISMAR 2012
Immersive interaction
telepresence
0.012003
Telepresence System Using High-Resolution Omnidirectional Movies and a Reactive Display · ISMAR 2003
Computer vision › 3D vision
3d reconstruction
0.012002
Dense 3-D Reconstruction of an Outdoor Scene by Hundreds-Baseline Stereo Using a Hand-Held Video Camera · Int. J. Comput. Vis. 2002
Computer vision › 3D vision › stereo vision
dense stereo reconstruction
0.012002
Dense 3-D Reconstruction of an Outdoor Scene by Hundreds-Baseline Stereo Using a Hand-Held Video Camera · Int. J. Comput. Vis. 2002
Virtual and augmented reality
locomotion
0.012003
Telepresence System Using High-Resolution Omnidirectional Movies and a Reactive Display · ISMAR 2003
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.012002
Dense 3-D Reconstruction of an Outdoor Scene by Hundreds-Baseline Stereo Using a Hand-Held Video Camera · Int. J. Comput. Vis. 2002

Methods — techniques the papers use, named apart from their topics

plane approximation · 0.7deep neural network · 0.7action recognition features · 0.7exemplar-based inpainting · 0.5sparse feature point analysis · 0.3homography · 0.3image selection · 0.2visual SLAM · 0.2texture synthesis · 0.2quantization · 0.2distance computation · 0.2depth map inflation · 0.2multi-screen projection · 0.03d position sensing · 0.0hundreds-baseline stereo · 0.0hand-held video camera · 0.0
YearPublicationVenuePosition
2018 Iterative applications of image completion with CNN-based failure detection
Takahiro Tanaka, Norihiko Kawai, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya
J. Vis. Commun. Image Represent.4
2018 Summarization of User-Generated Sports Video by Using Deep Action Recognition Features
abstract
Automatically generating a summary of a sports video poses the challenge of detecting interesting moments, or highlights, of a game. Traditional sports video summarization methods leverage editing conventions of broadcast sports video that facilitate the extraction of high-level semantics. However, user-generated videos are not edited and, thus, traditional methods are not suitable to generate a summary. In order to solve this problem, this paper proposes a novel video summarization method that uses players' actions as a cue to determine the highlights of the original video. A deep neural-network-based approach is used to extract two types of action-related features and to classify video segments into interesting or uninteresting parts. The proposed method can be applied to any sports in which games consist of a succession of actions. Especially, this paper considers the case of Kendo (Japanese fencing) as an example of a sport to evaluate the proposed method. The method is trained using Kendo videos with ground truth labels that indicate the video highlights. The labels are provided by annotators possessing a different experience with respect to Kendo to demonstrate how the proposed method adapts to different needs. The performance of the proposed method is compared with several combinations of different features, and the results show that it outperforms previous summarization methods.
Antonio Tejero-de-Pablos, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya, Marko Linna, Esa Rahtu
IEEE Trans. Multim.3
2017 Novel view synthesis with light-weight view-dependent texture mapping for a stereoscopic HMD
abstract
The proliferation of off-the-shelf head-mounted displays (HMDs) let end-users enjoy virtual reality applications, some of which render a real-world scene using a novel view synthesis (NVS) technique. View-dependent texture mapping (VDTM) has been studied for NVS due to its photo-realistic quality. The VDTM technique renders a novel view by adaptively selecting textures from the most appropriate images. However, this process is computationally expensive because VDTM scans every captured image. For stereoscopic HMDs, the situation is much worse because we need to render novel views once for each eye, almost doubling the cost. This paper proposes light-weight VDTM tailored for an HMD. In order to reduce the computational cost in VDTM, our method leverages the overlapping fields of view between a stereoscopic pair of HMD images and pruning the images to be scanned. We show that the proposed method drastically accelerates the VDTM process without spoiling the image quality through a user study.
Thiwat Rongsirigul, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya
ICME3
2017 ReMagicMirror: Action Learning Using Human Reenactment with the Mirror Metaphor
Fabian Lorenzo Dayrit, Ryosuke Kimura, Yuta Nakashima, Ambrosio Blanco, Hiroshi Kawasaki, Katsushi Ikeuchi, Tomokazu Sato, Naokazu Yokoya
MMM (1)7
2017 Increasing pose comprehension through augmented reality reenactment
Fabian Lorenzo Dayrit, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya
Multim. Tools Appl.3
2017 Video summarization using textual descriptions for authoring video blogs
Mayu Otani, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya
Multim. Tools Appl.3
2017 Augmented Reality Marker Hiding with Texture Deformation
abstract
Augmented reality (AR) marker hiding is a technique to visually remove AR markers in a real-time video stream. A conventional approach transforms a background image with a homography matrix calculated on the basis of a camera pose and overlays the transformed image on an AR marker region in a real-time frame, assuming that the AR marker is on a planar surface. However, this approach may cause discontinuities in textures around the boundary between the marker and its surrounding area when the planar surface assumption is not satisfied. This paper proposes a method for AR marker hiding without discontinuities around texture boundaries even under nonplanar background geometry without measuring it. For doing this, our method estimates the dense motion in the marker's background by analyzing the motion of sparse feature points around it, together with a smooth motion assumption, and deforms the background image according to it. Our experiments demonstrate the effectiveness of the proposed method in various environments with different background geometries and textures.
Norihiko Kawai, Tomokazu Sato, Yuta Nakashima, Naokazu Yokoya
IEEE Trans. Vis. Comput. Graph.2
2016 Ultra-Shallow DoF Imaging Using Faced Paraboloidal Mirrors
Ryoichiro Nishi, Takahito Aoto 0002, Norihiko Kawai, Tomokazu Sato, Yasuhiro Mukaigawa, Naokazu Yokoya
ACCV (3)4
2016 Human action recognition-based video summarization for RGB-D personal sports video
abstract
Automatic sports video summarization poses the challenge of acquiring semantics of the original video, and existing work leverages various knowledge in application domains, e.g., structure of games and editing conventions. In this paper, we propose a personal sports video summarization method for self-recorded RGB-D videos, which became available to the public due to the commodification of off-the-shelf RGB-D sensors. We focus on sports whose games consist of a succession of actions and, unlike previous research, we use human action recognition on the depth sequences in order to acquire higher level semantics of the video. The recognition results are used along with an entropy-based activity measure to train a hidden Markov model of the highlights of different games to extract a summary from the original RGB-D video. We trained our novel highlights model with the subjective opinion of users with different experience in the sport. We took Kendo, a martial art, as an example sport to evaluate our method, and objectively/subjectively investigated the accuracy and quality of the generated summaries.
Antonio Tejero-de-Pablos, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya
ICME3
2016 Moving object detection from a point cloud using photometric and depth consistencies
abstract
3D models of outdoor environments have been used for several applications such as a virtual earth system and a vision-based vehicle safety system. 3D data for constructing such 3D models are often measured by an on-vehicle system equipped with laser rangefinders, cameras, and GPS/IMU. However, 3D data of moving objects on streets lead to inaccurate 3D models when modeling outdoor environments. To solve this problem, this paper proposes a moving object detection method for point clouds by minimizing an energy function based on photometric and depth consistencies assuming that input data consist of synchronized point clouds, images, and camera poses from a single sequence captured with a moving on-vehicle system.
Atsushi Takabe, Hikari Takehara, Norihiko Kawai, Tomokazu Sato, Takashi Machida, Satoru Nakanishi, Naokazu Yokoya
ICPR4
2016 Diminished Reality Based on Image Inpainting Considering Background Geometry
abstract
Diminished reality aims to remove real objects from video images and fill in the missing regions with plausible background textures in real time. Most conventional methods based on image inpainting achieve diminished reality by assuming that the background around a target object is almost planar. This paper proposes a new diminished reality method that considers background geometries with less constraints than the conventional ones. In this study, we approximate the background geometry by combining local planes, and improve the quality of image inpainting by correcting the perspective distortion of texture and limiting the search area for finding similar textures as exemplars. The temporal coherence of texture is preserved using the geometries and camera pose estimated by visual-simultaneous localization and mapping (SLAM). The mask region that includes a target object is robustly set in each frame by projecting a 3D region, rather than tracking the object in 2D image space. The effectiveness of the proposed method is successfully demonstrated using several experimental environments.
Norihiko Kawai, Tomokazu Sato, Naokazu Yokoya
IEEE Trans. Vis. Comput. Graph.2
2015 Image resolution enhancement based on novel view synthesis
abstract
This paper proposes an example-based method to increase the resolution of a low-resolution image. In the proposed method, we generate example images by a novel view synthesis technique using 3D geometry reconstruction and camera pose estimation from a video or images capturing the same scene. We then increase the resolution by minimizing an energy function by searching for the optimal example from the generated example images. The proposed method has less limitations on camera positions and geometry of the target scene than those in conventional methods. Experiments demonstrate the effectiveness of the proposed method by qualitatively comparing the results of the proposed and conventional methods.
Yusuke Hayashi, Norihiko Kawai, Tomokazu Sato, Naokazu Yokoya
ICIP3
2015 Textual description-based video summarization for video blogs
abstract
Recent popularization of camera devices, including action cams and smartphones, enables us to record videos in everyday life and share them through the Internet. Video blog is a recent approach for sharing videos, in which users enjoy expressing themselves in blog posts with attractive videos. Generating such videos, however, requires users to review vast amount of raw videos and edit them appropriately, which keeps users away from doing so. In this paper, we propose a novel video summarization method for helping users to create a video blog post. Unlike typical video summarization methods, the proposed method utilizes the text, which is written for a video blog post, and makes the video summary consistent with the content of the text. For this, we perform video summarization by solving an optimization problem, in which an objective function involves the content similarity between the summarized video and the text. Our user study with 20 participants has demonstrated that our proposed method is suitable to create video blog posts compared with conventional methods for video summarization.
Mayu Otani, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya
ICME3
2015 Bundle adjustment using aerial images with two-stage geometric verification
Hideyuki Kume, Tomokazu Sato, Naokazu Yokoya
Comput. Vis. Image Underst.2
2014 Free-viewpoint AR human-motion reenactment based on a single RGB-D video stream
abstract
When observing a person (an actor) performing or demonstrating some activity for the purpose of learning the action, it is best for the viewers to be present at the same time and place as the actor. Otherwise, a video must be recorded. However, conventional video only provides two-dimensional (2D) motion, which lacks the original third dimension of motion. In the presence of some ambiguity, it may be hard for the viewer to comprehend the action with only two dimensions, making it harder to learn the action. This paper proposes an augmented reality system to reenact such actions at any time the viewer wants, in order to aid comprehension of 3D motion. In the proposed system, a user first captures the actor's motion and appearance, using a single RGB-D camera. Upon a viewer's request, our system displays the motion from an arbitrary viewpoint using a rough 3D model of the subject, made up of cylinders, and selecting the most appropriate textures based on the viewpoint and the subject's pose. We evaluate the usefulness of the system and the quality of the displayed images by user study.
Fabian Lorenzo Dayrit, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya
ICME3
2014 Vehicle Driver Face Detection in Various Sunlight Environments Using Composed Face Images
abstract
The purpose of this study is to increase the face detection accuracy in vehicle cabin. Although existing face detectors employed in consumer applications already have sufficient face detection accuracy for many situations, we revealed that detection rate of existing face detector is drastically decreased by shadow on the driver's face caused by sunlight whose relative direction to the driver is continuously changed while driving. In order to overcome this problem, we increase the number of driver's faces in training dataset by synthesizing the shadowed driver's faces from various directions of sunlight which are created using an image composing technique. In experiment, we found that the 20% to 40% shadowed faces should be blended into training dataset in terms of the generality and the adaptability for robust drivers' face detection.
Haruo Matsuo, Tomokazu Sato, Naokazu Yokoya
ICPR2
2014 Indirect augmented reality considering real-world illumination change
abstract
Indirect augmented reality (IAR) utilizes pre-captured omnidirectional images and offline superimposition of virtual objects for achieving high-quality geometric and photometric registration. Meanwhile, IAR causes inconsistency between the real world and the pre-captured image. This paper describes the first-ever study focusing on the temporal inconsistency issue in IAR. We propose a novel IAR system which reflects real-world illumination change by selecting an appropriate image from a set of images pre-captured under various illumination. Results of a public experiment show that the proposed system can improve the realism in IAR.
Fumio Okura, Takayuki Akaguma, Tomokazu Sato, Naokazu Yokoya
ISMAR3
2014 Evaluation of image processing algorithms on vehicle safety system based on free-viewpoint image rendering
abstract
Development of algorithms for vehicle safety systems, which support safety driving, takes a long period of time and a huge cost because it requires an evaluation stage where huge combinations of possible driving situations should be evaluated by using videos which are captured beforehand in real environments. In this paper, we address this problem by using free viewpoint images instead of the real images. More concretely, we generate free-viewpoint images from a combination of a 3D point cloud measured by laser scanners and an omni-directional image sequence acquired in a real environment. We basically rely on the 3D point cloud for geometrically correct virtual viewpoint images. In order to remove the holes caused by the unmeasured region of the 3D point cloud and to remove false silhouettes in surface reconstruction, we have developed a technique of free-viewpoint image generation that uses both a 3D point cloud and depth information extracted from images. In the experiments, we have evaluated our framework with a white line detection algorithm and experimental results have shown the applicability of free-viewpoint images for evaluation of algorithms.
Akitaka Oko, Tomokazu Sato, Hideyuki Kume, Takashi Machida, Naokazu Yokoya
Intelligent Vehicles Symposium2
2013 What is the Most EfficientWay to Select Nearest Neighbor Candidates for Fast Approximate Nearest Neighbor Search?
abstract
Approximate nearest neighbor search (ANNS) is a basic and important technique used in many tasks such as object recognition. It involves two processes: selecting nearest neighbor candidates and performing a brute-force search of these candidates. Only the former though has scope for improvement. In most existing methods, it approximates the space by quantization. It then calculates all the distances between the query and all the quantized values (e.g., clusters or bit sequences), and selects a fixed number of candidates close to the query. The performance of the method is evaluated based on accuracy as a function of the number of candidates. This evaluation seems rational but poses a serious problem; it ignores the computational cost of the process of selection. In this paper, we propose a new ANNS method that takes into account costs in the selection process. Whereas existing methods employ computationally expensive techniques such as comparative sort and heap, the proposed method does not. This realizes a significantly more efficient search. We have succeeded in reducing computation times by one-third compared with the state-of-theart on an experiment using 100 million SIFT features.
Masakazu Iwamura, Tomokazu Sato, Koichi Kise
ICCV2
2013 Key-Region Detection for Document Images - Application to Administrative Document Retrieval
abstract
In this paper we argue that a key-region detector designed to take into account the special characteristics of document images can result in the detection of less and more meaningful key-regions. We propose a fast key-region detector able to capture aspects of the structural information of the document, and demonstrate its efficiency by comparing against standard detectors in an administrative document retrieval scenario. We show that using the proposed detector results to a smaller number of detected key-regions and higher performance without any drop in speed compared to standard state of the art detectors.
Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Tomokazu Sato, Masakazu Iwamura, Koichi Kise
ICDAR5
2013 Detection of 3D points on moving objects from point cloud data for 3D modeling of outdoor environments
abstract
A 3D modeling technique for an urban environment can be applied to several applications such as landscape simulations, navigational systems, and mixed reality systems. In this field, the target environment is first measured using several types of sensors (laser rangefinders, cameras, GPS sensors, and gyroscopes). A 3D model of the environment is then constructed based on the results of the 3D measurements. In this 3D modeling process, 3D points that exist on moving objects become obstacles or outliers to enable the construction of an accurate 3D model. To solve this problem, we propose a method for detecting 3D points on moving objects from 3D point cloud data based on photometric consistency and knowledge of the road environment. In our method, 3D points on moving objects are detected based on luminance variations obtained by projecting 3D points onto omnidirectional images. After detecting 3D the points based on evaluation value, the points are detected using prior information of the road environment.
Tsunetake Kanatani, Hideyuki Kume, Takafumi Taketomi, Tomokazu Sato, Naokazu Yokoya
ICIP4
2013 Teleoperation of mobile robots by generating augmented free-viewpoint images
abstract
This paper proposes a teleoperation interface by which an operator can control a robot from freely configured viewpoints using realistic images of the physical world. The viewpoints generated by the proposed interface provide human operators with intuitive control using a head-mounted display and head tracker, and assist them to grasp the environment surrounding the robot. A state-of-the-art free-viewpoint image generation technique is employed to generate the scene presented to the operator. In addition, an augmented reality technique is used to superimpose a 3D model of the robot onto the generated scenes. Through evaluations under virtual and physical environments, we confirmed that the proposed interface improves the accuracy of teleoperation.
Fumio Okura, Yuko Ueda, Tomokazu Sato, Naokazu Yokoya
IROS3
2013 Diminished reality considering background structures
abstract
This paper proposes a new diminished reality method for 3D scenes considering background structures. Most conventional methods using image inpainting assumes that the background around a target object is almost planar. In this study, approximating the background structure by the combination of local planes, perspective distortion of texture is corrected and searching area is limited for improving the quality of image inpainting. The temporal coherence of texture is preserved using the estimated structures and camera pose estimated by Visual-SLAM.
Norihiko Kawai, Tomokazu Sato, Naokazu Yokoya
ISMAR2
2013 Augmented reality image generation with virtualized real objects using view-dependent texture and geometry
abstract
Augmented reality (AR) images with virtualized real objects can be used for various applications. However, such AR image generation requires hand-crafted 3D models of that objects, which are usually not available. This paper proposes a view-dependent texture (VDT)- and view-dependent geometry (VDG)-based method for generating high quality AR images, which uses 3D models automatically reconstructed from multiple images. Since the quality of reconstructed 3D models is usually insufficient, the proposed method inflates the objects in the depth map as VDG to repair chipped object boundaries and assigns a color to each pixel based on VDT to reproduce the detail of the objects. Background pixel exposure due to inflation is suppressed by the use of the foreground region extracted from the input images. Our experimental results have demonstrated that the proposed method can successfully reduce above visual artifacts.
Yuta Nakashima, Tomokazu Sato, Yusuke Uno, Naokazu Yokoya, Norihiko Kawai
ISMAR2
2013 Generation of a Super-Resolved Stereo Video Using Two Synchronized Videos with Different Magnifications
Yusuke Hayashi, Norihiko Kawai, Tomokazu Sato, Miyuki Okumoto, Naokazu Yokoya
PSIVT3
2012 Position estimation of near point light sources using a clear hollow sphere
Takahito Aoto 0002, Takafumi Taketomi, Tomokazu Sato, Yasuhiro Mukaigawa, Naokazu Yokoya
ICPR3
2012 AR marker hiding based on image inpainting and reflection of illumination changes
abstract
ISMAR 2012: IEEE International Symposium on Mixed and Augmented Reality , Nov 5-8, 2012 , Atlanta, GA, USA
Norihiko Kawai, Masayoshi Yamasaki, Tomokazu Sato, Naokazu Yokoya
ISMAR3
2011 Surface completion of shape and texture based on energy minimization
abstract
In this paper, we propose a novel surface completion method to generate plausible shapes and textures for missing regions of 3D models. The missing regions are filled in by minimizing two energy functions for shape and texture, which are both based on similarities between the missing region and the rest of the object; in doing so, we take into account the positive correlation between shape and texture. We demonstrate the effectiveness of the proposed method experimentally by applying it to two models.
Norihiko Kawai, Avideh Zakhor, Tomokazu Sato, Naokazu Yokoya
ICIP3
2011 Real-time and accurate extrinsic camera parameter estimation using feature landmark database for augmented reality
Takafumi Taketomi, Tomokazu Sato, Naokazu Yokoya
Comput. Graph.2
2010 Extrinsic Camera Parameter Estimation Using Video Images and GPS Considering GPS Positioning Accuracy
abstract
This paper proposes a method for estimating extrinsic camera parameters using video images and position data acquired by GPS. In conventional methods, the accuracy of the estimated camera position largely depends on the accuracy of GPS positioning data because they assume that GPS position error is very small or normally distributed. However, the actual error of GPS positioning easily grows to the 10m level and the distribution of these errors is changed depending on satellite positions and conditions of the environment. In order to achieve more accurate camera positioning in outdoor environments, in this study, we have employed a simple assumption that true GPS position exists within a certain range from the observed GPS position and the size of the range depends on the GPS positioning accuracy. Concretely, the proposed method estimates camera parameters by minimizing an energy function that is defined by using the reprojection error and the penalty term for GPS positioning.
Hideyuki Kume, Takafumi Taketomi, Tomokazu Sato, Naokazu Yokoya
ICPR3
2010 Efficient hundreds-baseline stereo by counting interest points for moving omni-directional multi-camera system
Tomokazu Sato, Naokazu Yokoya
J. Vis. Commun. Image Represent.1
2009 Generation of an Omnidirectional Video without Invisible Areas Using Image Inpainting
Norihiko Kawai, Kotaro Machikita, Tomokazu Sato, Naokazu Yokoya
ACCV (2)3
2009 Efficient surface completion using principal curvature and its evaluation
abstract
Surface completion is a technique for filling missing regions in 3D models measured by range scanners and videos. Conventionally, although missing regions were filled with the similar shape in a model, the completion process was fairly inefficient because the whole region in the model was searched for the similar shape. In this paper, the completion is efficiently performed using principal curvatures of local shape. In experiments, the effectiveness of the proposed method is successfully verified with subjective evaluation. In addition, the quantitative evaluation which has not been in the literature is newly performed.
Norihiko Kawai, Tomokazu Sato, Naokazu Yokoya
ICIP2
2009 Image Inpainting Considering Brightness Change and Spatial Locality of Textures and Its Evaluation
Norihiko Kawai, Tomokazu Sato, Naokazu Yokoya
PSIVT2
2008 Surface completion by minimizing energy based on similarity of shape
abstract
3D mesh models generated with range scanner or video images often have holes due to many occlusions by other objects and the object itself. This paper proposes a novel method to fill the missing parts in the incomplete models. The missing parts are filled by minimizing the energy function, which is defined based on similarity of local shape between the missing region and the rest of the object. The proposed method can generate complex and consistent shapes in the missing region. In the experiment, the effectiveness of the method is successfully demonstrated by applying it to complex shape objects with missing parts.
Norihiko Kawai, Tomokazu Sato, Naokazu Yokoya
ICIP2
2008 Real-time outdoor pre-visualization method for videographers - real-time geometric registration using point-based model -
abstract
This paper describes a real-time pre-visualization method using augmented reality techniques for videographers. It enables them to test camera work without real actors in a real environment. As a substitute for real actors, virtual ones are superimposed on a live video in real-time according to a real camera motion and an illumination condition. The key technique of this method is real-time motion estimation of a camera, which can be applied to unknown complex environments including natural objects. In our method, geometric and photometric registration problems for such unknown environments are solved to realize the above visualization. A prototype system demonstrates availability of the pre-visualization method.
Sei Ikeda, Takafumi Taketomi, Bunyo Okumura, Tomokazu Sato, Masayuki Kanbara, Naokazu Yokoya, Kunihiro Chihara
ICME4
2008 Real-time camera position and posture estimation using a feature landmark database with priorities
abstract
In the field of computer vision, many kinds of camera parameter estimation methods have been proposed. As one of these methods, an extrinsic camera parameter estimation method that uses pre-constructed feature landmark database has been studied. In this method, extrinsic camera parameters of video images are estimated from correspondences between landmarks and image features. Although this method can work in a large outdoor environment, its computational cost in matching process is expensive and it cannot work in real-time. In this paper, to achieve real-time camera parameter estimation, the number of matching candidates are reduced by using priorities of landmarks that are determined from previously captured video sequences.
Takafumi Taketomi, Tomokazu Sato, Naokazu Yokoya
ICPR2
2007 Video Mosaicing Based on Structure from Motion for Distortion-Free Document Digitization
Akihiko Iketani, Tomokazu Sato, Sei Ikeda, Masayuki Kanbara, Noboru Nakajima, Naokazu Yokoya
ACCV (2)2
2006 Super-Resolved Video Mosaicing for Documents Based on Extrinsic Camera Parameter Estimation
Akihiko Iketani, Tomokazu Sato, Sei Ikeda, Masayuki Kanbara, Noboru Nakajima, Naokazu Yokoya
ACCV (2)2
2006 Extrinsic Camera Parameter Estimation Based-on Feature Tracking and GPS Data
Yuji Yokochi, Sei Ikeda, Tomokazu Sato, Naokazu Yokoya
ACCV (1)3
2004 Extrinsic Camera Parameter Recovery from Multiple Image Sequences Captured by an Omni-Directional Multi-camera System
Tomokazu Sato, Sei Ikeda, Naokazu Yokoya
ECCV (2)1
2003 Telepresence System Using High-Resolution Omnidirectional Movies and a Reactive Display
abstract
This paper describes a novel telepresence system that uses high-resolution movies and a reactive display system with a treadmill. In this system, users can walk through a virtualized environment by actually walking on a treadmill. According to walking motion which is detected by using 3-D position sensors put on both legs, the virtualized environment captured by an omnidirectional multi-camera system is projected on a multi-screen display.
Sei Ikeda, Tomokazu Sato, Masayuki Kanbara, Naokazu Yokoya
ISMAR2
2002 Dense 3-D Reconstruction of an Outdoor Scene by Hundreds-Baseline Stereo Using a Hand-Held Video Camera
Tomokazu Sato, Masayuki Kanbara, Naokazu Yokoya, Haruo Takemura
Int. J. Comput. Vis.1