Haruo Takemura

dblp:t/HaruoTakemura · DBLP profile ↗
← Back
89ranked-venue papers
2as first author
13since 2021 · last 2025
0000-0002-4645-6439ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 66 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 32 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 30 · 6 since 2021Systems, architecture and hardware · 5 · 3 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 3
YearPublicationVenuePosition
2025 Detective Networks: Enhancing Disaster Recognition in Images Through Attention Shifting Using Optimal Masking
abstract
Aerial investigation is used for surveying damage and identifying post-disaster events through imagery data. However, the challenge lies in detecting disaster-related areas within aerial or shipborne images, as these can appear as minor regions, making recognition difficult. To address this challenge, we introduce the Detective Network (DeNet), designed to optimally mask images, thereby shifting the attention of machine learning models towards these small yet crucial regions. Utilizing the concepts of patch and anchor box, DeNet incorporates a masking candidate layer and a masking layer to facilitate optimal masking. Our experimental findings are compelling; by preprocessing images with DeNet before analysis using an image captioning model, we achieved a remarkable accuracy of 92.91% in landslide detection from side-view image captions and 87.50% for shipborne view detection. The result demonstrates the efficacy of DeNet in enhancing the recognition of disaster-related areas in challenging imaging conditions.
Narongthat Thanyawet, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura
WACV4
2025 User-Centric Locomotion Techniques for Virtual Reality Games: A Survey of User Needs and Issues
abstract
Virtual reality (VR) video games that are played on a VR headset are becoming increasingly common in households, and though many games require players to navigate vast virtual spaces, most homes cannot provide a large enough physical space to encompass the entire virtual space. Thus, VR video games that require locomotion often provide users with alternative locomotion techniques. While teleportation or steering is typically used as a standard, new techniques can overcome remaining problems, such as motion sickness. However, a holistic perspective of user needs and issues regarding these techniques in practical situations has not been studied on a broad basis. To address this gap in the literature and contribute to future VR video game development and research, we conducted 16 semi-structured interviews and surveyed 88 participants to help explore issues regarding existing locomotion techniques. Our results revealed preferences related to teleportation versus steering and the postures that users adopt while playing VR video games, along with user needs for locomotion techniques in each posture.
Daichi Hirobe, Shizuka Shirai, Jason Orlosky, Mehrasa Alizadeh, Masato Kobayashi 0001, Yuuki Uranishi, Photchara Ratsamee, Haruo Takemura
IEEE Trans. Games8
2025 CaliView: Continuous Viewpoint Calibration Using Dynamic Rotation Gain Control
abstract
Head tracking allows users of Virtual Reality (VR) to freely rotate their heads 360 degrees while exploring virtual environments. When using VR in a limited space, the ability to physically rotate one's head is limited to a specific range. To address this issue, previous studies have proposed methods to employ distinct rotation factors for real and imaginary rotations. However, its primary usage lies in redirected walking; thus, it is unsuitable for seated VR. In this article, we propose CaliView, which consistently adjusts the user's perspective to always face an optimal direction in VR, all while ensuring a comfortable posture. CaliView continuously controls the rotation gain to ensure that the disparity between the present body orientation and the optimal orientation is reduced to zero, encouraging the user to assume a forward-facing position with the target orientation. To assess the suggested approach, CaliView, we experimented to compare three conditions: CaliView, snap turning only (SnapTurn), and hybrid of CaliView and snap turning (Hybrid). The research findings suggest that CaliView functions as a useful reorientation technique, enabling implicit reorientation without sacrificing the user experience. Additionally, this study showcases its compatibility with various other techniques, such as the traditional snap-turn, thus emphasizing its versatility.
Donghae Lim, Shizuka Shirai, Masato Kobayashi 0001, Yuuki Uranishi, Haruo Takemura
IEEE Trans. Vis. Comput. Graph.5
2024 Solving Multi-Robot Task Allocation and Planning in Trans-media Scenarios
abstract
Trans-media robots, capable of operating across diverse environments, add significant complexity for multi-robot task allocation and planning problems. This paper introduces a novel approach to plan missions for such multi-robot systems, that addresses the associated specific complexities and constraints. It streamlines the overall mission planning process by decomposing it into tractable sub-problems, and addresses the issues of coalition formation, path planning, and task scheduling. It provides mission plans in very little computation time and allows to tackle large missions intractable by global planners, with negligible loss in plan optimality.
Virgile De La Rochefoucauld, Simon Lacroix, Photchara Ratsamee, Haruo Takemura
IROS4
2024 Panoptic-Level Image-to-Image Translation for Object Recognition and Visual Odometry Enhancement
abstract
Image-to-image translation methods have progressed from only considering the image-level information to integrating the global- and instance-level information. However, only the foreground instances are refined, and the background semantics are taken as an entire feature, which causes a substantial loss of the semantic information in the translation. Additionally, the insufficient quality of the translated semantic regions also leads to an unsatisfactory performance of the object recognition or visual odometry tasks in which the translated images/videos are further used. In this paper, we propose a novel generative adversarial network for panoptic-level image-to-image translation (PanopticGAN). The proposed method has three advantages: 1) the extracted panoptic perception (i.e., the foreground instances and background semantic regions) as content codes are aligned with the sampled panoptic style codes, which considers the panoptic-level information to avoid the semantic information loss, and the latent space of each object has a rich fusion of content and style codes to generate the higher-fidelity results; 2) a feature masking module is proposed to extract the representations within each object contour by masks for sharpening the object boundaries; 3) the improved fidelity of the translated semantic regions further contributes to enhancing the performance of the object recognition or visual odometry tasks that the translated images/videos are used in. In this paper, we also annotate a compact panoptic segmentation dataset for the thermal-to-color translation task. Extensive experiments are conducted to demonstrate the effectiveness of our PanopticGAN over the latest methods.
Photchara Ratsamee, Zhaojie Luo, Yuuki Uranishi, Manabu Higashida, Haruo Takemura
IEEE Trans. Circuits Syst. Video Technol.6
2023 Panoptic-aware Image-to-Image Translation
abstract
Despite remarkable progress in image translation, the complex scene with multiple discrepant objects remains a challenging problem. The translated images have low fidelity and tiny objects in fewer details causing unsatisfactory performance in object recognition. Without thorough object perception (i.e., bounding boxes, categories, and masks) of images as prior knowledge, the style transformation of each object will be difficult to track in translation. We propose panoptic-aware generative adversarial networks (PanopticGAN) for image-to-image translation together with a compact panoptic segmentation dataset. The panoptic perception (i.e., foreground instances and background semantics of the image scene) is extracted to achieve alignment between object content codes of the input domain and panoptic-level style codes sampled from the target style space, then refined by a proposed feature masking module for sharping object boundaries. The image-level combination between content and sampled style codes is also merged for higher fidelity image generation. Our proposed method was systematically compared with different competing methods and obtained significant improvement in both image quality and object recognition performance.
Photchara Ratsamee, Bowen Wang 0002, Zhaojie Luo, Yuuki Uranishi, Manabu Higashida, Haruo Takemura
WACV7
2023 Multi-modal humor segment prediction in video
abstract
Abstract Humor can be induced by various signals in the visual, linguistic, and vocal modalities emitted by humans. Finding humor in videos is an interesting but challenging task for an intelligent system. Previous methods predict humor in the sentence level given some text (e.g., speech transcript), sometimes together with other modalities, such as videos and speech. Such methods ignore humor caused by the visual modality in their design, since their prediction is made for a sentence. In this work, we first give new annotations to humor based on a sitcom by setting up temporal segments of ground truth humor derived from the laughter track. Then, we propose a method to find these temporal segments of humor. We adopt an approach based on sliding window, where the visual modality is described by pose and facial features along with the linguistic modality given as subtitles in each sliding window. We use long short-term memory networks to encode the temporal dependency in poses and facial features and pre-trained BERT to handle subtitles. Experimental results show that our method improves the performance of humor prediction.
Yuta Nakashima, Haruo Takemura
Multim. Syst.3
2021 Transferring Domain-Agnostic Knowledge in Video Question Answering
Tianran Wu, Noa Garcia, Mayu Otani, Chenhui Chu, Yuta Nakashima, Haruo Takemura
BMVC6
2021 UAV Target-Selection: 3D Pointing Interface System for Large-Scale Environment
abstract
This paper presents a 3D pointing interface application to signal a UAV’s target in a large-scale environment. This system enables UAVs equipped with a monocular camera to determine which window of a building is selected by a human user in large-scale indoor or outdoor environments. The 3D pointing interface consists of three parts: YOLO, Open- Pose, and ORB-SLAM. YOLO detects the target objects, e.g., windows, OpenPose extracts the user pose, and ORB-SLAM builds a scale-dependent 3D map, a set of 3D sparse feature points. To obtain the visual scale, it performs a calibration step with the user standing in front of the UAV at a certain distance. We detail how we chose the gesture, localize and detect objects, and transform between coordinate systems. The real- world experiment results showed that the 3D pointing interface obtained a 0.73 F1-score average and a 0.58 F1-Score at the maximum distance of 25 meters between UAV and building.
Anna Medeiros, Photchara Ratsamee, Jason Orlosky, Yuuki Uranishi, Manabu Higashida, Haruo Takemura
ICRA6
2021 Spherical Magnetic Joint for Inverted Locomotion of Multi-Legged Robot
abstract
In this paper, we present a spherical magnetic joint for the inverted locomotion of a multi-legged robot. The permanent magnet’s spherical shape allows the robot to attach its foot to a steel surface without energy consumption. However, the robot’s inverted locomotion requires foot flexibility for placement and gait construction of the robot. Therefore, the spherical magnetic joint mechanism was designed and implemented for the robot feet to deal with angular placement. For decoupling the foot from the steel surface, the attractive force is adjusted by tilting the adjustable sleeve mechanism at an adequate angle between the surface and foot tip. Experimental results show that the spherical magnetic joint can maintain the attractive force at any angle, and the sleeve mechanism can reduce 20% of the reaction force for pulling the legs from the steel surfaces. Furthermore, the designed gait for inverted locomotion with a spherical magnetic joint was tested and compared to prove the concept of the spherical magnetic joint and sleeve mechanism.
Harn Sison, Photchara Ratsamee, Manabu Higashida, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura
ICRA6
2021 A Case Study of Redesigning an Introductory CS Course into Fully Online and its Evaluation
abstract
This poster introduces a case study of redesigning an introductory CS course into fully online and its evaluation. The course redesigned is an introductory computer science course aiming to enhance non-CS major students' engagement. We applied flipped learning and active learning approach to the course design, utilizing interactive learning platforms. The proposed course offered a more immersive and satisfactory experience, even in a fully online course. The questionnaire results after the practice in the 2020 term showed that students rated significantly higher than that of 2019 in terms of difficulty level, amount of content, manner of delivery, course materials, consideration of course design, the structure according to the syllabus, learning outcomes, and satisfaction.
Shizuka Shirai, Hiroyuki Nagataki, Tomohiro Nishida, Haruo Takemura
SIGCSE4
2021 The Laughing Machine: Predicting Humor in Video
abstract
Humor is a very important communication tool; yet, it is an open problem for machines to understand humor. In this paper, we build a new multimodal dataset for humor prediction that includes subtitles and video frames, as well as humor labels associated with video's timestamps. On top of it, we present a model to predict whether a subtitle causes laughter. Our model uses the visual modality through facial expression and character name recognition, together with the verbal modality, to explore how the visual modality helps. In addition, we use an attention mechanism to adjust the weight for each modality to facilitate humor prediction. Interestingly, our experimental results show that the performance boost by combinations of different modalities, and the attention mechanism and the model mostly relies on the verbal modality.
Yuta Kayatani, Mayu Otani, Noa Garcia, Chenhui Chu, Yuta Nakashima, Haruo Takemura
WACV7
2021 A comparative study of language transformers for video question answering
Noa Garcia, Chenhui Chu, Mayu Otani, Yuta Nakashima, Haruo Takemura
Neurocomputing6
2020 BERT Representations for Video Question Answering
abstract
Visual question answering (VQA) aims at answering questions about the visual content of an image or a video. Currently, most work on VQA is focused on image-based question answering, and less attention has been paid into answering questions about videos. However, VQA in video presents some unique challenges that are worth studying: it not only requires to model a sequence of visual features over time, but often it also needs to reason about associated subtitles. In this work, we propose to use BERT, a sequential modelling technique based on Transformers, to encode the complex semantics from video clips. Our proposed model jointly captures the visual and language information of a video scene by encoding not only the subtitles but also a sequence of visual concepts with a pretrained language-based Transformer. In our experiments, we exhaustively study the performance of our model by taking different input arrangements, showing outstanding improvements when compared against previous work on two well-known video VQA datasets: TVQA and Pororo.
Noa Garcia, Chenhui Chu, Mayu Otani, Yuta Nakashima, Haruo Takemura
WACV6
2019 Evaluation of Pointing Interfaces with an AR Agent for Multi-section Information Guidance
abstract
In educational settings such as art galleries or museums, Augmented Reality (AR) has the potential to provide detailed information about exhibits. However, dealing with items that contain information in multiple sections or areas is still a significant challenge. For example, a large painting may contain many minute details, which requires a system that can explain its broader features rather than just a generic description. To address this challenge, we introduce an AR guidance system that uses an embodied agent to point out items and explain each piece and part of exhibit items in detail. We also designed and tested 3 different pointing interfaces for the embodied agent: gesture only, gesture with a dot laser, and gesture with line laser. To evaluate this interface, we conducted a user experiment simulating painting guidance to test interest and exhibit memory. During the experiment, the agent pointed to various areas of interest in the painting and provided a detailed description to participants. The result shows that the search times for target positions were the fastest with the line laser. However, no particular interface outperformed others in memory recall of exhibit content.
Nattaon Techasarntikul, Tomohiro Mashita, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura, Jason Orlosky, Kiyoshi Kiyokawa
VR5
2019 A Comparison of Adaptive View Techniques for Exploratory 3D Drone Teleoperation
abstract
Drone navigation in complex environments poses many problems to teleoperators. Especially in three dimensional (3D) structures such as buildings or tunnels, viewpoints are often limited to the drone’s current camera view, nearby objects can be collision hazards, and frequent occlusion can hinder accurate manipulation. To address these issues, we have developed a novel interface for teleoperation that provides a user with environment-adaptive viewpoints that are automatically configured to improve safety and provide smooth operation. This real-time adaptive viewpoint system takes robot position, orientation, and 3D point-cloud information into account to modify the user’s viewpoint to maximize visibility. Our prototype uses simultaneous localization and mapping (SLAM) based reconstruction with an omnidirectional camera, and we use the resulting models as well as simulations in a series of preliminary experiments testing navigation of various structures. Results suggest that automatic viewpoint generation can outperform first- and third-person view interfaces for virtual teleoperators in terms of ease of control and accuracy of robot operation.
John Thomason, Photchara Ratsamee, Jason Orlosky, Kiyoshi Kiyokawa, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura
ACM Trans. Interact. Intell. Syst.7
2018 Force Rendering and its Evaluation of a Friction-Based Walking Sensation Display for a Seated User
abstract
Most existing locomotion devices that represent the sensation of walking target a user who is actually performing a walking motion. Here, we attempted to represent the walking sensation, especially a kinesthetic sensation and advancing feeling (the sense of moving forward) while the user remains seated. To represent the walking sensation using a relatively simple device, we focused on the force rendering and its evaluation of the longitudinal friction force applied on the sole during walking. Based on the measurement of the friction force applied on the sole during actual walking, we developed a novel friction force display that can present the friction force without the influence of body weight. Using performance evaluation testing, we found that the proposed method can stably and rapidly display friction force. Also, we developed a virtual reality (VR) walk-through system that is able to present the friction force through the proposed device according to the avatar's walking motion in a virtual world. By evaluating the realism, we found that the proposed device can represent a more realistic advancing feeling than vibration feedback.
Ginga Kato, Yoshihiro Kuroda, Kiyoshi Kiyokawa, Haruo Takemura
IEEE Trans. Vis. Comput. Graph.4
2017 Exploring Proxemics for Human-Drone Interaction
abstract
We present a human-centered designed social drone aiming to be used in a human crowd environment. Based on design studies and focus groups, we created a prototype of a social drone with a social shape, face and voice for human interaction. We used the prototype for a proxemic study, comparing the required distance from the drone humans could comfortably accept compared with what they would require for a nonsocial drone. The social shaped design with greeting voice added decreased the acceptable distance markedly, as did present or previous pet ownership, and maleness. We also explored the proximity sphere around humans with a social shaped drone based on a validation study with variation of lateral distance and heights. Both lateral distance and the higher height of 1.8 m compared to the lower height of 1.2 m decreased the required comfortable distance as it approached.
Alexander Yeh, Photchara Ratsamee, Kiyoshi Kiyokawa, Yuuki Uranishi, Tomohiro Mashita, Haruo Takemura, Morten Fjeld, Mohammad Obaid
HAI6
2017 VisMerge: Light Adaptive Vision Augmentation via Spectral and Temporal Fusion of Non-visible Light
abstract
Low light situations pose a significant challenge to individuals working in a variety of different fields such as firefighting, rescue, maintenance and medicine. Tools like flashlights and infrared (IR) cameras have been used to augment light in the past, but they must often be operated manually, provide a field of view that is decoupled from the operator's own view, and utilize color schemes that can occlude content from the original scene. To help address these issues, we present VisMerge, a framework that combines a thermal imaging head mounted display (HMD) and algorithms that temporally and spectrally merge video streams of different light bands into the same field of view. For temporal synchronization, we first develop a variant of the time warping algorithm used in virtual reality (VR), but redesign it to merge video see-through (VST) cameras with different latencies. Next, using computer vision and image compositing we develop five new algorithms designed to merge non-uniform video streams from a standard RGB camera and small form-factor infrared (IR) camera. We then implement six other existing fusion methods, and conduct a series of comparative experiments, including a system level analysis of the augmented reality (AR) time warping algorithm, a pilot experiment to test perceptual consistency across all eleven merging algorithms, and an in-depth experiment on performance testing the top algorithms in a VR (simulated AR) search task. Results showed that we can reduce temporal registration error due to inter-camera latency by an average of 87.04%, that the wavelet and inverse stipple algorithms were perceptually rated the highest, that noise modulation performed best, and that freedom of user movement is significantly increased with visualizations engaged.
Jason Orlosky, Peter Kim, Kiyoshi Kiyokawa, Tomohiro Mashita, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura
ISMAR7
2017 Adaptive View Management for Drone Teleoperation in Complex 3D Structures
abstract
Drone navigation in complex environments poses many problems to teleoperators. Especially in 3D structures like buildings or tunnels, viewpoints are often limited to the drone's current camera view, nearby objects can be collision hazards, and frequent occlusion can hinder accurate manipulation. To address these issues, we have developed a novel interface for teleoperation that provides a user with environment-adaptive viewpoints that are automatically configured to improve safety and smooth user operation. This real-time adaptive viewpoint system takes robot position, orientation, and 3D pointcloud information into account to modify user-viewpoint to maximize visibility. Our prototype uses simultaneous localization and mapping (SLAM) based reconstruction with an omnidirectional camera and we use resulting models as well as simulations in a series of preliminary experiments testing navigation of various structures. Results suggest that automatic viewpoint generation can outperform first and third-person view interfaces for virtual teleoperators in terms of ease of control and accuracy of robot operation.
John Thomason, Photchara Ratsamee, Kiyoshi Kiyokawa, Pakpoom Kriengkomol, Jason Orlosky, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura
IUI8
2016 A bone marrow cavity segmentation method using wavelet-based texture feature
abstract
A better understanding of in vivo bio images is expected to contribute to the discovery of new drugs and mechanisms of disease. To improve the contributions of in vivo bioimaging, the extraction of a particular region is required in order to detect a particular cell's motion because manual image processing of a massive number of images is unrealistic. One of the issues for automatic image-segmentation is that conditions of image-taking are variable. Thus, some manual input and/or manual tuning of some parameters is required to adjust each image. To reduce manual operation for image processing of bone marrow cavity segmentation, we focused on the texture pattern of bone marrow cavity. In this paper, we propose a bone marrow cavity segmentation method using support vector machine and wavelet-based texture feature. The proposed method does not require manual inputs to obtain distribution of intensity before processing, because the texture patterns of bone marrow cavity regions are integrated into the system in advance. Moreover, it is applicable to a particular frame in an image sequence in which the condition of fluorescent material is variable because it does not require temporal variation or initial frame for the segmentation. In the experiment, we evaluated our method with nine types of mother wavelets and several sets of scale parameters. The bone marrow cavity segmentation, using graph-cuts with our texture pattern classification, performs well without manual inputs by a user.
Hironori Shigeta, Tomohiro Mashita, Junichi Kikuta, Shigeto Seno, Haruo Takemura, Hideo Matsuda, Masaru Ishii
ICPR5
2016 Spatial consistency perception in optical and video see-through head-mounted augmentations
abstract
Correct spatial alignment is an essential requirement for convincing augmented reality experiences. Registration error, caused by a variety of systematic, environmental, and user influences decreases the realism and utility of head mounted display AR applications. Focus is often given to rigorous calibration and prediction methods seeking to entirely remove misalignment error between virtual and real content. Unfortunately, producing perfect registration is often simply not possible. Our goal is to quantify the sensitivity of users to registration error in these systems, and identify acceptability thresholds at which users can no longer distinguish between the spatial positioning of virtual and real objects. We simulate both video see-through and optical see-through environments using a projector system and experimentally measure user perception of virtual content misalignment. Our results indicate that users are less perceptive to rotational errors over all and that translational accuracy is less important in optical see-through systems than in video see-through.
Alexander Plopski, Kenneth R. Moser, Kiyoshi Kiyokawa, J. Edward Swan II, Haruo Takemura
VR5
2016 SAGE-based Tiled Display Wall enhanced with dynamic routing functionality triggered by user interaction
Yoshiyuki Kido, Kohei Ichikawa, Susumu Date, Yasuhiro Watashiba, Hirotake Abe, Hiroaki Yamanaka, Eiji Kawai, Haruo Takemura, Shinji Shimojo
Future Gener. Comput. Syst.8
2015 HapSticks: A novel method to present vertical forces in tool-mediated interactions by a non-grounded rotation mechanism
abstract
Force feedback in tool-mediated interactions with the environment is important for successful performance of complex tasks in our daily life as well as in specialized fields like medicine. Stylus-based haptic devices are studied and used extensively, and most of these devices require either grounding or attachment to the body of the user. Recently, non-grounded haptic devices are getting an increasing attention. In this paper, we propose a novel method to represent the vertical forces that are applied on the tip of a tool: a non-grounded rotation mechanism that mimics the cutaneous sensation that is caused by these tool-tip forces. To evaluate this method, we developed a novel ungrounded haptic device - HapSticks - that renders the sensation of manipulating objects using chopsticks. First, we present the novel mechanism, and test the pressure that it applies on the hand of the user when rendering a force at the tip of the tool in comparison to applying a real force at the tip of the tool. Next, we used the mechanism to build the HapSticks device as an example of an application of the proposed method, and present a psychophysical evaluation of this device in a virtual weight discrimination task.
Ginga Kato, Yoshihiro Kuroda, Ilana Nisky, Kiyoshi Kiyokawa, Haruo Takemura
World Haptics5
2015 Corneal-Imaging Calibration for Optical See-Through Head-Mounted Displays
abstract
In recent years optical see-through head-mounted displays (OST-HMDs) have moved from conceptual research to a market of mass-produced devices with new models and applications being released continuously. It remains challenging to deploy augmented reality (AR) applications that require consistent spatial visualization. Examples include maintenance, training and medical tasks, as the view of the attached scene camera is shifted from the user's view. A calibration step can compute the relationship between the HMD-screen and the user's eye to align the digital content. However, this alignment is only viable as long as the display does not move, an assumption that rarely holds for an extended period of time. As a consequence, continuous recalibration is necessary. Manual calibration methods are tedious and rarely support practical applications. Existing automated methods do not account for user-specific parameters and are error prone. We propose the combination of a pre-calibrated display with a per-frame estimation of the user's cornea position to estimate the individual eye center and continuously recalibrate the system. With this, we also obtain the gaze direction, which allows for instantaneous uncalibrated eye gaze tracking, without the need for additional hardware and complex illumination. Contrary to existing methods, we use simple image processing and do not rely on iris tracking, which is typically noisy and can be ambiguous. Evaluation with simulated and real data shows that our approach achieves a more accurate and stable eye pose estimation, which results in an improved and practical calibration with a largely improved distribution of projection error.
Alexander Plopski, Yuta Itoh 0001, Christian Nitschke, Kiyoshi Kiyokawa, Gudrun Klinker, Haruo Takemura
IEEE Trans. Vis. Comput. Graph.6
2014 Performance Characteristics of an SDN-Enhanced Job Management System for Cluster Systems with Fat-Tree Interconnect
abstract
In the era of cloud computing, data centers that accommodate a series of user-requested jobs with a diversity of resource usage pattern need to have the capability of efficiently distributing resources to each user job, based on individual resource usage patterns. In particular, for high-performance computing as a cloud service which allows many users to benefit from a large-scale computing system, a new framework for resource management that treats not only the CPU resources, but also the network resources in the data center is essential. In this paper, an SDN-enhanced JMS that efficiently handles both network and CPU resources and as a result accelerates the execution time of user jobs is introduced as a building block technology for such a HPC cloud. Our evaluation shows that the SDN-enhanced JMS efficiently leverages the fat-tree interconnect of cluster systems running behind the cloud to suppress the collision of communications generated by different jobs.
Yasuhiro Watashiba, Susumu Date, Hirotake Abe, Yoshiyuki Kido, Kohei Ichikawa, Hiroaki Yamanaka, Eiji Kawai, Shinji Shimojo, Haruo Takemura
CloudCom9
2014 Analysing the effects of a wide field of view augmented reality display on search performance in divided attention tasks
abstract
A wide field of view augmented reality display is a special type of head-worn device that enables users to view augmentations in the peripheral visual field. However, the actual effects of a wide field of view display on the perception of augmentations have not been widely studied. To improve our understanding of this type of display when conducting divided attention search tasks, we conducted an in depth experiment testing two view management methods, in-view and in-situ labelling. With in-view labelling, search target annotations appear on the display border with a corresponding leader line, whereas in-situ annotations appear without a leader line, as if they are affixed to the referenced objects in the environment. Results show that target discovery rates consistently drop with in-view labelling and increase with in-situ labelling as display angle approaches 100 degrees of field of view. Past this point, the performances of the two view management methods begin to converge, suggesting equivalent discovery rates at approximately 130 degrees of field of view. Results also indicate that users exhibited lower discovery rates for targets appearing in peripheral vision, and that there is little impact of field of view on response time and mental workload.
Naohiro Kishishita, Kiyoshi Kiyokawa, Jason Orlosky, Tomohiro Mashita, Haruo Takemura, Ernst Kruijff
ISMAR5
2014 Corneal imaging in localization and HMD interaction
abstract
The human eyes perceive our surroundings and are one of, if not our most important sensory organs. Contrary to our other senses the eyes not only perceive but also provide information to a keen observer. However, thus far this has been mainly used to detect reflection of infrared light sources to estimate the user's gaze. The reflection of the visible spectrum on the other hand has rarely been utilized. In this dissertation we want to explore how the analysis of the corneal image can improve currently available eye-related solutions, such as calibration of optical see-through head-mounted devices or eye-gaze tracking and point of regard estimation in arbitrary environments. We also aim to study how corneal imaging can become an alternative for established augmented reality tasks such as tracking and localization.
Alexander Plopski, Kiyoshi Kiyokawa, Haruo Takemura, Christian Nitschke
ISMAR3
2014 The effectiveness of an AR-based context-aware assembly support system in object assembly
abstract
This study evaluates the effectiveness of an AR-based context-aware assembly support system with AR visualization modes proposed in object assembly. Although many AR-based assembly support systems have been proposed, few keep track of the assembly status in real-time and automatically recognize error and completion states at each step. Naturally, the effectiveness of such context-aware systems remains unexplored. Our test-bed system displays guidance information and error detection information corresponding to the recognized assembly status in the context of building block (LEGO) assembly. A user wearing a head mounted display (HMD) can intuitively build a building block structure on a table by visually confirming correct and incorrect blocks and locating where to attach new blocks. We proposed two AR visualization modes, one of them that displays guidance information directly overlaid on the physical model, and another one in which guidance information is rendered on a virtual model adjacent to the real model. An evaluation was conducted to comparatively evaluate these AR visualization modes as well as determine the effectiveness of context-aware error detection. Our experimental results indicate the visualization mode that shows target status next to real objects of concern outperforms the traditional direct overlay under moderate registration accuracy and marker-based tracking.
Bui Minh Khuong, Kiyoshi Kiyokawa, Joseph J. La Viola, Tomohiro Mashita, Haruo Takemura
VR6
2014 Reflectance and light source estimation for indoor AR Applications
abstract
We present an approach which enables real-time augmentation of an environment composed of materials with different texture and reflectance properties without the need of application-specific hardware or extensive preparation. Our solution uses a set of RGB images of a reconstructed model to optimize the reflectance parameters and light location. Each image is decomposed into its specular and diffuse components and we estimate the location of multiple light sources from specular highlights. The environment is stored in a voxel grid and we optimize the reflectance properties and colour of each voxel through inverse rendering. We verify our approach with a simulated environment and present results from a corresponding reconstructed environment.
Alexander Plopski, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
VR4
2013 OpenFlow Network Visualization Software with Flow Control Interface
abstract
Recently, the concept of Software-Defined Network (SDN), which allows us to administer and configure a network in a centralized and software-programming manner, has gathered network engineers' and researchers' attention rapidly. In particular, the expectation and concern to OpenFlow as an implementation of the SDN is remarkable. As a result, research activities, which include prototyping, implementation, demonstration and experiments, conducted over OpenFlow networks have been a worldwide tendency. In such research activities, however, the difficulty in understanding network topology, traffic amount and an actual path of a network flow on the OpenFlow network, and the intricacies in debugging software designed for OpenFlow are serious problems in the development process of OpenFlow controller. This research aims to realize a visualization software that facilitates researchers to perform OpenFlow controller development and demonstration experiments performed on an actual OpenFlow network. In this paper, the authors summarize the achievement of their research work in progress as well as the future direction.
Yasuhiro Watashiba, Seiichiro Hirabara, Susumu Date, Hirotake Abe, Kohei Ichikawa, Yoshiyuki Kido, Shinji Shimojo, Haruo Takemura
COMPSAC8
2013 In-situ lighting and reflectance estimations for indoor AR systems
abstract
We introduce an in-situ lighting and reflectance estimation method that does not require specific light probes and/or preliminary scanning. Our method uses images taken from multiple viewpoints while data accumulation and lighting and reflectance estimations run in the background of the primary AR system. As a result, our method requires little in the way of manipulations for image collection because it consists primarily of image processing and optimization. When used, lighting directions and initial optimization values are estimated via image processing. Eventually, the full parameters are obtained by optimization of the differences between real images. This system uses current best parameters because the parameter estimation and input image updates are run independently.
Tomohiro Mashita, Hiroyuki Yasuhara, Alexander Plopski, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR5
2013 Towards intelligent view management: A study of manual text placement tendencies in mobile environments using video see-through displays
abstract
When viewing content in a see-through head mounted display (HMD), displaying readable information is still difficult when text is overlayed onto a changing background or lighted surface. Moving text or content to a more appropriate place on the screen through automation or intelligent algorithms is one viable solution to this kind of issue. However, many of these algorithms fail to act as a human would when placing text in a more appropriate location in real time. In order to improve these text and view management algorithms, we report the results and analysis of an experiment designed to evaluate user tendencies when placing virtual text in the real world through an HMD. In the conducted experiment, 20 users manually overlayed text in real time onto 4 different videos taken from the first-person perspective of a pedestrian. We find that users have a tendency to place overlayed text in locations near the center of the viewing field, gravitating towards a point just below the horizon. Common locations for text overlay such as walls, shaded areas, and pavement are classified and discussed.
Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR3
2013 Management and manipulation of text in dynamic mixed reality workspaces
abstract
Viewing and interacting with text based content safely and easily while mobile has been an issue with see-through displays for many years. For example, in order to effectively use optical see through Head Mounted Displays (HMDs) in constantly changing dynamic environments, variables like lighting conditions, human or vehicular obstructions in a user's path, and scene variation must be dealt with effectively. My PhD research focuses on answering the following questions: 1) What are appropriate methods to intelligently move digital content such as e-mail, SMS messeges, and news articles, throughout the real world? 2) Once a user stops moving, in what way should dynamics of the current workspace change when migrated to a new static environment? 3) Lastly, how can users manipulate mobile content using the fewest number of interactions possible? My strategy for developing solutions to these problems primarily involves automatic or semi-automatic movement of digital content throughout the real world using camera tracking. I have already developed an intelligent text management system that actively manages movement of text in a user's field of view while mobile [11]. I am optimizing and expanding on this type of management system, developing appropriate interaction methodology, and conducting experiments to verify effectiveness, usability, and safety when used with an HMD in various environments.
Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR3
2013 Dynamic text management for see-through wearable and heads-up display systems
abstract
Reading text safely and easily while mobile has been an issue with see-through displays for many years. For example, in order to effectively use optical see through Head Mounted Displays (HMDs) or Heads Up Display (HUD) systems in constantly changing dynamic environments, variables like lighting conditions, human or vehicular obstructions in a user's path, and scene variation must be dealt with effectively.
Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura
IUI3
2013 A next location prediction method for smartphones using blockmodels
abstract
Context aware systems on smart-phones aim to provide useful information by analysing and recognizing users' situations from built-in sensors logs. Especially, predicting user actions is one of the important functions for the context aware systems on smart-phones because this function enables context aware systems to provide proactive and responsive services. Therefore the next location prediction is also an important function for context aware systems. This paper introduces a next location prediction method based on context recognition. In this method, we define a context as combinations of features which are extracted from a set of relational data generated from a phone's sensor logs. We applied Mixed Membership Stochastic Blockmodels (MMSB) to context extraction. We then collected sensor logs of a single user over a period of three months and conducted an evaluation using this collected dataset. An an evaluation using the dataset was conducted and the result shows that 60% of the test dataset ranked in the top 30% of all candidates of the next locations.
Jun Fukano, Tomohiro Mashita, Takahiro Hara, Kiyoshi Kiyokawa, Haruo Takemura, Shojiro Nishio
VR5
2013 Pinch-n-Paste: Direct texture transfer interaction in augmented reality
abstract
Our Pinch-n-Paste allows a user to touch or pinch one part of an object, copy and move its texture, and paste it onto another object, directly with his or her hand, in an augmented reality environment. To transfer texture appropriately from one part of an object to another, two texture images are generated by the Least Square Conformal Map (LSCM) technique. Two regions in the texture images corresponding to source and target areas of interest are then obtained using cross-boundary brushes. Target texel values are sampled from corresponding source texels by Moving Least Squares (MLS), and are finally mapped onto the target object. In this poster, we will describe the basic idea, implementation details, and example interaction results and a preliminary user study.
Atsushi Umakatsu, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
VR4
2013 A content search system considering the activity and context of a mobile user
Mayu Iwata, Hiroki Miyamoto, Takahiro Hara, Daijiro Komaki, Kentaro Shimatani, Tomohiro Mashita, Kiyoshi Kiyokawa, Toshiaki Uemukai, Gen Hattori, Shojiro Nishio, Haruo Takemura
Pers. Ubiquitous Comput.11
2012 A waist-mounted ProCam system for remote collaboration
abstract
We propose a waist-mounted projector-camera (ProCam) system for asymmetric remote collaboration. A wearable camera is often used to transmit a worker's situation to a remote instructor, however 3D structure of the worker's environment is not always available and the instructor has a minimal flexibility in changing the camera's viewpoint. A stationary 3D measurement system is also commonly used for remote collaboration, however a narrow measurement area and occlusion from a worker's body can be a severe problem. Our waist-mounted ProCam system reconstructs worker's environment in real-time without occlusion from the worker's body. The remote instructor can give instructions simply by drawing annotations on the reconstructed environment on screen, and they are properly projected in front of the worker. Structured-light based reconstruction, vision-based localization, and visual annotation projection, are processed synchronously with a camera and modified shutter glasses so that both the camera and the worker observe only the information they need. Experimental results show that users prefer our system to a stationary ProCam system.
Shigeki Morishima, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR4
2012 Touch-n-Paste: Direct texture transfer interaction in AR environments
abstract
Our Touch-n-Paste allows a user to touch one part of an object, copy and move its texture, and paste it onto another object, directly with his or her hand, in an augmented reality environment. To transfer texture appropriately from one part of an object to another, two texture images are generated by the Least Square Conformal Map (LSCM) technique. Two regions in the texture images corresponding to source and target areas of interest are then obtained using cross-boundary brushes. Target texel values are sampled from corresponding source texels by Moving Least Squares (MLS), and are finally mapped onto the target object.
Atsushi Umakatsu, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR4
2012 Subjective evaluations on perceptual depth of stereo image and effective field of view of a wide-view head mounted projective display with a semi-transparent retro-reflective screen
abstract
We report two user studies on a wearable hyperboloidal head mounted projective display (HHMPD) with a semi-transparent retro-reflective screen. First experiment revealed that a virtual image is perceived at a similar distance as the real image only when the observation distance is within 2.5m with monocular vision, whereas its threshold is further than 3m with stereo (binocular) vision. Second experiment revealed that users are able to identify visual stimuli in the periphery of the visual field up to ±50 degrees in horizontal, while paying attention to a real object in frontal direction.
Duc Nguyen Van, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR4
2012 Human-computer dance interaction with realtime accelerometer control
abstract
Motion-capture-based character animations are widely used in computer graphics and interactive games.In this paper, we show a novel approach to create dancing character animations that react to input music and an accelerometer manipulated by a user. Since the sensor reads express intensities of users' body movements, the system can synthesize character motions whose intensities are synchronized to those of users. Our system consists of analysis phase and synthesis phase. In the analysis phase, the musical beat and segments are detected from input sound, and motion rhythm and intensities are found from motion capture data. With the results of this analysis, we generate a motion graph that can generate character motions matched to the musical rhythm. In synthesis phase, the system receives the output data from an accelerometer and traverses the motion graph according to the matching result between the sensor data and the motion intensity. As the result, our system adds an interactive component to live dancing performed by virtual characters.
Takuya Yasunaga, Atsushi Nakazawa, Haruo Takemura
ACM Multimedia3
2012 A content search system for mobile devices based on user context recognition
abstract
People carry around mobile devices all the time in their daily life and get various information from the Internet in various situations. When searching for information (content) by using mobile devices, users' activities (e.g., walking and standing) and their situations (e.g., commuting in the morning and going out downtown in the evening) often change and this change may affect their degree of concentration on the display of mobile devices and their information needs. Therefore, search systems should provide users with an amount of information suitable for their activities and with a type of information suitable for their situations. In this paper, we present the design and implementation of a content search system considering mobile users' activities and situations, which aims to reduce users' load of operations in content searching. Our system recognizes user's activities and switches between two kinds of content search systems according to the user's activity: the location-based content search system runs when the user is standing, while the menu-based content search system runs when the user is walking. Both systems present information based on the user's situation. We also introduce a user's activity recognition method for mobile devices. This method classifies user's activity into standing, walking, and running using the sensors equipped in mobile devices.
Tomohiro Mashita, Daijiro Komaki, Mayu Iwata, Kentaro Shimatani, Hiroki Miyamoto, Takahiro Hara, Kiyoshi Kiyokawa, Haruo Takemura, Shojiro Nishio
VR8
2012 Human activity recognition for a content search system considering situations of smartphone users
abstract
Smart-phone users can search for information about surrounding facilities or a route to their destination. However, it is difficult to get or search for information while walking because of low legibility. To address this problem, users have to stop walking or enlarge the screen. Our previously proposed system for smart-phone switches the information presentation policies in response to the user's context. In this paper we describe our context recognition mechanism for this system. This mechanism estimates user context from sensors embedded in a smart-phone. We use a Support Vector Machine for the context classification and compare four types of feature values consisting of FFT and 3 types of Wavelet Transforms. Experimental results show that recognition rates are 87.2 % with FFT, 90.9 % with Gabor Wavelet, 91.8 % with Haar Wavelet, and 92.1 % with MexicanHat Wavelet.
Tomohiro Mashita, Kentaro Shimatani, Mayu Iwata, Hiroki Miyamoto, Daijiro Komaki, Takahiro Hara, Kiyoshi Kiyokawa, Haruo Takemura, Shojiro Nishio
VR8
2012 Pinch-n-paste: direct texture transfer interaction in augmented reality
abstract
Our Pinch-n-Paste allows a user to touch or pinch one part of an object, copy and move its texture, and paste it onto an other object, directly with his or her hand, in an augmented reality environment. To transfer texture appropriately from one part of an object to another, two texture images are generated by the Least Square Conformal Map (LSCM) technique. Two regions in the texture images corresponding to source and target areas of interest are then obtained using cross-boundary brushes. Target texel values are sampled from corresponding source texels by Moving Least Squares (MLS), and are finally mapped onto the target object.
Atsushi Umakatsu, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
VRST4
2011 Display-camera calibration using eye reflections and geometry constraints
Christian Nitschke, Atsushi Nakazawa, Haruo Takemura
Comput. Vis. Image Underst.3
2011 A Wide-View Parallax-Free Eye-Mark Recorder with a Hyperboloidal Half-Silvered Mirror and Appearance-Based Gaze Estimation
abstract
In this paper, we propose a wide-view parallax-free eye-mark recorder with a hyperboloidal half-silvered mirror and a gaze estimation method suitable for the device. Our eye-mark recorder provides a wide field-of-view video recording of the user's exact view by positioning the focal point of the mirror at the user's viewpoint. The vertical angle of view of the prototype is 122 degree (elevation and depression angles are 38 and 84 degree, respectively) and its horizontal view angle is 116 degree (nasal and temporal view angles are 38 and 78 degree, respectively). We implemented and evaluated a gaze estimation method for our eye-mark recorder. We use an appearance-based approach for our eye-mark recorder to support a wide field-of-view. We apply principal component analysis (PCA) and multiple regression analysis (MRA) to determine the relationship between the captured images and their corresponding gaze points. Experimental results verify that our eye-mark recorder successfully captures a wide field-of-view of a user and estimates gaze direction with an angular accuracy of around 2 to 4 degree.
Hiroki Mori, Erika Sumiya, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
IEEE Trans. Vis. Comput. Graph.5
2009 Display-camera calibration from eye reflections
abstract
We present a novel technique for calibrating display-camera systems from reflections in the user's eyes. Display-camera systems enable a range of vision applications that need controlled illumination, including 3D object reconstruction, facial modeling and human computer interaction. One important issue, though, is the geometric calibration of the display, which requires additional hardware and tedious user interaction. The proposed approach eliminates this requirement by analyzing patterns that are reflected in the cornea, a mirroring device that naturally exists in any display-camera system. We introduce an optimization strategy that is able to refine eye and spherical mirror calibration results. When applied to the eye, it even outperforms spherical mirror calibration unoptimized. Furthermore, we obtain a robust estimation of eye poses which can be used for eye tracking applications. Despite the difficult working conditions, the calibration results are good and should be sufficient for many applications.
Christian Nitschke, Atsushi Nakazawa, Haruo Takemura
ICCV3
2009 Large-scale 3D scene modeling by registration of laser range data with Google Maps images
abstract
This work presents a novel approach to registering multiple range images on top of a Google Maps image. The fundamental concept behind the method is matching completely different types of input with each other using classification as a middleman. Range images and Google Maps images are separated into classes, and the range image is also projected into a 2D top-down template image. The template image can then be matched against the Google Maps image to find its location and orientation on the map, which can be used for registering the range images. An experiment comparing this technique against using GPS to find position and orientation showed that it is effective at automatically constructing a reasonable large-scale 3D model whereas GPS would be completely ineffective.
Anuraag Agrawal, Miki Matsumura, Atsushi Nakazawa, Haruo Takemura
ICIP4
2009 Eye reflection analysis and application to display-camera calibration
abstract
We present a novel technique for calibrating display-camera systems from reflections in the user's eyes. Display-camera systems enable a range of vision applications that need controlled illumination, including 3D object reconstruction, facial modeling and human computer interaction. One important issue, though, is the geometric calibration of the display, which requires additional hardware and tedious user interaction. The proposed approach eliminates this requirement by analyzing patterns that are reflected in the cornea, a mirroring device that naturally exists in any display-camera system. By applying this strategy we also obtain a continuous estimation of eye poses which facilitates further applications. We investigate the effect of display size, camera-eye distance and individual eye anatomy experimentally using only off-the-shelf components. Results are promising and show the general feasibility of the approach.
Christian Nitschke, Atsushi Nakazawa, Haruo Takemura
ICIP3
2009 MMM-classification of 3D range data
abstract
This paper presents a method for accurately segmenting and classifying 3D range data into particular object classes. Object classification of input images is necessary for applications including robot navigation and automation, in particular with respect to path planning. To achieve robust object classification, we propose the idea of an object feature which represents a distribution of neighboring points around a target point. In addition, rather than processing raw points, we reconstruct polygons from the point data, introducing connectivity to the points. With these ideas, we can refine the Markov Random Field (MRF) calculation with more relevant information with regards to determining ldquorelated pointsrdquo. The algorithm was tested against five outdoor scenes and provided accurate classification even in the presence of many classes of interest.
Anuraag Agrawal, Atsushi Nakazawa, Haruo Takemura
ICRA3
2009 A wide-view parallax-free eye-mark recorder with a hyperboloidal half-silvered mirror
abstract
In this paper, we propose a wide-view parallax-free eye-mark recorder with a hyperboloidal half-silvered mirror. Our eye-mark recorder provides a wide field-of-view (FOV) video recording of the user's exact view by positioning the focal point of the mirror at the user's viewpoint. The vertical view angle of the prototype is 122 [deg] (elevation and depression angles are 38 and 84 [deg], respectively) and its horizontal view angle is 116 [deg] (nasal and temporal view angles are 38 and 78 [deg], respectively). We have implemented and evaluated a gaze estimation method for our eyemark recorder. Experimental results have verified that our eye-mark recorder successfully captures a wide FOV of a user and estimates a rough gaze direction.
Erika Sumiya, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
VRST4
2008 Optimized Rendering for a Three-Dimensional Videoconferencing System
abstract
Industry widely employs the two-dimensional videoconferencing system as a long distance communication tool, but current limitations such as its tendency to misrepresent eye contact prevent it from becoming more widely adopted. We are exploring the possibility of a three-dimensional videoconferencing system for future interactive streaming of point cloud data, and present the preliminary research results in this paper. We have tested thus far with one sender and one receiver, using pre-recorded data for the sender. The sender, encircled by high-definition cameras, stands and speaks in a room. A cluster of computers reconstructs each frame of the camera images into a 3D point cloud and streams it across a high-speed, low-latency network. On the receiving end, a splat-based renderer employs a new algorithm to efficiently resample the points in real-time, maintaining a user-specified frame rate. Parallel hardware projects onto multiple screens while head tracking equipment records the viewer's movements, allowing the receiver to view a stereoscopic 3D representation of the sender from multiple angles. We can combine these visuals with appropriate use of multiple audio channels to forge an unparalleled virtual experience. This next step towards immersive 3D videoconferencing brings us closer to empowering worldwide collaboration between research departments.
Rachel Chu, Daniel Tenedorio, Jürgen P. Schulze, Susumu Date, Seiki Kuwabara, Atsushi Nakazawa, Haruo Takemura, Fang-Pang Lin
eScience7
2008 PRIUS: An Educational Framework on PRAGMA Fostering Globally-Leading Researchers in Integrated Sciences
abstract
In 2005, Osaka University, in Japan, started an international educational program called, Pacific Rim International University (PRIUS), on top of the Pacific Rim Application and Grid Middleware Assembly (PRAGMA) research framework. The PRIUS framework is based on and similar to that of the PRIME program at the University of California San Diego. Through the PRIUS program, Osaka University has explored a new structure of higher education for graduate students by combining lectures given by PRAGMA researchers and scientists as well as internship abroad opportunities to PRAGMA member institutions and universities. In this paper, we describe the goals and framework of the PRIUS program and discuss issues for the improvement of PRIUS. We also present two examples of interns' achievements as well as other educational effects brought through collaboration with PRAGMA.
Susumu Date, Shoji Miyanaga, Kohei Ichikawa, Shinji Shimojo, Haruo Takemura, Toru Fujiwara
eScience5
2008 Mutual occlusions on table-top displays in mixed reality applications
abstract
This paper describes an approach to dealing with mutual occlusions between virtual and real objects on a table-top display. Display tables use stereoscopy to make virtual content appear to exist in 3 dimensions on or above a table top. The actual image, however, lies on the physical plane of the display table. Any real physical object introduced above this plane therefore obstructs our view of the display surface and disrupts the illusion of the virtual scene. The occlusions result between real objects and the display surface, not between real objects and virtual objects. For the same reason virtual objects cannot occlude real ones. Our approach uses an additional projector located near the user's head to project those parts of virtual objects that should occlude real ones directly onto the real objects. We describe possible applications and limitations of the approach and its current implementation. Despite its limitations, we believe that the proposed approach can significantly improve interaction quality and performance for mixed reality scenarios.
Daniel Kurz, Kiyoshi Kiyokawa, Haruo Takemura
VRST3
2007 Human Pose Estimation from Volume Data and Topological Graph Database
Hidenori Tanaka, Atsushi Nakazawa, Haruo Takemura
ACCV (1)3
2007 A 2D-3D integrated tabletop environment for multi-user collaboration
abstract
Abstract This paper proposes a novel tabletop display system for natural communication and flexible information sharing. The proposed system is specifically designed to integrate two‐dimensional (2D) and three‐dimensional (3D) user interfaces by using a multi‐user stereoscopic display, IllusionHole. The proposed system takes awareness into consideration and provides both 2D and 3D information and user interfaces. On the display, a number of standard Windows desktop environments are provided as personal workspaces, as well as a shared workspace with a dedicated graphical user interface. In the personal workspaces, users can simultaneously access existing applications and data, and exchange information between personal and shared workspaces. In this way, the proposed system can seamlessly integrate personal, shared, 2D, and 3D workspaces with conventional user interfaces and effectively support communication and information sharing. To demonstrate the capabilities of the proposed display system, a modeling application was implemented. A preliminary experiment confirmed the effectiveness of this system. Copyright © 2006 John Wiley & Sons, Ltd.
Kousuke Nakashima, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
Comput. Animat. Virtual Worlds4
2006 Spatial Reflectance Recovery under Complex Illumination from Sparse Images
abstract
A major challenge in inverse reflectometry is the acquisition of spatially varying materials. In this paper, we introduce a method to recover spatial reflectance from a sparse set of images under general illumination. Specifically, we first remove the high-frequency varying diffuse reflection term by using a low-order spherical harmonic approximation. This allows us to directly estimate the specular properties with a cluster fitting process, which simplifies the fitting processes and addresses the problem of data inadequacy for sparse images. As a result, we can reconstruct a truly spatially varying BRDF model of the surface from less than 10 images. Experimental results will be presented in order to demonstrate the effectiveness of the proposed algorithm.
Li Shen 0006, Haruo Takemura
CVPR (2)2
2006 GPU Accelerated Inverse Photon Mapping for Real-Time Surface Reflectance Modeling
abstract
This paper investigates the problem of object surface reflectance modeling, which is sometimes referred to as inverse reflectometry, for photorealistic rendering and effective multimedia applications. A number of methods have been developed for estimating object surface reflectance properties in order to render real objects under arbitrary illumination conditions. However, it is still difficult to densely estimate surface reflectance properties in real-time. This paper describes a new method for real-time estimation of the non-uniform surface reflectance properties in the inverse rendering framework. Experiments are conducted in order to demonstrate the usefulness and the advantage of the proposed methods through comparative study
Takashi Machida, Naokazu Yokoya, Haruo Takemura
ICME3
2006 A 2D-3D integrated interface for mobile robot control using omnidirectional images and 3D geometric models
abstract
This paper proposes a novel visualization and interaction technique for remote surveillance using both 2D and 3D scene data acquired by a mobile robot equipped with an omnidirectional camera and an omnidirectional laser range sensor. In a normal situation, telepresence with an egocentric-view is provided using high resolution omnidirectional live video on a hemispherical screen. As depth information of the remote environment is acquired, additional 3D information can be overlaid onto the 2D video image such as passable area and roughness of the terrain in a manner of video see-through augmented reality. A few functions to interact with the 3D environment through the 2D live video are provided, such as path-drawing and path-preview. Path-drawing function allows to plan a robot's path by simply specifying 3D points on the path on screen. Path- preview function provides a realistic image sequence seen from the planned path using a texture-mapped 3D geometric model in a manner of virtualized reality. In addition, a miniaturized 3D model is overlaid on the screen providing an exocentric view, which is a common technique in virtual reality. In this way, our technique allows an operator to recognize the remote place and navigate the robot intuitively by seamlessly using a variety of mixed reality techniques on a spectrum of Milgram's real-virtual continuum.
Kensaku Saitoh, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR4
2005 A Hybrid Image-Based and Model-Based Telepresence System Using Two-Pass Video Projection onto a 3D Scene Model
abstract
A telepresence system is presented that has the advantage of both model-based and image-based approaches, namely, free viewpoint control and real-time color update with live video. A remote place is presented as a virtual environment by using live video projection captured by a head-worn camera onto the static 3D geometry. The observer can then observe the remote place in cooperation with the remote camera man, and give him a set of 3D instructions by a mouse.
Takefumi Ogawa, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR3
2005 A Display Table for Strategic Collaboration Preserving Private and Public Information
Yoshifumi Kitamura, Wataru Osawa, Tokuo Yamaguchi, Haruo Takemura, Fumio Kishino
ICEC4
2005 A Study of Depth Visualization Techniques for Virtual Annotations in Augmented Reality
abstract
In this paper, we discuss depth visualization techniques for virtual annotations to alleviate the depth ambiguity problem. We begin by describing the depth ambiguity problem with regard to virtual annotations. Then a number of possible solutions are discussed by introducing a metaphor of monocular depth cues, and the effectiveness and characteristics of the three visualization techniques are shown through the preliminary experiments.
Kengo Uratani, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
VR4
2005 A 2D-3D integrated environment for cooperative work
abstract
This paper proposes a novel tabletop display system for natural communication and flexible information sharing. The proposed system is specifically designed for integration of 2D and 3D user interfaces, using a multi-user stereoscopic display, IllusionHole. The proposed system takes awareness into consideration and provides both 2D and 3D information and user interfaces. On the display, a number of standard Windows desktop environments are provided as personal workspaces, as well as a shared workspace with a dedicated graphical user interface. In personal workspaces, users can simultaneously access existing applications and data, and exchange information between personal and shared workspaces. In this way, the proposed system can seamlessly integrate personal, shared, 2D and 3D workspaces with conventional user interfaces and effectively support communication and information sharing. To demonstrate capabilities of the proposed display system, a modeling application has been implemented. A preliminary experiment confirmed the effectiveness of the system.
Kousuke Nakashima, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
VRST4
2004 Unified Gesture-Based Interaction Techniques for Object Manipulation and Navigation in a Large-Scale Virtual Environment
Yusuke Tomozoe, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
VR4
2003 Surface Reflectance Modeling of Real Objects with Interreflections
abstract
In mixed reality, especially in augmented virtuality which virtualizes real objects, it is important to estimate object surface reflectance properties to render the objects under arbitrary illumination conditions. Though several methods have been explored to estimate the surface reflectance properties, it is still difficult to estimate surface reflectance parameters faithfully for complex objects which have nonuniform surface reflectance properties and exhibit interreflections. We describe a new method for densely estimating nonuniform surface reflectance properties of real objects constructed of convex and concave surfaces with interreflections. We use registered range and surface color texture images obtained by a laser rangefinder. Experiments show the usefulness of the proposed method.
Takashi Machida, Naokazu Yokoya, Haruo Takemura
ICCV3
2003 A High Immersive Tele- Directing System Using CyberDome
Tomoaki Adachi, Takefumi Ogawa, Kiyoshi Kiyokawa, Haruo Takemura
INTERACT4
2002 Dense 3-D Reconstruction of an Outdoor Scene by Hundreds-Baseline Stereo Using a Hand-Held Video Camera
Tomokazu Sato, Masayuki Kanbara, Naokazu Yokoya, Haruo Takemura
Int. J. Comput. Vis.4
2001 Workshop 1: Virtual Reality and its Application for Human Centered System
Haruo Takemura, Makoto Sato
VR1
2000 A Stereo Vision-Based Augmented Reality System with a Wide Range of Registration
abstract
Proposes a vision-based augmented reality system with a wide range of registration. To realize an augmented reality system, it is required to geometrically register real and virtual worlds. In the case of a vision-based augmented reality with marker tracking, its measurement range is usually limited because markers placed in the real world should be captured by cameras. The proposed method realizes a stereo vision-based augmented reality system with a wide range of registration by automatically detecting and tracking new, markers that come into sight. The feasibility of the system has been successfully demonstrated through experiments.
Masayuki Kanbara, Haruo Takemura, Naokazu Yokoya, Hidehiko Iwasa
ICPR2
2000 Real-Time Camera Parameter Estimation from Images for a Mixed Reality System
abstract
This paper describes a method of estimating the position and orientation of a camera for constructing a mixed reality (MR) system. In an MR system 3D virtual objects should be merged into a 3D real environment at a right position in real time. To acquire the user's viewing position and orientation is the main technical problem of constructing an MR system. The user's viewpoint can be determined by estimating the position and orientation of a camera using images taken at the viewpoint. Our method estimates the camera pose using screen coordinates of captured color fiducial markers whose 3D positions are known. The method consists of three algorithms for perspective n-points problems and rises each algorithm selectively. The method also estimates the screen coordinates of untracked markers that are occluded or are out of the view. It has been found that an experimental MR system that is based on the proposed method can seamlessly merge 3D virtual objects into a 3D real environment at right position in real-time and allows users to look around an area in which numbers are placed.
Takashi Okuma, Katsuhiko Sakaue, Haruo Takemura, Naokazu Yokoya
ICPR3
2000 Construction and Presentation of a Virtual Environment Using Panoramic Stereo Images of a Real Scene and Computer Graphics Models
abstract
The progress in computer graphics has made it possible to construct various virtual environments such as urban or natural scenes. The paper proposes a hybrid method to construct a realistic virtual environment containing an existing real scene. The proposed method combines two different types of 3-D models. A 3-D geometric model is used to represent virtual objects in the user's vicinity, enabling a user to handle virtual objects. A texture mapped cylindrical 2.5-D model of a real scene is used to render the background of the environment, maintaining real-time rendering and increasing realistic sensation. The cylindrical 2.5-D model is generated from cylindrical stereo images captured by an omnidirectional stereo imaging sensor. A prototype mixed reality system has been developed to confirm the feasibility of the method, in which panoramic binocular stereo images are projected on a cylindrical immersive projective display depending on the user's view point in real time.
J. Shimamura, Haruo Takemura, Naokazu Yokoya, Kazumasa Yamazawa
ICPR2
2000 Real-Time Generation and Presentation of View-Dependent Binocular Stereo Images Using a Sequence of Omnidirectional Images
abstract
This paper presents a new method to generate and present arbitrarily directional binocular stereo images from a sequence of omnidirectional images. A sequence of omnidirectional images is taken by moving an omnidirectional image sensor in a static real environment. The motion of the omnidirectional image sensor is constrained to a plane. The sensor's route and speed are known. In the proposed method, a fixed length of the sequence is buffered in a computer to generate arbitrarily directional binocular stereo images by combining captured rays. Using the method a user can look around a scene in the distance with rich 3D sensation without significant time delay. This paper describes the principle of real-time generation of binocular stereo images. In addition, we introduce a prototype telepresence system of view-dependent stereo image generation and presentation.
Koichiro Yamaguchi, Haruo Takemura, Kazumasa Yamazawa, Naokazu Yokoya
ICPR2
2000 A Stereoscopic Video See-Through Augmented Reality System Based on Real-Time Vision-Based Registration
abstract
In an augmented reality system, it is required to obtain the position and orientation of the user's viewpoint in order to display the composed image while maintaining a correct registration between the real and virtual worlds. All the procedures must be done in real time. This paper proposes a method for augmented reality with a stereo vision sensor and a video see-through head-mounted display (HMD). It can synchronize the display timing between the virtual and real worlds so that the alignment error is reduced. The method calculates camera parameters from three markers in image sequences captured by a pair of stereo cameras mounted on the HMD. In addition, it estimates the real-world depth from a pair of stereo images in order to generate a composed image maintaining consistent occlusions between real and virtual objects. The depth estimation region is efficiently limited by calculating the position of the virtual object by using the camera parameters. Finally, we have developed a video see-through augmented reality system which mainly consists of a pair of stereo cameras mounted on the HMD and a standard graphics workstation. The feasibility of the system has been successfully demonstrated with experiments.
Masayuki Kanbara, Takashi Okuma, Haruo Takemura, Naokazu Yokoya
VR3
2000 An immersive modeling system for 3D free-form design using implicit surfaces
abstract
We present a new free-form interactive modeling technique based on the metaphor of clay work. This paper discusses design issues and an immersive modeling system which enables a user to design intuitively and interactively 3D solid objects with curved surfaces by using one's finger. Shape deformation is expressed by simple formulas without complex calculation because of skeletal implicit surfaces employed to represent smooth free-form surfaces. A polygonization algorithm that generates polygonal representation from implicit surfaces is developed to reduce the time required for rendering curved surfaces, since conventional graphics hardware is optimized for displaying polygons. The prototype system has shown that a user can design 3D solid objects composed of curved surfaces in a short time by deforming objects intuitively using one's finger in real time.
Masatoshi Matsumiya, Haruo Takemura, Naokazu Yokoya
VRST2
1998 Acquisition of Three-Dimensional Information Using Omnidirectional Stereo Vision
Atsushi Chaen, Kazumasa Yamazawa, Naokazu Yokoya, Haruo Takemura
ACCV (1)4
1998 Memory-based self-localization using omnidirectional images
abstract
This paper proposes a new self-localization method using an omnidirectional image sensor which can observe a surrounding environment with 360-degree of view. The method extracts information which is identical for the position of a sensor and invariant against the rotation of the sensor by generating an autocorrelation image from an observed omnidirectional image. The location of the sensor is estimated by evaluating the similarity among the autocorrelation image of an observed image and stored autocorrelation images. The similarity of autocorrelation images is evaluated in low dimensional eigenspaces generated with stored autocorrelation images. We have conducted experiments with real images and examined the performance of the proposed method. The results show that accurate and robust estimation of the sensor's position is possible with our method.
Nobuhiro Aihara, Hidehiko Iwasa, Naokazu Yokoya, Haruo Takemura
ICPR4
1998 Real-time tracking of multiple moving objects in moving camera image sequences using robust statistics
abstract
In this paper, we propose a new method for detection and tracking of moving objects from a moving camera image sequence using robust statistics and active contour models. We assume that the apparent background motion between two consecutive image frames can be approximated by affine transformation. In order to register the static background, we estimate affine transformation parameters using LMedS (least median of squares) method which is a kind of robust statistics. Split-and-merge contour models are employed for tracking multiple moving objects which have been recently proposed by the authors. Image energy of contour models is defined based on the image which is obtained by subtracting the previous frame transformed with estimated affine parameters from the current frame. We have implemented the method on an image processing system which consists of DSP boards for real-time tracking of moving objects from a moving camera image sequence.
Shoichi Araki, Takashi Matsuoka, Haruo Takemura, Naokazu Yokoya
ICPR3
1998 A factorization method using 3D linear combination for shape and motion recovery
abstract
This study proposes a new factorization method for shape and motion recovery. In the past, the fourth greatest singular value of the measurement matrix was ignored. But when noise is large enough so that the fourth greatest singular value can not be ignored, it would be difficult to get reliable results by using the traditional factorization method. In order to acquire reliable results, We start with adopting an orthogonalization method to find a matrix which is composed of three mutually orthogonal vectors. By using this matrix, another matrix can be obtained. Then, the two expected matrices which represent shape of object and motion of camera/object, can be obtained through normalization. This study also conducts several experiments to discuss the feasibility of the proposed method.
Kuo-Chang Hwang, Naokazu Yokoya, Haruo Takemura, Kazumasa Yamazawa
ICPR3
1998 Generation of high-resolution stereo panoramic images by omnidirectional imaging sensor using hexagonal pyramidal mirrors
abstract
We have developed a high-resolution omnidirectional stereo imaging sensor that can take images at video-rate. The sensor system takes an omnidirectional view by a component constructed of six cameras and a hexagonal pyramidal mirror and acquires stereo views by symmetrically connecting two sensor components. The paper describes a method of generating stereo panoramic images by using our sensor. First, the sensor system is calibrated; that is, twelve cameras are correctly aligned with pyramidal mirrors and the Tsai's method restores the radial distortion of each camera image. Stereo panoramic images are then computed by registering the camera images captured at the same time.
Takahito Kawanishi, Kazumasa Yamazawa, Hidehiko Iwasa, Haruo Takemura, Naokazu Yokoya
ICPR4
1998 An augmented reality system using a real-time vision based registration
abstract
Describes a prototype of an augmented reality system using vision-based registration. In order to build an augmented reality system with video see-through image composition, camera parameters for generating virtual objects must be obtained at video-rate. The system estimates camera parameters from four known markers in an image sequence captured by a small CCD camera mounted on a HMD (head mounted display). Virtual objects are overlaid upon a captured image sequence in real-time.
Takashi Okuma, Kiyoshi Kiyokawa, Haruo Takemura, Naokazu Yokoya
ICPR3
1998 Visual surveillance and monitoring system using an omnidirectional video camera
abstract
This paper describes a visual surveillance and monitoring system which is based on omnidirectional imaging and view-dependent image generation from omnidirectional video streams. While conventional visual surveillance and monitoring systems usually consist of either a number of fixed regular cameras or a mechanically controlled camera, the proposed system has a single omnidirectional video camera using a hyperboloidal mirror. This approach has an advantage of less latency in looking around a large field of view. In a prototype system developed, the viewing direction is determined by viewers' head tracking, by using a mouse, or by moving object trading in the omnidirectional image.
Yoshio Onoe, Naokazu Yokoya, Kazumasa Yamazawa, Haruo Takemura
ICPR4
1998 Telepresence by Real-Time View-Dependent Image Generation from Omnidirectional Video Streams
Yoshio Onoe, Kazumasa Yamazawa, Haruo Takemura, Naokazu Yokoya
Comput. Vis. Image Underst.3
1996 Facial component extraction by cooperative active nets with global constraints
abstract
This paper describes a new method for extracting facial components from a color image using cooperative active nets with global constraints. We utilize an active net model which is a region extraction method based on energy minimization principle. Each net deforms with its own energy being minimized and the position of the nets are controlled by minimizing an additional energy defined by global constraints on their placement. Our method has been experimentally shown to be robust to variations of facial size, position and rotation.
Ryuji Funayama, Naokazu Yokoya, Hidehiko Iwasa, Haruo Takemura
ICPR4
1996 Analysis and synthesis of six primary facial expressions using range images
abstract
This paper describes a synthesis-by-analysis approach using cylindrical range images for producing human facial images with realistic expressions. First view-independent representations of 3D locations of facial feature points are obtained by using an object-centered coordinate system defined on a face. Then we quantify facial feature points locations for the neutral expression, and six primary expressions: "anger", "disgust", "surprise", "fear", "happiness" and "sadness" Applying an image warping technique to both range and texture images, we finally generate 3D facial expression images from neutral expression images using motion vectors of facial feature points.
Yumiko Tatsuno, Naokazu Yokoya, Hidehiko Iwasa, Haruo Takemura
ICPR5
1996 VLEGO: a simple two-handed modeling environment based on toy blocks
abstract
This paper describes a case study of building a prototype of an immersive three dimensional (3-D) modeler which supports simple two-handed operations. Designing 3-D objects in a virtual environment has a number of advantages for 3-D geometry creation over designing with traditional computer aided design (CAD) tools. In order to enhance the human-computer interaction in a virtual workspace, two-handed spatial input has been incorporated into a few 3-D designing applications. However, existing 3-D designing tools do not utilize two handed interaction for enhancing the interface sufficiently. Our prototype immersive modeler, VLEGO, employs some features of toy blocks to give flexible two-handed interaction for 3-D design. Features of VLEGO can be summarized as follows: Firstly, VLEGO supports various two-handed operations and hence it makes design environment intuitive and efficient. Secondly, possible location and orientation of primitives are discretely limited so that the user can arrange objects accurately with ease. Finally, the system automatically avoids collisions among primitives and adjusts their positions. As a result, precise design of 3-D objects can be achieved easily by using a set of two-handed operations in intuitive way. This paper describes the design and implementation of VLEGO as well as an experiment for examining the effectiveness of two-handed interaction.
Kiyoshi Kiyokawa, Haruo Takemura, Yoshiaki Katayama, Hidehiko Iwasa, Naokazu Yokoya
VRST2
1995 Virtual Space Teleconferencing: Real-Time Reproduction of 3D Human Images
Jun Ohya, Yasuichi Kitamura, Fumio Kishino, Nobuyoshi Terashima, Haruo Takemura, Hirofumi Ishii
J. Vis. Commun. Image Represent.5
1994 Efficient collision detection among objects in arbitrary motion using multiple shape representations
abstract
We propose an efficient method for detecting potential collisions among multiple objects with arbitrary motion (translation and rotation) in 3D space. The method is useful for online monitoring and path planning in a 3D environment in which there are multiple independently-moving objects. The method consists of two main stages: 1) the coarse stage, an approximate test is performed to identify interfering objects in the entire workspace using octree representation of object shapes; and 2) the fine stage, polyhedral representation of object shapes is used to more accurately identify any object parts that might cause interference and collisions. For this purpose, specific pairs of faces belonging to any of the interfering objects found in the first stage are tested, thus performing detailed computation on a reduced amount of data. Experimental results, which demonstrate the efficiency of the proposed collision detection method, are given.
Yoshifumi Kitamura, Haruo Takemura, Narendra Ahuja, Fumio Kishino
ICPR (1)2
1992 Cooperative Work Environment Using Virtual Workspace
abstract
A virtual environment, which is created by computer graphics and an appropriate user interface, can be used in many application fields, such as teleopetution, telecommunication and real time simulation.Furthermore, if this environment could be shared by multiple users, there would be more potential applications.Discussed in this paper is a case study of building a prototype of a cooperative work environment using a virtual environment, where more than two people can solve problemscooperatively, including design strategies and implementirig issues.An environment where two operators can directly grasp, move or release stereoscopic computer graphics images by hand is implemented.The system is built by combining head position tracking stereoscopic displays, hand gesture input devices and graphics workstations.Our design goal is to utilize this type of interface for a future teleconferencing system.In order to provide good interactivity for users, we discuss potential bottlenecks and their solutions.The system allows two users to share a virtual environment and to organize 3-D objects cooperatively.
Haruo Takemura, Fumio Kishino
CSCW1