Kiyoshi Kiyokawa

dblp:19/6374 · DBLP profile ↗
← Back
123ranked-venue papers
11as first author
44since 2021 · last 2026
0000-0003-2260-1707ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 104 · 11 first-author · 37 since 2021Human-computer interaction and ubiquitous computing · 75 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Challenges in Synchronous & Remote Collaboration Around Visualization
abstract
We characterize 16 challenges faced by those investigating and developing remote and synchronous collaborative experiences around visualization. Our work reflects the perspectives and prior research efforts of an international group of 29 experts from across human-computer interaction and visualization sub-communities. The challenges are anchored around five collaborative activities that exhibit a centrality of visualization and multimodal communication. These activities include exploratory data analysis, creative ideation, visualization-rich presentations, joint decision making grounded in data, and real-time data monitoring. The challenges also reflect the changing dynamics of these activities in the face of recent advances in extended reality (XR) and artificial intelligence (AI). As an organizing scheme for future research at the intersection of visualization and computer-supported cooperative work, we align the challenges with a sequence of four sets of research and development activities: technological choices, social factors, AI assistance, and evaluation.
Matthew Brehmer, Maxime Cordeil, Christophe Hurter, Takayuki Itoh, Wolfgang Büschel, Mahmood Jasim, Arnaud Prouzeau, David Saffo, Lyn Bartram, Sheelagh Carpendale, Chen Zhu-Tian, Andrew Cunningham, Tim Dwyer, Samuel Huron, Masahiko Itoh, Alark Joshi, Kiyoshi Kiyokawa, Hideaki Kuzuoka, Bongshin Lee, Gabriela Molina León, Harald Reiterer, Bektur Ryskeldiev, Jonathan A. Schwabish, Brian A. Smith 0001, Yasuyuki Sumi, Ryo Suzuki 0001, Anthony Tang 0001, Yalong Yang 0001, Jian Zhao 0010
CHI17
2026 EMA: Effort Metric Attention for Anatomical Effort-Guided Human Motion Diffusion
Joshua Siy, Huakun Liu, Yutaro Hirao, Monica Perusquía-Hernández, Hideaki Uchiyama, Kiyoshi Kiyokawa
FG6
2026 Electrooculography-Based Detection of Refractive Vision Problems
abstract
Early detection of visual impairments remains a persistent challenge, especially due to the subtle and often unnoticed nature of early-stage symptoms. Recent works have attempted to transition clinical tests to home-based services or develop innovative diagnostic methods, but most approaches remain self-initiated and discrete. In this study, we focused on refractive disorders and explored the feasibility of using electrooculography (EOG) to detect changes in refractive power passively. Thirty-nine participants used optometry trial lenses to simulate different refractive conditions. Participants performed a series of visual tasks while their EOG signals were recorded. We trained classification models to predict simulated refractive power levels relative to baseline visual condition across multiple evaluation settings, including within-subject, temporal generalization, and across-subject scenarios. The findings reveal that refractive power classification models achieve a mean accuracy of $0.950 \pm 0.034$ in within-subject, within-condition scenarios. Within-subject models tested on data from a different time point showed highly variable performance. While some participants achieved promising results, overall accuracy remained low, with a mean of $0.159 \pm 0.285$. We employed three strategies to evaluate the across-subject models. Naive models performed poorly ($0.161 \pm 0.063$) and linear normalization provided limited improvement ($0.175 \pm 0.062$). However, the fine-tuning strategy substantially improved the model's performance ($0.785 \pm 0.123$). EOG signals contain useful information for refractive power classification, particularly in personalized contexts. However, generalizing across time and individuals remains challenging. Overall, this work offers valuable insights for advancing EOG-based systems aimed at passive, real-time monitoring of visual conditions.
Xin Wei 0007, Huakun Liu, Yutaro Hirao, Monica Perusquía-Hernández, Katsutoshi Masai, Hideaki Uchiyama, Kiyoshi Kiyokawa
IEEE J. Biomed. Health Informatics7
2026 Visual and Somatosensory Integration With Higher Sitting Posture Enhances the Sense of Standing and Self-Motion in Seated VR
abstract
Users are often seated in the real environment, while their virtual avatars either remain standing stationary or move in virtual reality (VR). This creates posture inconsistencies between the real and virtual embodiment representations. The relationship between posture consistency in locomotion techniques and sense of presence in VR is still unclear. This study investigates how visual and somatosensory integration affects the sense of standing (SoSt) and the sense of self-motion (SoSm) when the sitting posture is varied slightly, including highlighting the importance of sitting posture for locomotion design in VR. The degree and occurrence of SoSt and SoSm were assessed by subjective experiments, and it was found that higher sitting and lower sitting postures present higher SoSt and lower SoSm, respectively. Invocation of SoSt also influences postural perception. Perception of travel distance varied according to the posture condition when identical visual flow was presented. The findings suggest that visual and somatosensory integration related to posture enhances SoSt and SoSm, and a sitting posture with a higher seating position is recommended in seated VR locomotion design.
Daiki Hagimori, Naoya Isoyama, Monica Perusquía-Hernández, Shunsuke Yoshimoto, Hideaki Uchiyama, Nobuchika Sakata, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.7
2026 Tag-Along Virtual Windows Increase Perceived Resistance and Task Load in Augmented Reality
abstract
Augmented Reality (AR) can enhance accessibility by anchoring virtual windows to the user's body. Among common approaches, head-following windows help maintain floating virtual windows within the user's field of view. Previous studies have actively explored this new design space to improve user experience and efficiency. In contrast, this study focuses on the perceived resistance of head-following windows in AR, despite their lack of physical mass. We conducted a within-subject experiment with 24 participants, manipulating Follow-Up Delay, Window Size, and UI Type. We measured subjective resistance ratings, NASA-TLX (Raw TLX Scores), and the gaze-head angular offset. The results showed that both a certain level of Follow-Up Delay and the Tag-Along elicited significantly stronger perceived resistance as well as task load. Although Window Size alone did not show a significant effect on resistance ratings, we observed an interaction between the size and UI Type. These findings extend existing pseudo-haptics research by revealing the previously unexplored domain of resistance in head-based interactions with head-following virtual windows. We further provide design implications for head-following windows in AR.
Motoki Kagami, Yuta Kataoka, Yutaro Hirao, Monica Perusquía-Hernández, Satoshi Hashiguchi, Hideaki Uchiyama, Kiyoshi Kiyokawa, Shohei Mori
IEEE Trans. Vis. Comput. Graph.7
2026 HybridSphere: Enhancing Hybrid Meetings with Avatar-Based VR Environments
abstract
With recent advances in information and communication technologies, Hybrid meetings, where local attendees are physically present and remote participants join virtually, have become increasingly common. However, remote participants often experience reduced contextual awareness and a sense of isolation. To address these issues, we propose HybridSphere, a hybrid meeting system that reconstructs a shared virtual reality environment for remote participants. The system employs a 360-degree camera and pose estimation to generate real-time avatar representations of local attendees to allow remote users with head-mounted displays to experience the meeting as if all participants are in the same virtual space. We conducted a user study comparing HybridSphere to a baseline condition in which remote participants viewed an unmodified 360-degree video. Although the avatars were rated lower in perceived trustworthiness and likability due to limited visual fidelity, participants appreciated the seated-6DoF function. These results suggest that immersive and spatially flexible VR representations can enhance remote engagement in hybrid meetings.
Koji Momota, Shizuka Shirai, Masato Kobayashi 0001, Naoya Chiba, Photchara Ratsamee, Kiyoshi Kiyokawa, Yuuki Uranishi
IEEE Trans. Vis. Comput. Graph.6
2026 IEEE VR 2026 Introducing the Special Issue
Kiyoshi Kiyokawa, Maud Marchal
IEEE Trans. Vis. Comput. Graph.2
2026 Mask Balancing: Perception-Driven Dynamic Visibility Enhancement for Occlusion-Capable Optical See-Through Head-Mounted Displays
abstract
The poor transparency of occlusion-capable optical see-through head-mounted displays (OC-OSTHMDs) deteriorates the visibility of the real scene, hindering the practical application of the devices. Previous works mitigate the issue by upgrading the transmittance of the spatial light modulator (SLM). However, the strategy soon reaches a limit because further optimization requires improving the transmittance of all optical elements, e.g., lenses and beam splitters. Moreover, pixelated occlusion usually relies on polarizing the real scene light, inevitably cutting the input optical power by half. To overcome this limitation, we propose a mask balancing method that improves real-scene brightness through polarization blending. Specifically, the s-polarized component, which passes through the optical system to provide occlusion-capable vision, is blended with the p-polarized component, which bypasses the system to preserve the raw view of the real scene. The blending is realized by simply modulating the cross-angle between a polarizing beam splitter and a linear polarizer, benefiting the robustness and versatility of the proposed method. We introduce a perception-driven blending approach, where the cross-angle is optimized in real-time to balance the visibility of the real scene and the texture and lighting of the virtual object. A benchtop prototype is built. A user study with 12 participants is conducted to quantify the visibility threshold of the texture and lighting of virtual objects. Then, a user study with 12 participants proves that the proposed method improves the visibility of the real scene while keeping a good appearance of the virtual object. We believe the proposed method is an important step toward developing practical solutions for OC-OSTHMDs.
Yan Zhang 0101, Rundong Chu, Qingtai Dong, Xiaodan Hu, Keyao You, Zixuan Guo 0003, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang
IEEE Trans. Vis. Comput. Graph.9
2025 UMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband Units
abstract
Sparse wearable inertial measurement units (IMUs) have gained popularity for estimating 3D human motion. However, challenges such as pose ambiguity, data drift, and limited adaptability to diverse bodies persist. To address these issues, we propose UMotion, an uncertainty-driven, online fusing-all state estimation framework for 3D human shape and pose estimation, supported by six integrated, body-worn ultra-wideband (UWB) distance sensors with IMUs. UWB sensors measure inter-node distances to infer spatial relationships, aiding in resolving pose ambiguities and body shape variations when combined with anthropometric data. Unfortunately, IMUs are prone to drift, and UWB sensors are affected by body occlusions. Consequently, we develop a tightly coupled Unscented Kalman Filter (UKF) framework that fuses uncertainties from sensor data and estimated human motion based on individual body shape. The UKF iteratively refines IMU and UWB measurements by aligning them with uncertain human motion constraints in real-time, producing optimal estimates for each. Experiments on both synthetic and real-world datasets demonstrate the effectiveness of UMotion in stabilizing sensor data and the improvement over state of the art in pose accuracy. Code is available at: https://github.com/kk9six/umotion.
Huakun Liu, Hiroki Ota, Xin Wei 0007, Yutaro Hirao, Monica Perusquía-Hernández, Hideaki Uchiyama, Kiyoshi Kiyokawa
CVPR7
2025 Have a Seat: An Enhanced Reactive Alignment of a Single Target's Position and Angle from the User's Perspective in VR
abstract
Redirected Walking (RDW) techniques allow users to explore virtually infinite environments within constrained physical spaces. However, achieving precise alignment between physical and virtual targets remains a significant challenge. In particular, when both position and orientation of the targets need to align in order to get a proper haptic feedback like siting on a virtual chair. This paper introduces a revised version of the Reactive Alignment (REA) controller that simultaneously minimizes the Angular and Positional Distance Errors between a physical and a virtual target. The proposed method enhances spatial alignment and optimizes user navigation using a novel rotation gain control algorithm that takes angular misalignment into account. In addition, a new metric,$\Delta p$, is proposed to quantify the angular alignment, complementing the redefined Physical Distance Error (PDE) for positional accuracy. We implemented the algorithm on Oculus Quest head-mounted display and utilized the HMD's physical space tracking to locate the physical prop's location without the need of any external tracking. We also incorporated saccadic redirection by utilizing the HMD's eyetracking functionality to complement the revised REA approach. A user study demonstrates that the revised REA controller outperforms the original REA by reducing Physical Distance Error, angular error, and reset counts. It also enhanced user interaction with physical props by enabling users to successfully sit on a physical chair 60% of the time compared to 0% with the original REA when$\Delta p$is zero.
Habiba H. AbdelAziz, Yutaro Hirao, Monica Perusquía-Hernández, Hideaki Uchiyama, Kiyoshi Kiyokawa
ISMAR5
2025 X-Mask: Improving Soft-Edge Occlusion in Optical See-Through Displays with Cross-Shaped Pinholes
abstract
Placing a transparent liquid crystal display (LCD) into the light path is a simple approach to create occlusion-capable optical seethrough head-mounted displays (OST-HMDs) that suffers from defocused (soft-edge) occlusion where the mask leakage partially occludes surrounding content as well. Creating a focused (hard-edge) occlusion that does not suffer from mask leakage requires complicated, bulky optical setups. We present X-Mask, a pinhole-arraybased OST-HMD that creates a sharp occlusion mask without the need for a bulky setup requiring only two transparent LCD layers. By rendering a pinhole array on the layer closer to the user's eye, our system functions as a programmable aperture layer that extends the effective depth of field and improves the sharpness of the occlusion mask rendered on the second LCD layer. Utilizing a conventional circular pinhole would result in non-uniform brightness and contrast. By changing the pinhole shape to a cross enables nearoptimal retinal tiling with reduced overlaps and gaps. To accommodate pupil size variation, focus distance, and gaze direction, our system design allows for gaze-contingent adjustment of both LCD layers. We validate X-Mask in simulations and a physical prototype showing improved occlusion sharpness and visual uniformity.
Xiaodan Hu, Christoph Ebner, Yan Zhang 0101, Kiyoshi Kiyokawa, Alexander Plopski
ISMAR4
2025 The Awe-Some Spectrum: Self-Reported Awe Varies by Eliciting Scenery and Presence in Virtual Reality, and the User's Nationality
abstract
Awe is a multifaceted emotion often associated with the perception of vastness, that challenges existing mental frameworks. Despite its growing relevance in affective computing and psychological research, awe remains difficult to elicit and measure. This raises the research questions of how awe can be effectively elicited, which factors are associated with the experience of awe, and whether it can reliably be measured using biosensors. For this study, we designed 10 immersive Virtual Reality (VR) scenes with dynamic transitions from narrow to vast environments. These scenes were used to explore how awe relates to environmental features (abstract, human-made, nature), personality traits, and country of origin. We collected skin conductance, respiration, self-reported awe and presence data from participants from Germany, Japan, and Jordan. Our results indicate that self-reported awe varies significantly across countries and scene types. In particular, a scene depicting outer space elicited the strongest awe. Scenes that elicited high selfreported awe also induced a stronger sense of presence. However, we found no evidence that awe ratings are correlated with physiological responses. These findings challenge the assumption that awe is reliably reflected in autonomic arousal and underscore the importance of cultural and perceptual context. Our study offers new insights into how immersive VR can be designed to elicit awe, and suggests that subjective reports - rather than physiological signals - remain the most consistent indicators of emotional impact.
Melissa Steininger, Alexander Marquardt, Monica Perusquía-Hernández, Marvin Lehnort, Hiromu Otsubo, Felix Dollack, Ernst Kruijff, Björn Krüger, Kiyoshi Kiyokawa, Bernhard E. Riecke
ISMAR9
2025 Mind Your Vision: A Passive Multimodal Framework for Refractive Disorders Measurement Combining Electrooculography and Eye Tracking
abstract
Refractive errors are among the most common visual impairments globally, yet their diagnosis often relies on active user participation and clinical oversight. This study explores a passive method for estimating refractive power using two eye movement recording techniques: electrooculography (EOG) and video-based eye tracking. Using a publicly available dataset recorded under varying diopter conditions, we trained Long Short-Term Memory (LSTM) models to classify refractive power from unimodal (EOG or video-based eye tracking) and multimodal configurations. In the context of eye movement analysis, EOG captures fine-grained electrical signals, while video-based tracking provides rich features such as pupil dynamics and gaze behavior, making the two modalities complementary. We assess performance in both subject-dependent and subject-independent settings to evaluate model personalization and generalizability across individuals. Results show that the multimodal model consistently outperforms unimodal models, achieving the highest average accuracy in both settings: 96.568% in the subject-dependent scenario and 9.344% in the subject-independent scenario. Statistical comparisons in the subject-dependent setting confirmed that both unimodal and multimodal models significantly exceeded the chance level. Among them, the multimodal model significantly outperformed the EOG and eye-tracking models. The strong performance of subject-dependent models highlights the potential for developing personalized models tailored to the target user for refractive power monitoring. However, generalization remains limited, with classification accuracy only marginally above chance in the subject-independent evaluations. Our findings demonstrate both the potential and current limitations of eye movement data-based refractive error estimation, contributing to the development of continuous, non-invasive screening methods using EOG signals and eye-tracking data.
Xin Wei 0007, Huakun Liu, Yutaro Hirao, Monica Perusquía-Hernández, Katsutoshi Masai, Hideaki Uchiyama, Kiyoshi Kiyokawa
MUM7
2025 Color Correction for Occlusion-Capable Optical See-Through Head-Mounted Displays by Using Phase-Modulation
abstract
Occlusion-capable optical see-through head-mounted displays (OC-OSTHMDs) overcome the deficiency of semi-transparent virtual images by selectively cutting off light emitted from the physical background, considerably improving the graphics performance of augmented reality (AR). Existing OC-OSTHMDs achieve compact form factors by compressing the optical system based on the modulation of light polarization. However, the wavelength sensitivity of polarizing optical elements (POEs) causes color aberration in the see-through view. In this paper, we propose the spectrum- tuning method that mitigates color aberration of the see-through view caused by the wavelength sensitivity of OC-OSTHMDs. The methods operate on a spectrum-based color perception model that formulates the variation of the visible spectrum through OC-OSTHMDs. The optimization is performed globally, requiring minimal computation at runtime. A bench-top prototype of the OC-OSTHMD was built to validate these methods. Experimental results demonstrate that the spectrum-tuning method reduces the color difference by 18.1%. Additionally, the advantages of OC-OSTHMDs in presenting fluid animations in AR scenarios are demonstrated based on the prototype and a multi-buffer mask synthesis method.
Yan Zhang 0101, Shulin Hong, Weike Qian, Keyao You, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang
VR6
2025 Perception-Driven Soft-Edge Occlusion for Optical See-Through Head-Mounted Displays
abstract
Systems with occlusion capabilities, such as those used in vision augmentation, image processing, and optical see-through head-mounted display (OST-HMD), have gained popularity. Achieving precise (hard-edge) occlusion in these systems is challenging, often requiring complex optical designs and bulky volumes. On the other hand, utilizing a single transparent liquid crystal display (LCD) is a simple approach to create occlusion masks. However, the generated mask will appear defocused (soft-edge) resulting in insufficient blocking or occlusion leakage. In our work, we delve into the perception of soft-edge occlusion by the human visual system and present a preference-based optimal expansion method that minimizes perceived occlusion leakage. In a user study involving 20 participants, we made a noteworthy observation that the human eye perceives a sharper edge blur of the occlusion mask when individuals see through it and gaze at a far distance, in contrast to the camera system's observation. Moreover, our study revealed significant individual differences in the perception of soft-edge masks in human vision when focusing. These differences may lead to varying degrees of demand for mask size among individuals. Our evaluation demonstrates that our method successfully accounts for individual differences and achieves optimal masking effects at arbitrary distances and pupil sizes.
Xiaodan Hu, Yan Zhang 0101, Alexander Plopski, Yuta Itoh 0001, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.8
2025 Application of Transitional Mixed Reality Interfaces: A Co-Design Study with Flood-Prone Communities
abstract
Flood risk communication in disaster-prone communities often relies on traditional tools (e.g., paper and browser-based hazard/flood maps) that struggle to engage community stakeholders and reflect intuitive flood situations. In this paper, we applied the transitional mixed reality (MR) interface concept from pioneering work and extended it for flood risk communication scenarios through co-design with community stakeholders to help vulnerable residents understand flood risk and facilitate preparedness. Starting with an initial transitional MR prototype, we conducted three iterative workshops - each dedicated to device usability, visualization techniques, and interaction methods. We collaborated with diverse community stakeholders in flood-prone areas, collecting feedback to refine the system according to community needs. Our preliminary evaluation indicates that this co-designed system significantly improves user understanding and engagement compared to traditional tools, though some older residents faced usability challenges. We detailed this iterative co-design process, critical insights and design implications, offering our work as a practical case of mixed reality application in strengthening flood risk communication. We also discuss the system's potential to support community-driven collaboration in flood preparedness.
Zhiling Jie, Geert Lugtenberg, Armin Teubert, Makoto Fujisawa, Hideaki Uchiyama, Kiyoshi Kiyokawa, Isidro Butaslac, Taishi Sawabe, Hirokazu Kato 0001
IEEE Trans. Vis. Comput. Graph.7
2025 IEEE VR 2025 Introducing the Special Issue
Han-Wei Shen, Kiyoshi Kiyokawa, Maud Marchal
IEEE Trans. Vis. Comput. Graph.2
2025 IEEE ISMAR 2025 Introducing the Special Issue
Han-Wei Shen, Kiyoshi Kiyokawa, Maud Marchal
IEEE Trans. Vis. Comput. Graph.2
2024 ReAR Indicators: Peripheral Cycling Indicators for Rear-Approaching Hazards
abstract
During cycling activities, cyclists often focus on pedestrians, vehicles or road conditions in front of their bicycle. Because of this forward focus, approaching vehicles from behind can easily be missed, which can result in accidents, injury, or death. Although rear information can viewed with handle-mounted mirrors or monitors, looking down can distract the cyclist from other hazards.
Guanghan Zhao, Xiaodan Hu, Jason Orlosky, Kiyoshi Kiyokawa
AVI4
2024 U2R: Underwater Ultrasonic Reflection Wave Dataset Toward Pose-Invariant Material Recognition
abstract
In underwater environments, the reflected ultrasonic waves from objects generally provide more than just information about their color and shape for object recognition. Previous studies have overlooked the influence of object pose on these wave components. It is crucial to investigate how these poses affect the reflected wave components because object poses can vary widely and are often unpredictable in real-world scenarios. In this work, we introduce a novel dataset comprising reflected wave components collected from objects made of various materials and observed from various angles. We also show the preliminary evaluations on the performance of machine learning-based material classification on object pose. Our results indicate that the accuracy is consistently high (≥ 91%) for known angles but significantly drops (< 60%) when dealing with unknown angles in most cases. Based on these evaluations, we suggest several directions for future research. Our dataset is available at https://github.com/Nyamotaro/U2R.
Mayuka Kono, Yutaro Hirao, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Nobuchika Sakata, Jun Takamatsu, Kiyoshi Kiyokawa
ICASSP8
2024 First-Person Perspective Induces Stronger Feelings of Awe and Presence Compared to Third-Person Perspective in Virtual Reality
abstract
Awe is a complex emotion described as a perception of vastness and a need for accommodation to integrate new, overwhelming experiences. Virtual Reality (VR) has recently gained attention as a convenient means to facilitate experiences of awe. In VR, a first-person perspective might increase awe due to its immersive nature, while a third-person perspective might enhance the perception of vastness. However, the impact of VR perspectives on experiencing awe has not been thoroughly examined. We created two types of VR scenes: one with elements designed to induce high awe, such as a snowy mountain, and a low awe scene without such elements. We compared first-person and third-person perspectives in each scene. Forty-two participants explored the VR scenes, with their physiological responses captured by electrocardiogram (ECG) and face tracking (FT). Subsequently, participants self-reported their experience of awe (AWE-S) and presence (IPQ) within VR. The results revealed that the first-person perspective induced stronger feelings of awe and presence than the third-person perspective. The findings of this study provide useful guidelines for designing VR content that enhances emotional experiences.
Hiromu Otsubo, Alexander Marquardt, Melissa Steininger, Marvin Lehnort, Felix Dollack, Yutaro Hirao, Monica Perusquía-Hernández, Hideaki Uchiyama, Ernst Kruijff, Bernhard E. Riecke, Kiyoshi Kiyokawa
ICMI11
2024 GlanXR: A Hands-Free Fast Switching System for Virtual Screens
abstract
To date, virtual and augmented reality technologies enable users to view multiple, large virtual screens in their workspaces. However, users must frequently rotate their heads to shift focus among these screens. This paper presents GlanXR, a fast and robust handsfree approach for screen switching in virtual reality. GlanXR incorporates a peripheral interface that remains fixed within the user’s view, in which screens can be dynamically selected based on the user’s eye-head position beyond an adaptive range. Additionally, the user triggers the switch to the screen chosen by making an opposing head rotation in the direction of the eye-head position to minimize false triggers. We conducted an experiment including a fast-switching scenario and a working simulation scenario with 24 participants to assess the effectiveness of GlanXR as compared to a baseline (taskbar), an expansive multi-screen setup, and a gazebased screen selection method. The results indicate that GlanXR facilitates precise screen-switching, minimizes the necessity for head rotation, and allows users to maintain a neutral head position.
Guanghan Zhao, Jason Orlosky, Kiyoshi Kiyokawa, Yuuki Uranishi
ISMAR3
2024 IEEE VR 2024 Steering Committee Message
abstract
The IEEE VR Steering Committee congratulates and offers gratitude to the enormous efforts of the IEEE VR Conference Organizing Committees—with special recognition for the multiple years of hard work of the General Chairs, Carolina Cruz-Neira, Greg Welch, and Xubo Yang. To realize the conference requires a sizeable, motivated team of volunteers to work with the IEEE conference management team. Maintaining the conference’s status as the premier international virtual reality conference is a testament to the commitment of all the organizers and participants that make up the IEEE Virtual Reality community.
Mark Billinghurst, Sabine Coquillart, Kiyoshi Kiyokawa, Gudrun Klinker, Anatole Lécuyer, Betty J. Mohler, Amela Sadagic, J. Edward Swan II, Mary C. Whitton
VR4
2024 Hap'n'Roll: A Scroll-inspired Device for Delivering Diverse Haptic Feedback with a Single Actuator
abstract
Hap’n’Roll is a wearable device that leverages the concept of a scroll to present, with a single motor, tactile sensations of various sizes, shapes, and textures. Hap’n’Roll is composed of two axes, a sheet, and one motor. By changing the number of sheet wraps, the thickness within the user’s hand can be adjusted. Additionally, using holes on the sheet to secure the fingertips, it can present a wide range of sizes and shapes. Unlike typical existing handheld shape-changing devices, Hap’n’Roll is not limited to cylindrical forms. Furthermore, by moving different materials attached on the sheet to the fingertips, it can also express different textures. A user study showed that Hap’n’Roll can convey at least three sizes (small, medium, and large) and four types of shapes (a cylinder, a rectangle, a cone, and a cup), with a shape and size identification accuracy of approx. 76.1%. The identification accuracy for shape alone was approx. 98.5%. Moreover, several applications were developed to showcase the effectiveness of Hap’n’Roll’s mechanism for various haptic feedback.
Hiroki Ota, Daiki Hagimori, Monica Perusquía-Hernández, Naoya Isoyama, Yutaro Hirao, Hideaki Uchiyama, Kiyoshi Kiyokawa
VR7
2024 Retinotopic Foveated Rendering
abstract
Foveated rendering (FR) improves the rendering performance of virtual reality (VR) by allocating fewer computational loads in the peripheral field of view (FOV). Existing FR techniques are built based on the radially symmetric regression model of human visual acuity. However, horizontal-vertical asymmetry (HVA) and vertical meridian asymmetry (VMA) in the cortical magnification factor (CMF) of the human visual system have been evidenced by retinotopy research of neuroscience, suggesting the radially asymmetric regression of visual acuity. In this paper, we begin with functional magnetic resonance imaging (fMRI) data, construct an anisotropic CMF model of the human visual system, and then introduce the first radially asymmetric regression model of the rendering precision for FR applications. We conducted a pilot experiment to adapt the proposed model to VR head-mounted displays (HMDs). A user study demonstrates that retinotopic foveated rendering (RFR) provides participants with perceptually equal image quality compared to typical FR methods while reducing fragments shading by 27.2% averagely, leading to the acceleration of 1/6 for graphics rendering. We anticipate that our study will enhance the rendering performance of VR by bridging the gap between retinotopy research in neuroscience and computer graphics in VR.
Yan Zhang 0101, Keyao You, Xiaodan Hu, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang
VR5
2024 Message from the Editor-in-Chief and from the Associate Editor-in-Chief
abstract
Welcome to the13th IEEE Transactions on Visualization and Computer Graphics (TVCG)special issue on IEEE Virtual Reality and 3D User Interfaces. This volume contains a total of 80 full papers selected for and presented at the IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR 2024), held in Orlando, Florida, USA, from March 16 to 21, 2024.
Han-Wei Shen, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.2
2024 Message from the Editor-in-Chief and from the Associate Editor-in-Chief
abstract
Welcome to the 10th IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue on IEEE International Symposium on Mixed and Augmented Reality (ISMAR). This volume contains a total of 44 full papers selected for and presented at ISMAR 2024, held from October 21 to 25, 2024 in the Greater Seattle Area, USA, in a hybrid mode.
Han-Wei Shen, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.2
2024 HazARdSnap: Gazed-Based Augmentation Delivery for Safe Information Access While Cycling
abstract
During cycling activities, cyclists often monitor a variety of information such as heart rate, distance, and navigation using a bike-mounted phone or cyclocomputer. In many cases, cyclists also ride on sidewalks or paths that contain pedestrians and other obstructions such as potholes, so monitoring information on a bike-mounted interface can slow the cyclist down or cause accidents and injury. In this article, we present HazARdSnap, an augmented reality-based information delivery approach that improves the ease of access to cycling information and at the same time preserves the user's awareness of hazards. To do so, we implemented real-time outdoor hazard detection using a combination of computer vision and motion and position data from a head mounted display (HMD). We then developed an algorithm that snaps information to detected hazards when they are also viewed so that users can simultaneously view both rendered virtual cycling information and the real-world cues such as depth, position, time to hazard, and speed that are needed to assess and avoid hazards. Results from a study with 24 participants that made use of real-world cycling and virtual hazards showed that both HazARdSnap and forward-fixed augmented reality (AR) user interfaces (UIs) can effectively help cyclists access virtual information without having to look down, which resulted in fewer collisions (51% and 43% reduction compared to baseline, respectively) with virtual hazards.
Guanghan Zhao, Jason Orlosky, Joseph L. Gabbard, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.4
2023 Effects of Visual Presentation Near the Mouth on Cross-Modal Effects of Multisensory Flavor Perception and Ease of Eating
abstract
Various studies have suggested that altering the appearance of food can impact multisensory flavor perception. The cross-modal effect of such visual changes on gustation may allow for the presentation of food tastes that are difficult to express with simple combinations of taste stimuli. This cross-modal effect of visual changes on gustation holds potential for applications in gustatory displays. However, the current limitation of existing Head-Mounted Displays (HMDs) is their restricted vertical Field of View (FoV), which prohibits the display of images near the mouth while eating. This limitation may impede the cross-modal effect of visual changes on multisensory flavor perception. Additionally, the lack of visibility around the mouth area challenges the ease of eating. To address these issues, we design a Video See-Through (VST)-HMD with an expanded vertical FoV (approx. 100 [deg]). Using the HMD, we investigated how presenting visual information near the mouth affects the cross-modal effects of flavor perception and ease of eating. In our experiment, machine learning techniques were utilized to alter the appearance of food. However, the result showed no significant differences in the amount of cross-modal effects or the ease of eating between the groups with and without visual information near the mouth. As a discussion of this result, the participants may not direct their visual attention to the food when they put the food in their mouths. The experiment also examined whether visual changes alter the taste as well as the smell and texture of the food. The findings demonstrated that visual changes could present the smell and texture of the food following the modifications. This result was confirmed irrespective of the visibility near the mouth.
Kizashi Nakano, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Kiyoshi Kiyokawa
ISMAR5
2023 A Two-layer Haptic Device for Presenting a Wide Range of Softness and Hardness Using a Pneumatic Balloon and a Mechanical Piston
abstract
Although a variety of haptic devices are used for virtual reality (VR) and augmented reality (AR) experiences, few can present a wide range of softness-hardness of the surface of the virtual objects. We propose a haptic device that can present a wide range of softness-hardness by using a two-layered structure consisting of a pneumatic balloon and a mechanical piston. Through a series of user studies, we confirmed that the prototype can present five levels of softness and three levels of hardness, and that the prototype device improves the VR experience in terms of realism, enjoyment, and comfort for virtual objects with a variety of softness/hardness.
Takuya Sasaki, Daiki Hagimori, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Kiyoshi Kiyokawa, Yoshihiro Kuroda
RO-MAN6
2023 IEEE VR 2023 Steering Committee Message
abstract
The IEEE VR Steering Committee congratulates and offers gratitude to the enormous efforts of the IEEE VR Conference Organizing Committees — with special recognition for the multiple years of hard work of the General Chairs, Xubo Yang, Kun Zhou, Tobias Langlotz, and Stephan Lukosch. To realize the conference requires a sizeable, motivated team of volunteers to work with the IEEE conference management team. Maintaining the conference's status as the premier international virtual reality conference is a testament to the commitment of all the organizers and participants that make up the IEEE Virtual Reality community.
Mark Billinghurst, Sabine Coquillart, Kiyoshi Kiyokawa, Gudrun Klinker, Anatole Lécuyer, Betty J. Mohler, Amela Sadagic, J. Edward Swan II, Mary C. Whitton
VR4
2023 I'm Transforming! Effects of Visual Transitions to Change of Avatar on the Sense of Embodiment in AR
abstract
Virtual avatars are more and more often featured in Virtual Reality (VR) and Augmented Reality (AR) applications. When embodying a virtual avatar, one may desire to change of appearance over the course of the embodiment. However, switching suddenly from one appearance to another can break the continuity of the user experience and potentially impact the sense of embodiment (SoE), especially when the new appearance is very different. In this paper, we explore how applying smooth visual transitions at the moment of the change can help to maintain the SoE and benefit the general user experience. To address this, we implemented an AR system allowing users to embody a regular-shaped avatar that can be transformed into a muscular one through a visual effect. The avatar's transformation can be triggered either by the user through physical action (“active” transition), or automatically launched by the system (“passive” transition). We conducted a user study to evaluate the effects of these two types of transformations on the SoE by comparing them to control conditions where there was no visual feedback of the transformation. Our results show that changing the appearance of one's avatar with an active transition (with visual feedback), compared to a passive transition, helps to maintain the user's sense of agency, a component of the SoE. They also partially suggest that the Proteus effects experienced during the embodiment were enhanced by these transitions. Therefore, we conclude that visual effects controlled by the user when changing their avatar's appearance can benefit their experience by preserving the SoE and intensifying the Proteus effects.
Riku Otono, Adélaïde Genay, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Martin Hachet, Anatole Lécuyer, Kiyoshi Kiyokawa
VR8
2023 IEEE VR 2023 Introducing the Special Issue
abstract
Welcome to the 12th IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue on IEEE Virtual Reality and 3D User Interfaces. This volume contains a total of 61 full papers selected for and presented at the IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR 2023), held in a hybrid style physically in Shanghai, China and virtually online, from March 25 to 29, 2023.
Han-Wei Shen, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.2
2023 Message from the Editor-in-Chief and from the Associate Editor-in-Chief
abstract
Welcome to the November 2023 issue of theIEEE Transactions on Visualization and Computer Graphics (TVCG). This issue contains selected papers accepted at the IEEE International Symposium on Mixed and Augmented Reality (ISMAR). The conference took place from October 16 to 20, 2023 in Sydney, Australia, in a hybrid mode.
Han-Wei Shen, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.2
2023 Add-on Occlusion: Turning Off-the-Shelf Optical See-through Head-mounted Displays Occlusion-capable
abstract
The occlusion-capable optical see-through head-mounted display (OC-OSTHMD) is actively developed in recent years since it allows mutual occlusion between virtual objects and the physical world to be correctly presented in augmented reality (AR). However, implementing occlusion with the special type of OSTHMDs prevents the appealing feature from the wide application. In this paper, a novel approach for realizing mutual occlusion for common OSTHMDs is proposed. A wearable device with per-pixel occlusion capability is designed. OSTHMD devices are upgraded to be occlusion-capable by attaching the device before optical combiners. A prototype with HoloLens 1 is built. The virtual display with mutual occlusion is demonstrated in real-time. A color correction algorithm is proposed to mitigate the color aberration caused by the occlusion device. Potential applications, including the texture replacement of real objects and the more realistic semi-transparent objects display, are demonstrated. The proposed system is expected to realize a universal implementation of mutual occlusion in AR.
Yan Zhang 0101, Xiaodan Hu, Kiyoshi Kiyokawa, Xubo Yang
IEEE Trans. Vis. Comput. Graph.3
2022 An Object Synthesis Method to Enhance Visuo-Haptic Consistency
abstract
The sense of reality is enhanced by presenting appropriate haptic feedback when the user interacts with a virtual object in virtual reality (VR). To present appropriate feedback, we often use a real object that resembles the virtual one to manipulate. However, such a real object is not always available. The user may feel a visuohaptic inconsistency between real and virtual objects when their shapes are different. To alleviate such an inconsistency, we propose a novel object synthesis method that combines the shape of the real object that the user manipulates in reality and the shape of the virtual object which was to be presented to the user in VR. In other words, this synthesized object, a chimera object, is created by transforming the part of the virtual object that the user would touch into a shape that is similar to the corresponding part of the real object while maintaining the other parts of the virtual object intact. The expected haptic sensation from the appearance of our chimera object is more consistent with the one produced by the real object. Therefore, the visuo-haptic inconsistency is expected to alleviate in VR, compared to the original virtual object. In addition, we propose an interactive system to support the design of a chimera object for users. To investigate the effectiveness of our method, we conducted two user studies. The first experiment confirmed that our proposed chimera object helps enhance visuo-haptic consistency. The second experiment confirmed that our system was effective for chimera object creation by users with acceptable system usability.
Naoya Fukumoto, Naoya Isoyama, Hideaki Uchiyama, Nobuchika Sakata, Kiyoshi Kiyokawa
ISMAR5
2022 Recent advances in vision-based indoor navigation: A systematic literature review
Dawar Khan, Zhanglin Cheng, Hideaki Uchiyama, Sikandar Ali 0002, Muhammad Asshad, Kiyoshi Kiyokawa
Comput. Graph.6
2022 VGTC Virtual Reality Technical Achievement Award
abstract
The 2022 IEEE VGTC Virtual Reality Technical Achievement Award goes to Kiyoshi Kiyokawa of the Nara Institute of Science and Technology, in recognition of his pioneering research contributions in the development of advanced head mounted display systems, vision augmentation and assistive interfaces, immersive modeling, collaborative virtual and augmented reality, seamless transitional interfaces, and multimodal interfaces.
Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.1
2022 VGTC Virtual Reality Service Award
abstract
The 2022 IEEE VGTC Virtual Reality Service Award goes to Kiyoshi Kiyokawa of Nara Institute of Science and Technology in recognition of his dedication, support, and many years of service contributions to the VR/AR academic community in operational capacity and in numerous leadership roles. Specifically, Prof. Kiyokawa has served on Steering Committees of IEEE VR, IEEE ISMAR, IEEE 3DUI, and several other conferences and has chaired numerous conferences successfully, including IEEE VR 2019 in Osaka, Japan. The IEEE VGTC is pleased to award Kiyoshi Kiyokawa the 2022 Virtual Reality Service Award.
Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.1
2022 Multisensory Proximity and Transition Cues for Improving Target Awareness in Narrow Field of View Augmented Reality Displays
abstract
Augmented reality applications allow users to enrich their real surroundings with additional digital content. However, due to the limited field of view of augmented reality devices, it can sometimes be difficult to become aware of newly emerging information inside or outside the field of view. Typical visual conflicts like clutter and occlusion of augmentations occur and can be further aggravated especially in the context of dense information spaces. In this article, we evaluate how multisensory cue combinations can improve the awareness for moving out-of-view objects in narrow field of view augmented reality displays. We distinguish between proximity and transition cues in either visual, auditory or tactile manner. Proximity cues are intended to enhance spatial awareness of approaching out-of-view objects while transition cues inform the user that the object just entered the field of view. In study 1, user preference was determined for 6 different cue combinations via forced-choice decisions. In study 2, the 3 most preferred modes were then evaluated with respect to performance and awareness measures in a divided attention reaction task. Both studies were conducted under varying noise levels. We show that on average the Visual-Tactile combination leads to 63% and Audio-Tactile to 65% faster reactions to incoming out-of-view augmentations than their Visual-Audio counterpart, indicating a high usefulness of tactile transition cues. We further show a detrimental effect of visual and audio noise on performance when feedback included visual proximity cues. Based on these results, we make recommendations to determine which cue combination is appropriate for which application.
Christina Trepkowski, Alexander Marquardt, David Eibich, Yusuke Shikanai, Jens Maiero, Kiyoshi Kiyokawa, Ernst Kruijff, Johannes Schöning, Peter König
IEEE Trans. Vis. Comput. Graph.6
2021 ModularHMD: A Reconfigurable Mobile Head-Mounted Display Enabling Ad-hoc Peripheral Interactions with the Real World
abstract
We propose ModularHMD, a new mobile head-mounted display concept, which adopts a modular mechanism and allows a user to perform ad-hoc peripheral interaction with real-world devices or people during VR experiences. ModularHMD is comprised of a central HMD and three removable module devices installed in the periphery of the HMD cowl. Each module has four main states: occluding, extended VR view, video see-through (VST), and removed/reused. Among different combinations of module states, a user can quickly setup the necessary HMD forms, functions, and real-world visions for ad-hoc peripheral interactions without removing the headset. For instance, an HMD user can see her surroundings by switching a module into the VST mode. She can also physically remove a module to obtain direct peripheral visions of the real world. The removed module can be reused as an instant interaction device (e.g., touch keyboards) for subsequent peripheral interactions. Users can end the peripheral interaction and revert to a full VR experience by re-mounting the module. We design ModularHMD’s configuration and peripheral interactions with real-world objects and people. We also implement a proof-of-concept prototype of ModularHMD to validate its interactions capabilities through a user study. Results show that ModularHMD is an effective solution that enables both immersive VR and ad-hoc peripheral interactions.
Isamu Endo, Kazuki Takashima, Maakito Inoue, Kazuyuki Fujita, Kiyoshi Kiyokawa, Yoshifumi Kitamura
UIST5
2021 Foreword to the special section on 2019 international conference on cyberworlds (Cyberworlds 2019)
Takayuki Itoh, Kiyoshi Kiyokawa
Comput. Graph.2
2021 Computational Phase-Modulated Eyeglasses
abstract
We present computational phase-modulated eyeglasses, a see-through optical system that modulates the view of the user using phase-only spatial light modulators (PSLM). A PSLM is a programmable reflective device that can selectively retardate, or delay, the incoming light rays. As a result, a PSLM works as a computational dynamic lens device. We demonstrate our computational phase-modulated eyeglasses with either a single PSLM or dual PSLMs and show that the concept can realize various optical operations including focus correction, bi-focus, image shift, and field of view manipulation, namely optical zoom. Compared to other programmable optics, computational phase-modulated eyeglasses have the advantage in terms of its versatility. In addition, we also presents some prototypical focus-loop applications where the lens is dynamically optimized based on distances of objects observed by a scene camera. We further discuss the implementation, applications but also discuss limitations of the current prototypes and remaining issues that need to be addressed in future research.
Yuta Itoh 0001, Tobias Langlotz, Stefanie Zollmann, Daisuke Iwai, Kiyoshi Kiyokawa, Toshiyuki Amano
IEEE Trans. Vis. Comput. Graph.5
2021 Head-Mounted Display with Increased Downward Field of View Improves Presence and Sense of Self-Location
abstract
Common existing head-mounted displays (HMDs) for virtual reality (VR) provide users with a high presence and embodiment. However, the field of view (FoV) of a typical HMD for VR is about 90 to 110 [deg] in the diagonal direction and about 70 to 90 [deg] in the vertical direction, which is narrower than that of humans. Specifically, the downward FoV of conventional HMDs is too narrow to present the user avatar's body and feet. To address this problem, we have developed a novel HMD with a pair of additional display units to increase the downward FoV by approximately 60 ( 10+50) [deg]. We comprehensively investigated the effects of the increased downward FoV on the sense of immersion that includes presence, sense of self-location (SoSL), sense of agency (SoA), and sense of body ownership (SoBO) during VR experience and on patterns of head movements and cybersickness as its secondary effects. As a result, it was clarified that the HMD with an increased FoV improved presence and SoSL. Also, it was confirmed that the user could see the object below with a head movement pattern close to the real behavior, and did not suffer from cybersickness. Moreover, the effect of the increased downward FoV on SoBO and SoA was limited since it was easier to perceive the misalignment between the real and virtual bodies.
Kizashi Nakano, Naoya Isoyama, Diego Monteiro 0001, Nobuchika Sakata, Kiyoshi Kiyokawa, Takuji Narumi
IEEE Trans. Vis. Comput. Graph.5
2020 Super Wide-view Optical See-through Head Mounted Displays with Per-pixel Occlusion Capability
abstract
Augmented reality (AR) has been widely used that combines human with the digital world tightly at an unprecedented level, and various types of optical see-through head-mounted displays (OSTHMDs) have been actively developed to meet the requirement of everyday AR use. Correct mutual occlusion between real and virtual objects is often necessary for displaying realistic virtual images for users in the AR application scenarios. Some optical designs have been proposed to realize mutual occlusion by means of an OSTHMD in recent years. However, all of them support a limited field of view (FOV) that is much narrower than that of a natural human. The main limiting factor of the FOV of general OSTHMDs is the limited numerical aperture (NA) of the lenses. To address the problem, we propose an OSTHMD based on the double ellipsoidal mirror structure to avoid a stack of lenses and to achieve a wide FOV close to that of the naked eye with small distortion. A pair of imaging lenses are carefully arranged, then assembled between the two ellipsoidal mirrors with a pinhole mask to improve image quality. An experiment using a monocular prototype shows that a sharp see through view is achieved which remains clear regardless of the focus of the eye-simulating camera. An image undistortion algorithm is also developed to obtain a full-view display, enabling a virtual image to be displayed spanning a super wide FOV of H160°×V74°. Finally, per-pixel mutual occlusion of a wide FOV of H122°×V74° is realized by placing a spatial light modulator (SLM) in front of the entrance pupil.
Yan Zhang 0101, Naoya Isoyama, Nobuchika Sakata, Kiyoshi Kiyokawa, Hong Hua
ISMAR4
2019 Combining Tendon Vibration and Visual Stimulation Enhances Kinesthetic Illusions
abstract
In recent years, human augmentation has attracted much attention. One type of human augmentation, motion augmentation makes perceived motion larger than in reality, and it can be used for a variety of applications such as rehabilitation of motor functions of stroke patients and a more realistic experience in virtual reality (VR) such as redirected walking (RDW). However, as augmented motion becomes larger than the real motion, a variety of senses that accompany will be more inconsistent with those perceived from somatic sensations, which will cause a severe sense of discomfort. To address the problem, we focus on kinesthetic illusions that are psychological phenomena where a person feels as if his or her own body is moving. Kinesthetic illusions are expected to fill the gap between the intended augmented motion and perceived physical motion. However, it has not been explored if and how large kinesthetic illusions are produced while a user is moving their limbs voluntarily in VR. To expand the knowledge on kinesthetic illusions, we have conducted two user studies on the impact of tendon vibration and visual stimuli on kinesthetic illusions. First experiment confirmed that the perceived elbow angle becomes larger than the actual angle when presented with tendon vibration. Second experiment revealed that the increase of the perceived elbow angle was about 20 degrees when both tendon vibration and visual stimuli were presented whereas it was about 10 degrees when only visual stimuli were presented. Through these experiments, it has been confirmed that combining tendon vibration and visual stimulation enhances kinesthetic illusions.
Daiki Hagimori, Naoya Isoyama, Shunsuke Yoshimoto, Nobuchika Sakata, Kiyoshi Kiyokawa
CW5
2019 DeepTaste: Augmented Reality Gustatory Manipulation with GAN-Based Real-Time Food-to-Food Translation
abstract
We have been studying augmented reality (AR)-based gustatory manipulation interfaces and previously proposed a gustatory manipulation interface using generative adversarial network (GAN)-based real time image-to-image translation. Unlike three-dimensional (3D) food model-based systems that only change the color or texture pattern of a particular type of food in an inflexible manner, our GAN-based system changes the appearance of food into multiple types of food in real time flexibly, dynamically, and interactively. In the present paper, we first describe in detail a user study on a vision-induced gustatory manipulation system using a 3D food model and report its successful experimental results. We then summarize identified problems of the 3D model-based system and describe implementation details of the GAN-based system. We finally report in detail the main user study in which we investigated the impact of the GAN-based system on gustatory sensations and food recognition when somen noodles were turned into ramen noodles or fried noodles, and steamed rice into curry and rice or fried rice. The experimental results revealed that our system successfully manipulates gustatory sensations to some extent and that the effectiveness seems to depend on the original and target types of food as well as the experience of each individual with the food.
Kizashi Nakano, Daichi Horita, Nobuchika Sakata, Kiyoshi Kiyokawa, Keiji Yanai, Takuji Narumi
ISMAR4
2019 Tendon Vibration Increases Vision-induced Kinesthetic IIIusions in a Virtual Environment
abstract
In virtual reality (VR) systems, a user avatar is typically manipulated based on actual body motion in order to present natural immersive sensations resulted from proprioceptors. It is useful to be able to feel a large body motion in a virtual environment while it is actually smaller in the real environment. As a method to achieve this, we have been working on kinesthetic illusions induced by tendon vibration. In this article, we report on a user study focusing on the effects of a combination of visual stimulus and tendon vibration. In the user study, we measured subjects' perceived elbow angles while they observed the corresponding virtual arm bending to 90 degrees, with different rotation amplification ratios between 1.0 and 1.5, with and without tendon vibration on the wrist joint. The results show that the perceived angle is increased by up to 8 degrees when tendon vibration is applied, roughly in proportion to the rotation amplification ratio. Our study suggests that the vision-induced kinesthetic illusion is further increased by tendon vibration, and that our technique can be applied to VR applications to make users feel larger body motion than that in the real environment.
Daiki Hagimori, Shunsuke Yoshimoto, Nobuchika Sakata, Kiyoshi Kiyokawa
VR4
2019 Augmented Concentration: Concentration Improvement by Visual Noise Reduction with a Video See-Through HMD
abstract
We propose a concentration improvement technique by using a video see-through head mounted display (HMD). Our technique reduces the visual noise or disturbing visual stimulus, such as moving objects near the central visual field, by lowering its visual saliency in real-time. Earphones and noise cancelling headphones are often used to shutdown auditory noise from surroundings when we need to concentrate on the job. Several studies have proven the effectiveness of such noise reduction on improving the concentration and learning efficiencies. We apply this analogy to the visual noise and propose a visual noise reduction HMD to improve concentration. In this article, we report on two preliminary user studies we conducted to investigate the effectiveness of our method. In the prototype system, we manually specify a part of the visual field as the working area and apply grayscale and blur filters outside it in real-time. We compared two conditions “no effect” and “strong blur” (for Exp 1) or “weak blur” (for Exp 2), using a simple math task. Our preliminary results show that visual noise reduction improves the task completion time by 10% for Exp 1 and 7.9% for Exp 2 on average.
Masaki Koshi, Nobuchika Sakata, Kiyoshi Kiyokawa
VR3
2019 Enchanting Your Noodles: GAN-based Real-time Food-to-Food Translation and Its Impact on Vision-induced Gustatory Manipulation
abstract
We propose a novel gustatory manipulation interface which utilizes the cross-modal effect of vision on taste elicited with augmented reality (AR)-based real-time food appearance modulation using a generative adversarial network (GAN). Unlike existing systems which only change color or texture pattern of a particular type of food in an inflexible manner, our system changes the appearance of food into multiple types of food in real-time flexibly, dynamically and interactively in accordance with the deformation of the food that the user is actually eating by using GAN-based image-to-image translation. The experimental results reveal that our system successfully manipulates gustatory sensations to some extent and that the effectiveness depends on the original and target types of food as well as each user's food experience.
Kizashi Nakano, Kiyoshi Kiyokawa, Daichi Horita, Keiji Yanai, Nobuchika Sakata, Takuji Narumi
VR2
2019 Enchanting Your Noodles: A Gustatory Manipulation Interface by Using GAN-based Real-time Food-to-Food Translation
abstract
In this demonstration, we present a novel gustatory manipulation interface which utilizes the cross-modal effect of vision on taste elicited with real-time food appearance modulation using a generative adversarial network (GAN). Unlike existing systems which only change color or texture pattern of a particular type of food in an inflexible manner, our system changes the appearance of food into multiple types of food in real-time flexibly, dynamically and interactively in accordance with the deformation of the food that the user is actually eating by using GAN-based image-to-image translation. Our system can turn somen noodles into ramen noodles or fried noodles, or steamed rice into curry and rice or fried rice. Users of our demonstration system will taste what is visually presented to some extent rather than what they are actually eating.
Kizashi Nakano, Kiyoshi Kiyokawa, Daichi Horita, Keiji Yanai, Nobuchika Sakata, Takuji Narurni
VR2
2019 Speech-Driven Facial Animation by LSTM-RNN for Communication Use
abstract
The goal of this research is developing a system that a rich facial animation can be used in communication is generated from only speech. Generally, a source of the generating facial animation is a camera. Using cameras as an input source, it causes limitations of the angle of view of the camera or problems that cannot be aware of the human face, depending on the orientation of the face. Therefore, it is reasonable for developing a system for generating a facial animation using only voice. In this study, we generate facial expressions from only speech using LSTM-RNN. Comparing 3 patterns of speech analysis data, we showed that the proposed method using A-weighting is effective for facial expression estimation.
Ryosuke Nishimura, Nobuchika Sakata, Kensuke Harada, Tomu Tominaga, Kiyoshi Kiyokawa, Yoshinori Hijikata
VR5
2019 Evaluation of Pointing Interfaces with an AR Agent for Multi-section Information Guidance
abstract
In educational settings such as art galleries or museums, Augmented Reality (AR) has the potential to provide detailed information about exhibits. However, dealing with items that contain information in multiple sections or areas is still a significant challenge. For example, a large painting may contain many minute details, which requires a system that can explain its broader features rather than just a generic description. To address this challenge, we introduce an AR guidance system that uses an embodied agent to point out items and explain each piece and part of exhibit items in detail. We also designed and tested 3 different pointing interfaces for the embodied agent: gesture only, gesture with a dot laser, and gesture with line laser. To evaluate this interface, we conducted a user experiment simulating painting guidance to test interest and exhibit memory. During the experiment, the agent pointed to various areas of interest in the painting and provided a detailed description to participants. The result shows that the search times for target positions were the fastest with the line laser. However, no particular interface outperformed others in memory recall of exhibit content.
Nattaon Techasarntikul, Tomohiro Mashita, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura, Jason Orlosky, Kiyoshi Kiyokawa
VR7
2019 Foreword to the Special Section on Cyberworlds 2018
Kiyoshi Kiyokawa
Comput. Graph.1
2019 A Comparison of Adaptive View Techniques for Exploratory 3D Drone Teleoperation
abstract
Drone navigation in complex environments poses many problems to teleoperators. Especially in three dimensional (3D) structures such as buildings or tunnels, viewpoints are often limited to the drone’s current camera view, nearby objects can be collision hazards, and frequent occlusion can hinder accurate manipulation. To address these issues, we have developed a novel interface for teleoperation that provides a user with environment-adaptive viewpoints that are automatically configured to improve safety and provide smooth operation. This real-time adaptive viewpoint system takes robot position, orientation, and 3D point-cloud information into account to modify the user’s viewpoint to maximize visibility. Our prototype uses simultaneous localization and mapping (SLAM) based reconstruction with an omnidirectional camera, and we use the resulting models as well as simulations in a series of preliminary experiments testing navigation of various structures. Results suggest that automatic viewpoint generation can outperform first- and third-person view interfaces for virtual teleoperators in terms of ease of control and accuracy of robot operation.
John Thomason, Photchara Ratsamee, Jason Orlosky, Kiyoshi Kiyokawa, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura
ACM Trans. Interact. Intell. Syst.4
2019 Light Attenuation Display: Subtractive See-Through Near-Eye Display via Spatial Color Filtering
abstract
We present a display for optical see-through near-eye displays based on light attenuation, a new paradigm that forms images by spatially subtracting colors of light. Existing optical see-through head-mounted displays (OST-HMDs) form virtual images in an additive manner-they optically combine the light from an embedded light source such as a microdisplay into the users' field of view (FoV). Instead, our light attenuation display filters the color of the real background light pixel-wise in the users' see-through view, resulting in an image as a spatial color filter. Our image formation is complementary to existing light-additive OST-HMDs. The core optical component in our system is a phase-only spatial light modulator (PSLM), a liquid crystal module that can control the phase of the light in each pixel. By combining PSLMs with polarization optics, our system realizes a spatially programmable color filter. In this paper, we introduce our optics design, evaluate the spatial color filter, consider applications including image rendering and FoV color control, and discuss the limitations of the current prototype.
Yuta Itoh 0001, Tobias Langlotz, Daisuke Iwai, Kiyoshi Kiyokawa, Toshiyuki Amano
IEEE Trans. Vis. Comput. Graph.4
2019 The Influence of Label Design on Search Performance and Noticeability in Wide Field of View Augmented Reality Displays
abstract
In Augmented Reality (AR), search performance for outdoor tasks is an important metric for evaluating the success of a large number of AR applications. Users must be able to find content quickly, labels and indicators must not be invasive but still clearly noticeable, and the user interface should maximize search performance in a variety of conditions. To address these issues, we have set up a series of experiments to test the influence of virtual characteristics such as color, size, and leader lines on the performance of search tasks and noticeability in both real and simulated environments. We evaluate two primary areas, including 1) the effects of peripheral field of view (FOV) limitations and labeling techniques on target acquisition during outdoor mobile search, and 2) the influence of local characteristics such as color, size, and motion on text labels over dynamic backgrounds. The first experiment showed that limited FOV will severely limit search performance, but that appropriate placement of labels and leaders within the periphery can alleviate this problem without interfering with walking or decreasing user comfort. In the second experiment, we found that different types of motion are more noticeable in optical versus video see-through displays, but that blue coloration is most noticeable in both. Results can aid in designing more effective view management techniques, especially for wider field of view displays.
Ernst Kruijff, Jason Orlosky, Naohiro Kishishita, Christina Trepkowski, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.5
2018 Obstacle Avoidance Method in Real Space for Virtual Reality Immersion
abstract
Typical Head-Mounted Displays (HMDs) that provide a highly immersive Virtual Reality (VR) experience make any interaction between a user and real space difficult by occluding the user's entire field of view. Video see-through type HMDs can solve this problem by superimposing real-space information on the VR environment. The existing method of supporting interactions with the real space is superimposition of boundary lines of the real space on the virtual space in the HMD. However, overlaying the boundary lines on the entire field of view may reduce the user's immersive feeling. In this paper, we propose two methods to support interactions with the real world while playing immersive VR games without reducing the user's immersive feeling as much as possible, even when the user wanders. The first method is to superimpose a 3D point cloud of real space around the user on the virtual space in the HMD. The second method is to deploy familiar objects (e.g., furniture in his/her room) in the virtual space in the HMD. The user traces the familiar objects as subgoals to reach the goal. We implement the two methods and conduct a user study to compare interaction performance. As a result of the user study, we find that the second method provides better spatial information about the real space without reducing the user's immersive feeling, compared to the existing method.
Kohei Kanamori, Nobuchika Sakata, Tomu Tominaga, Yoshinori Hijikata, Kensuke Harada, Kiyoshi Kiyokawa
ISMAR6
2018 IntelliPupil: Pupillometric Light Modulation for Optical See-Through Head-Mounted Displays
abstract
In practical use of optical see-through head-mounted displays, users often have to adjust the brightness of virtual content to ensure that it is at the optimal level. Automatic adjustment is still a challenging problem, largely due to the bidirectional nature of the structure of the human eye, complexity of real world lighting, and user perception. Allowing the right amount of light to pass through to the retina requires a constant balance of incoming light from the real world, additional light from the virtual image, pupil contraction, and feedback from the user. While some automatic light adjustment methods exist, none have completely tackled this complex input-output system. As a step towards overcoming this issue, we introduce IntelliPupil, an approach that uses eye tracking to properly modulate augmentation lighting for a variety of lighting conditions and real scenes. We first take the data from a small form factor light sensor and changes in pupil diameter from an eye tracking camera as passive inputs. This data is coupled with user-controlled brightness selections, allowing us to fit a brightness model to user preference using a feed-forward neural network. Using a small amount of training data, both scene luminance and pupil size are used as inputs into the neural network, which can then automatically adjust to a user's personal brightness preferences in real time. Experiments in a high dynamic range AR scenario with varied lighting show that pupil size is just as important as environment light for optimizing brightness and that our system outperforms linear models.
Chang Liu 0081, Alexander Plopski, Kiyoshi Kiyokawa, Photchara Ratsamee, Jason Orlosky
ISMAR3
2018 Force Rendering and its Evaluation of a Friction-Based Walking Sensation Display for a Seated User
abstract
Most existing locomotion devices that represent the sensation of walking target a user who is actually performing a walking motion. Here, we attempted to represent the walking sensation, especially a kinesthetic sensation and advancing feeling (the sense of moving forward) while the user remains seated. To represent the walking sensation using a relatively simple device, we focused on the force rendering and its evaluation of the longitudinal friction force applied on the sole during walking. Based on the measurement of the friction force applied on the sole during actual walking, we developed a novel friction force display that can present the friction force without the influence of body weight. Using performance evaluation testing, we found that the proposed method can stably and rapidly display friction force. Also, we developed a virtual reality (VR) walk-through system that is able to present the friction force through the proposed device according to the avatar's walking motion in a virtual world. By evaluating the realism, we found that the proposed device can represent a more realistic advancing feeling than vibration feedback.
Ginga Kato, Yoshihiro Kuroda, Kiyoshi Kiyokawa, Haruo Takemura
IEEE Trans. Vis. Comput. Graph.3
2018 Preface
abstract
We are pleased to present the technical papers for IEEE VR 2018: the 25th IEEE Conference on Virtual Reality and 3D User Interfaces, held March 18-22, 2018 in Reutlingen, Germany. IEEE VR 2018 features two categories of submissions: (i) VR Journal Papers, and (ii) VR Conference Papers. Both categories have their own program committees, submission processes, and review processes. This year, 178 submissions were submitted to the Journal track from which 29 were accepted as articles to IEEE TVCG (16.3%). All of these will be presented at the IEEE VR 2018 conference, along with 6 additional papers in the VR area that were published in IEEE TVCG during the past year. Furthermore, 6 submission (3.4%) were recommended for a regular issue of TVCG with major revisions with reviewer continuity. Each of the papers in this special issue went through a rigorous two-round review process. All accepted journal papers are published in a special issue of IEEE Transactions on Visualization and Computer Graphics (TVCG). The Proceedings of the IEEE Conference on Virtual Reality and 3D User Interfaces contains all of the accepted conference papers and the poster abstracts.
Kiyoshi Kiyokawa, Frank Steinicke, Bruce H. Thomas, Greg Welch
IEEE Trans. Vis. Comput. Graph.1
2017 Exploring Proxemics for Human-Drone Interaction
abstract
We present a human-centered designed social drone aiming to be used in a human crowd environment. Based on design studies and focus groups, we created a prototype of a social drone with a social shape, face and voice for human interaction. We used the prototype for a proxemic study, comparing the required distance from the drone humans could comfortably accept compared with what they would require for a nonsocial drone. The social shaped design with greeting voice added decreased the acceptable distance markedly, as did present or previous pet ownership, and maleness. We also explored the proximity sphere around humans with a social shaped drone based on a validation study with variation of lateral distance and heights. Both lateral distance and the higher height of 1.8 m compared to the lower height of 1.2 m decreased the required comfortable distance as it approached.
Alexander Yeh, Photchara Ratsamee, Kiyoshi Kiyokawa, Yuuki Uranishi, Tomohiro Mashita, Haruo Takemura, Morten Fjeld, Mohammad Obaid
HAI3
2017 VisMerge: Light Adaptive Vision Augmentation via Spectral and Temporal Fusion of Non-visible Light
abstract
Low light situations pose a significant challenge to individuals working in a variety of different fields such as firefighting, rescue, maintenance and medicine. Tools like flashlights and infrared (IR) cameras have been used to augment light in the past, but they must often be operated manually, provide a field of view that is decoupled from the operator's own view, and utilize color schemes that can occlude content from the original scene. To help address these issues, we present VisMerge, a framework that combines a thermal imaging head mounted display (HMD) and algorithms that temporally and spectrally merge video streams of different light bands into the same field of view. For temporal synchronization, we first develop a variant of the time warping algorithm used in virtual reality (VR), but redesign it to merge video see-through (VST) cameras with different latencies. Next, using computer vision and image compositing we develop five new algorithms designed to merge non-uniform video streams from a standard RGB camera and small form-factor infrared (IR) camera. We then implement six other existing fusion methods, and conduct a series of comparative experiments, including a system level analysis of the augmented reality (AR) time warping algorithm, a pilot experiment to test perceptual consistency across all eleven merging algorithms, and an in-depth experiment on performance testing the top algorithms in a VR (simulated AR) search task. Results showed that we can reduce temporal registration error due to inter-camera latency by an average of 87.04%, that the wavelet and inverse stipple algorithms were perceptually rated the highest, that noise modulation performed best, and that freedom of user movement is significantly increased with visualizations engaged.
Jason Orlosky, Peter Kim, Kiyoshi Kiyokawa, Tomohiro Mashita, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura
ISMAR3
2017 Adaptive View Management for Drone Teleoperation in Complex 3D Structures
abstract
Drone navigation in complex environments poses many problems to teleoperators. Especially in 3D structures like buildings or tunnels, viewpoints are often limited to the drone's current camera view, nearby objects can be collision hazards, and frequent occlusion can hinder accurate manipulation. To address these issues, we have developed a novel interface for teleoperation that provides a user with environment-adaptive viewpoints that are automatically configured to improve safety and smooth user operation. This real-time adaptive viewpoint system takes robot position, orientation, and 3D pointcloud information into account to modify user-viewpoint to maximize visibility. Our prototype uses simultaneous localization and mapping (SLAM) based reconstruction with an omnidirectional camera and we use resulting models as well as simulations in a series of preliminary experiments testing navigation of various structures. Results suggest that automatic viewpoint generation can outperform first and third-person view interfaces for virtual teleoperators in terms of ease of control and accuracy of robot operation.
John Thomason, Photchara Ratsamee, Kiyoshi Kiyokawa, Pakpoom Kriengkomol, Jason Orlosky, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura
IUI3
2017 Monocular focus estimation method for a freely-orienting eye using Purkinje-Sanson images
abstract
We present a method for focal distance estimation of a freely-orienting eye using Purkinje-Sanson (PS) images, which are reflections of light on the inner structures of the eye. Using an infrared camera with a rigidly-fixed LED, our method creates an estimation model based on 3D gaze and the distance between reflections in the PS images that occur on the corneal surface and anterior surface of the eye lens. The distance between these two reflections changes with focus, so we associate that information to the focal distance on a user. Unlike conventional methods that mainly relies on 2D pupil size which is sensitive to scene lighting and the fourth PS image, our method detects the third PS image which is more representative of accommodation. Our feasibility study on a single user with a focal range from 15–45 cm shows that our method achieves mean and median absolute errors of 3.15 and 1.93 cm for a 10-degree viewing angle. The study shows that our method is also tolerant against environment lighting changes.
Yuta Itoh 0001, Jason Orlosky, Kiyoshi Kiyokawa, Toshiyuki Amano, Maki Sugimoto
VR3
2017 Emulation of Physician Tasks in Eye-Tracked Virtual Reality for Remote Diagnosis of Neurodegenerative Disease
abstract
For neurodegenerative conditions like Parkinson's disease, early and accurate diagnosis is still a difficult task. Evaluations can be time consuming, patients must often travel to metropolitan areas or different cities to see experts, and misdiagnosis can result in improper treatment. To date, only a handful of assistive or remote methods exist to help physicians evaluate patients with suspected neurological disease in a convenient and consistent way. In this paper, we present a low-cost VR interface designed to support evaluation and diagnosis of neurodegenerative disease and test its use in a clinical setting. Using a commercially available VR display with an infrared camera integrated into the lens, we have constructed a 3D virtual environment designed to emulate common tasks used to evaluate patients, such as fixating on a point, conducting smooth pursuit of an object, or executing saccades. These virtual tasks are designed to elicit eye movements commonly associated with neurodegenerative disease, such as abnormal saccades, square wave jerks, and ocular tremor. Next, we conducted experiments with 9 patients with a diagnosis of Parkinson's disease and 7 healthy controls to test the system's potential to emulate tasks for clinical diagnosis. We then applied eye tracking algorithms and image enhancement to the eye recordings taken during the experiment and conducted a short follow-up study with two physicians for evaluation. Results showed that our VR interface was able to elicit five common types of movements usable for evaluation, physicians were able to confirm three out of four abnormalities, and visualizations were rated as potentially useful for diagnosis.
Jason Orlosky, Yuta Itoh 0001, Maud Ranchet, Kiyoshi Kiyokawa, John Morgan, Hannes Devos
IEEE Trans. Vis. Comput. Graph.4
2016 Automated Spatial Calibration of HMD Systems with Unconstrained Eye-cameras
abstract
Properly calibrating an optical see-through head-mounted display (OST-HMD) and maintaining a consistent calibration over time can be a very challenging task. Automated methods need an accurate model of both the OST-HMD screen and the user's constantly changing eye-position to correctly project virtual information. While some automated methods exist, they often have restrictions, including fixed eye-cameras that cannot be adjusted for different users.To address this problem, we have developed a method that automatically determines the position of an adjustable eye-tracking camera and its unconstrained position relative to the display. Unlike methods that require a fixed pose between the HMD and eye camera, our framework allows for automatic calibration even after adjustments of the camera to a particular individual's eye and even after the HMD moves on the user's face. Using two sets of IR-LEDs rigidly attached to the camera and OST-HMD frame, we can calculate the correct projection for different eye positions in real time and changes in HMD position within several frames. To verify the accuracy of our method, we conducted two experiments with a commercial HMD by calibrating a number of different eye and camera positions. Ground truth was measured through markers on both the camera and HMD screens, and we achieve a viewing accuracy of 1.66 degrees for the eyes of 5 different experiment participants.
Alexander Plopski, Jason Orlosky, Yuta Itoh 0001, Christian Nitschke, Kiyoshi Kiyokawa, Gudrun Klinker
ISMAR5
2016 OST Rift: Temporally consistent augmented reality with a consumer optical see-through head-mounted display
abstract
We present an off-the-shelf, low-latency Optical See-through Head-Mounted Displays (OST-HMD) for Augmented Reality (AR). Temporally consistent visualization is crucial for realizing immersive AR experiences. This is challenging since it requires both accurate head-tracking and low-latency rendering of AR content. Building a system which meets both constraints usually requires experts on computer vision/graphics and expensive display hardware. This work demonstrates that such high spatio-temporal fidelity is achievable with commodity hardware available today. We build a custom OST-HMD system that consists of a virtual reality HMD, i.e., the Oculus Rift DK2, and half-mirror optics, and adapt the rendering pipeline in order to integrate the OST-HMD calibration framework. An evaluation with a user-perspective camera shows that the system achieves mean temporal error of <;1 ms (95% reduction of the latency from naive, no-predictive rendering), and median spatial error <;0.3° in the viewing angle with maximum error at most 1.0°.
Yuta Itoh 0001, Jason Orlosky, Manuel J. Huber, Kiyoshi Kiyokawa, Gudrun Klinker
VR4
2016 Spatial consistency perception in optical and video see-through head-mounted augmentations
abstract
Correct spatial alignment is an essential requirement for convincing augmented reality experiences. Registration error, caused by a variety of systematic, environmental, and user influences decreases the realism and utility of head mounted display AR applications. Focus is often given to rigorous calibration and prediction methods seeking to entirely remove misalignment error between virtual and real content. Unfortunately, producing perfect registration is often simply not possible. Our goal is to quantify the sensitivity of users to registration error in these systems, and identify acceptability thresholds at which users can no longer distinguish between the spatial positioning of virtual and real objects. We simulate both video see-through and optical see-through environments using a projector system and experimentally measure user perception of virtual content misalignment. Our results indicate that users are less perceptive to rotational errors over all and that translational accuracy is less important in optical see-through systems than in video see-through.
Alexander Plopski, Kenneth R. Moser, Kiyoshi Kiyokawa, J. Edward Swan II, Haruo Takemura
VR3
2015 HapSticks: A novel method to present vertical forces in tool-mediated interactions by a non-grounded rotation mechanism
abstract
Force feedback in tool-mediated interactions with the environment is important for successful performance of complex tasks in our daily life as well as in specialized fields like medicine. Stylus-based haptic devices are studied and used extensively, and most of these devices require either grounding or attachment to the body of the user. Recently, non-grounded haptic devices are getting an increasing attention. In this paper, we propose a novel method to represent the vertical forces that are applied on the tip of a tool: a non-grounded rotation mechanism that mimics the cutaneous sensation that is caused by these tool-tip forces. To evaluate this method, we developed a novel ungrounded haptic device - HapSticks - that renders the sensation of manipulating objects using chopsticks. First, we present the novel mechanism, and test the pressure that it applies on the hand of the user when rendering a force at the tip of the tool in comparison to applying a real force at the tip of the tool. Next, we used the mechanism to build the HapSticks device as an example of an application of the proposed method, and present a psychophysical evaluation of this device in a virtual weight discrimination task.
Ginga Kato, Yoshihiro Kuroda, Ilana Nisky, Kiyoshi Kiyokawa, Haruo Takemura
World Haptics4
2015 Halo Content: Context-aware Viewspace Management for Non-invasive Augmented Reality
abstract
In mobile augmented reality, text and content placed in a user's immediate field of view through a head worn display can interfere with day to day activities. In particular, messages, notifications, or navigation instructions overlaid in the central field of view can become a barrier to effective face-to-face meetings and everyday conversation. Many text and view management methods attempt to improve text viewability, but fail to provide a non-invasive personal experience for the user.
Jason Orlosky, Kiyoshi Kiyokawa, Takumi Toyama, Daniel Sonntag
IUI2
2015 Attention Engagement and Cognitive State Analysis for Augmented Reality Text Display Functions
abstract
Human eye gaze has recently been used as an effective input interface for wearable displays. In this paper, we propose a gaze-based interaction framework for optical see-through displays. The proposed system can automatically judge whether a user is engaged with virtual content in the display or focused on the real environment and can determine his or her cognitive state. With these analytic capacities, we implement several proactive system functions including adaptive brightness, scrolling, messaging, notification, and highlighting, which would otherwise require manual interaction. The goal is to manage the relationship between virtual and real, creating a more cohesive and seamless experience for the user. We conduct user experiments including attention engagement and cognitive state analysis, such as reading detection and gaze position estimation in a wearable display towards the design of augmented reality text display applications. The results from the experiments show robustness of the attention engagement and cognitive state analysis methods. A majority of the experiment participants (8/12) stated the proactive system functions are beneficial.
Takumi Toyama, Daniel Sonntag, Jason Orlosky, Kiyoshi Kiyokawa
IUI4
2015 Collaboration in Augmented Reality
abstract
Augmented Reality (AR) is a technology that allows users to view and interact in real time with virtual images seamlessly superimposed over the real world. AR systems can be used to create unique collaborative experiences. For example, co-located users can see shared 3D virtual objects that they interact with, or a user can annotate the live video view of a remote worker, enabling them to collaborate at a distance. The overall goal is to augment the face-to-face collaborative experience, or to enable remote people to feel that they are virtually co-located. In this special issue on collaboration in augmented reality, we begin with the visions of science fiction authors of future technologies that might significantly improve collaboration, then introduce research articles which describe progress towards these visions, finally we outline a research agenda discussing the work still to be done.
Stephan G. Lukosch, Mark Billinghurst, Leila Alem, Kiyoshi Kiyokawa
Comput. Support. Cooperative Work.4
2015 Guest Editor's Introduction to the Special Section on the International Symposium on Mixed and Augmented Reality 2013
abstract
The articles in this special section were presented at the 2013 IEEE International Symposium on Mixed and Augmented Reality (ISMAR).
Maribeth Gandy Coleman, Simon J. Julier, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.3
2015 ModulAR: Eye-Controlled Vision Augmentations for Head Mounted Displays
abstract
In the last few years, the advancement of head mounted display technology and optics has opened up many new possibilities for the field of Augmented Reality. However, many commercial and prototype systems often have a single display modality, fixed field of view, or inflexible form factor. In this paper, we introduce Modular Augmented Reality (ModulAR), a hardware and software framework designed to improve flexibility and hands-free control of video see-through augmented reality displays and augmentative functionality. To accomplish this goal, we introduce the use of integrated eye tracking for on-demand control of vision augmentations such as optical zoom or field of view expansion. Physical modification of the device's configuration can be accomplished on the fly using interchangeable camera-lens modules that provide different types of vision enhancements. We implement and test functionality for several primary configurations using telescopic and fisheye camera-lens systems, though many other customizations are possible. We also implement a number of eye-based interactions in order to engage and control the vision augmentations in real time, and explore different methods for merging streams of augmented vision into the user's normal field of view. In a series of experiments, we conduct an in depth analysis of visual acuity and head and eye movement during search and recognition tasks. Results show that methods with larger field of view that utilize binary on/off and gradual zoom mechanisms outperform snapshot and sub-windowed methods and that type of eye engagement has little effect on performance.
Jason Orlosky, Takumi Toyama, Kiyoshi Kiyokawa, Daniel Sonntag
IEEE Trans. Vis. Comput. Graph.3
2015 Corneal-Imaging Calibration for Optical See-Through Head-Mounted Displays
abstract
In recent years optical see-through head-mounted displays (OST-HMDs) have moved from conceptual research to a market of mass-produced devices with new models and applications being released continuously. It remains challenging to deploy augmented reality (AR) applications that require consistent spatial visualization. Examples include maintenance, training and medical tasks, as the view of the attached scene camera is shifted from the user's view. A calibration step can compute the relationship between the HMD-screen and the user's eye to align the digital content. However, this alignment is only viable as long as the display does not move, an assumption that rarely holds for an extended period of time. As a consequence, continuous recalibration is necessary. Manual calibration methods are tedious and rarely support practical applications. Existing automated methods do not account for user-specific parameters and are error prone. We propose the combination of a pre-calibrated display with a per-frame estimation of the user's cornea position to estimate the individual eye center and continuously recalibrate the system. With this, we also obtain the gaze direction, which allows for instantaneous uncalibrated eye gaze tracking, without the need for additional hardware and complex illumination. Contrary to existing methods, we use simple image processing and do not rely on iris tracking, which is typically noisy and can be ambiguous. Evaluation with simulated and real data shows that our approach achieves a more accurate and stable eye pose estimation, which results in an improved and practical calibration with a largely improved distribution of projection error.
Alexander Plopski, Yuta Itoh 0001, Christian Nitschke, Kiyoshi Kiyokawa, Gudrun Klinker, Haruo Takemura
IEEE Trans. Vis. Comput. Graph.4
2014 A natural interface for multi-focal plane head mounted displays using 3D gaze
abstract
In mobile augmented reality (AR), it is important to develop interfaces for wearable displays that not only reduce distraction, but that can be used quickly and in a natural manner. In this paper, we propose a focal-plane based interaction approach with several advantages over traditional methods designed for head mounted displays (HMDs) with only one focal plane. Using a novel prototype that combines a monoscopic multi-focal plane HMD and eye tracker, we facilitate interaction with virtual elements such as text or buttons by measuring eye convergence on objects at different depths. This can prevent virtual information from being unnecessarily overlaid onto real world objects that are at a different range, but in the same line of sight. We then use our prototype in a series of experiments testing the feasibility of interaction. Despite only being presented with monocular depth cues, users have the ability to correctly select virtual icons in near, mid, and far planes in 98.6% of cases.
Takumi Toyama, Daniel Sonntag, Jason Orlosky, Kiyoshi Kiyokawa
AVI4
2014 Analysing the effects of a wide field of view augmented reality display on search performance in divided attention tasks
abstract
A wide field of view augmented reality display is a special type of head-worn device that enables users to view augmentations in the peripheral visual field. However, the actual effects of a wide field of view display on the perception of augmentations have not been widely studied. To improve our understanding of this type of display when conducting divided attention search tasks, we conducted an in depth experiment testing two view management methods, in-view and in-situ labelling. With in-view labelling, search target annotations appear on the display border with a corresponding leader line, whereas in-situ annotations appear without a leader line, as if they are affixed to the referenced objects in the environment. Results show that target discovery rates consistently drop with in-view labelling and increase with in-situ labelling as display angle approaches 100 degrees of field of view. Past this point, the performances of the two view management methods begin to converge, suggesting equivalent discovery rates at approximately 130 degrees of field of view. Results also indicate that users exhibited lower discovery rates for targets appearing in peripheral vision, and that there is little impact of field of view on response time and mental workload.
Naohiro Kishishita, Kiyoshi Kiyokawa, Jason Orlosky, Tomohiro Mashita, Haruo Takemura, Ernst Kruijff
ISMAR2
2014 Collaboration in mediated and augmented reality
abstract
In this half-day workshop we will explore how Augmented Reality (AR) and Mediated Reality (MR) can be used to develop radically new types of collaborative experiences that overcome some of the limitations of current conferencing systems. In combination, AR and MR technologies could be used to merge the shared perceived realities of different users as well as enriching their own individual experience in a collaborative task. The goal of the workshop is to bring together researchers who are interested developing collaborative systems using AR and MR technologies. They will build a picture of current and prior research on collaboration in AR and MR as well as set up a common research agenda for work going forward. Topics of the workshop will address open research issues and include but are not restricted to the following: •Case studies on using MR/AR for collaboration •Tools for building collaborative MR/AR systems •Effects of MR/AR on trust, presence, and coordination •Interaction models for collaboration in MR/AR •Tools for collaboration in MR/AR •Collaboration awareness in MR/AR.
Stephan G. Lukosch, Mark Billinghurst, Kiyoshi Kiyokawa, Leila Alem
ISMAR3
2014 Corneal imaging in localization and HMD interaction
abstract
The human eyes perceive our surroundings and are one of, if not our most important sensory organs. Contrary to our other senses the eyes not only perceive but also provide information to a keen observer. However, thus far this has been mainly used to detect reflection of infrared light sources to estimate the user's gaze. The reflection of the visible spectrum on the other hand has rarely been utilized. In this dissertation we want to explore how the analysis of the corneal image can improve currently available eye-related solutions, such as calibration of optical see-through head-mounted devices or eye-gaze tracking and point of regard estimation in arbitrary environments. We also aim to study how corneal imaging can become an alternative for established augmented reality tasks such as tracking and localization.
Alexander Plopski, Kiyoshi Kiyokawa, Haruo Takemura, Christian Nitschke
ISMAR2
2014 The effectiveness of an AR-based context-aware assembly support system in object assembly
abstract
This study evaluates the effectiveness of an AR-based context-aware assembly support system with AR visualization modes proposed in object assembly. Although many AR-based assembly support systems have been proposed, few keep track of the assembly status in real-time and automatically recognize error and completion states at each step. Naturally, the effectiveness of such context-aware systems remains unexplored. Our test-bed system displays guidance information and error detection information corresponding to the recognized assembly status in the context of building block (LEGO) assembly. A user wearing a head mounted display (HMD) can intuitively build a building block structure on a table by visually confirming correct and incorrect blocks and locating where to attach new blocks. We proposed two AR visualization modes, one of them that displays guidance information directly overlaid on the physical model, and another one in which guidance information is rendered on a virtual model adjacent to the real model. An evaluation was conducted to comparatively evaluate these AR visualization modes as well as determine the effectiveness of context-aware error detection. Our experimental results indicate the visualization mode that shows target status next to real objects of concern outperforms the traditional direct overlay under moderate registration accuracy and marker-based tracking.
Bui Minh Khuong, Kiyoshi Kiyokawa, Joseph J. La Viola, Tomohiro Mashita, Haruo Takemura
VR2
2014 Reflectance and light source estimation for indoor AR Applications
abstract
We present an approach which enables real-time augmentation of an environment composed of materials with different texture and reflectance properties without the need of application-specific hardware or extensive preparation. Our solution uses a set of RGB images of a reconstructed model to optimize the reflectance parameters and light location. Each image is decomposed into its specular and diffuse components and we estimate the location of multiple light sources from specular highlights. The environment is stored in a voxel grid and we optimize the reflectance properties and colour of each voxel through inverse rendering. We verify our approach with a simulated environment and present results from a corresponding reconstructed environment.
Alexander Plopski, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
VR3
2014 Message from the Paper Chairs and Guest Editors
abstract
In this special issue of IEEE Transactions on Visualization and Computer Graphics (TVCG), we are pleased to present the long papers from the IEEE Virtual Reality Conference 2014 (IEEE VR 2014), held March 29–April 2, 2014 in Minneapolis, Minnesota, USA.
Sabine Coquillart, Kiyoshi Kiyokawa, J. Edward Swan II, Doug A. Bowman
IEEE Trans. Vis. Comput. Graph.2
2013 Navigation Maps for Virtual Travelers
abstract
Navigation is one of the most fundamental tasks in virtual environments, that requires timely collision detection and collision avoidance mechanisms. It was shown that when navigation is constrained to moving on a terrain, the problem of collision detection can be reduced from 3D to 2D case, by encoding all 3D objects that constitute travel obstacles into a terrain elevation image map, using a dedicated color. The original approach was developed for static environments explored by a single user. The new improved system, presented in this article, is capable of processing collisions for multiple travelers, including autonomous agents, in dynamically changing environments. Implementation details of the new system are presented, with a number of case studies and a discussion of future work.
Andrei Sherstyuk, Kiyoshi Kiyokawa
CW2
2013 3D Collision Detection with 2D Terrain Maps
abstract
Terrain elevation maps can be efficiently used for real-time collision processing for ground-bound travelers Maps. We extend this approach to handle objects that may fly freely in 3D space. The new algorithm makes effective use of allocated terrain image memory and does not require extra storage. It is easy to implement into any travel processing system.
Andrei Sherstyuk, Kiyoshi Kiyokawa, Yuki Yano
CW2
2013 Program chairs
abstract
We are delighted to welcome you to ISMAR 2013, the 12th symposium on Mixed and Augmented Reality! This year's symposium continues a long tradition of ISMAR meetings, a series that itself followed a related series of IWAR, ISMR, and ISAR meetings.
Maribeth Gandy Coleman, Simon J. Julier, Kiyoshi Kiyokawa
ISMAR3
2013 In-situ lighting and reflectance estimations for indoor AR systems
abstract
We introduce an in-situ lighting and reflectance estimation method that does not require specific light probes and/or preliminary scanning. Our method uses images taken from multiple viewpoints while data accumulation and lighting and reflectance estimations run in the background of the primary AR system. As a result, our method requires little in the way of manipulations for image collection because it consists primarily of image processing and optimization. When used, lighting directions and initial optimization values are estimated via image processing. Eventually, the full parameters are obtained by optimization of the differences between real images. This system uses current best parameters because the parameter estimation and input image updates are run independently.
Tomohiro Mashita, Hiroyuki Yasuhara, Alexander Plopski, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR4
2013 Towards intelligent view management: A study of manual text placement tendencies in mobile environments using video see-through displays
abstract
When viewing content in a see-through head mounted display (HMD), displaying readable information is still difficult when text is overlayed onto a changing background or lighted surface. Moving text or content to a more appropriate place on the screen through automation or intelligent algorithms is one viable solution to this kind of issue. However, many of these algorithms fail to act as a human would when placing text in a more appropriate location in real time. In order to improve these text and view management algorithms, we report the results and analysis of an experiment designed to evaluate user tendencies when placing virtual text in the real world through an HMD. In the conducted experiment, 20 users manually overlayed text in real time onto 4 different videos taken from the first-person perspective of a pedestrian. We find that users have a tendency to place overlayed text in locations near the center of the viewing field, gravitating towards a point just below the horizon. Common locations for text overlay such as walls, shaded areas, and pavement are classified and discussed.
Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR2
2013 Management and manipulation of text in dynamic mixed reality workspaces
abstract
Viewing and interacting with text based content safely and easily while mobile has been an issue with see-through displays for many years. For example, in order to effectively use optical see through Head Mounted Displays (HMDs) in constantly changing dynamic environments, variables like lighting conditions, human or vehicular obstructions in a user's path, and scene variation must be dealt with effectively. My PhD research focuses on answering the following questions: 1) What are appropriate methods to intelligently move digital content such as e-mail, SMS messeges, and news articles, throughout the real world? 2) Once a user stops moving, in what way should dynamics of the current workspace change when migrated to a new static environment? 3) Lastly, how can users manipulate mobile content using the fewest number of interactions possible? My strategy for developing solutions to these problems primarily involves automatic or semi-automatic movement of digital content throughout the real world using camera tracking. I have already developed an intelligent text management system that actively manages movement of text in a user's field of view while mobile [11]. I am optimizing and expanding on this type of management system, developing appropriate interaction methodology, and conducting experiments to verify effectiveness, usability, and safety when used with an HMD in various environments.
Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR2
2013 Dynamic text management for see-through wearable and heads-up display systems
abstract
Reading text safely and easily while mobile has been an issue with see-through displays for many years. For example, in order to effectively use optical see through Head Mounted Displays (HMDs) or Heads Up Display (HUD) systems in constantly changing dynamic environments, variables like lighting conditions, human or vehicular obstructions in a user's path, and scene variation must be dealt with effectively.
Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura
IUI2
2013 A next location prediction method for smartphones using blockmodels
abstract
Context aware systems on smart-phones aim to provide useful information by analysing and recognizing users' situations from built-in sensors logs. Especially, predicting user actions is one of the important functions for the context aware systems on smart-phones because this function enables context aware systems to provide proactive and responsive services. Therefore the next location prediction is also an important function for context aware systems. This paper introduces a next location prediction method based on context recognition. In this method, we define a context as combinations of features which are extracted from a set of relational data generated from a phone's sensor logs. We applied Mixed Membership Stochastic Blockmodels (MMSB) to context extraction. We then collected sensor logs of a single user over a period of three months and conducted an evaluation using this collected dataset. An an evaluation using the dataset was conducted and the result shows that 60% of the test dataset ranked in the top 30% of all candidates of the next locations.
Jun Fukano, Tomohiro Mashita, Takahiro Hara, Kiyoshi Kiyokawa, Haruo Takemura, Shojiro Nishio
VR4
2013 Learning system for adapting users with user's state classification by vital sensing
abstract
In ambient information systems, not only extracting human behavior with a sensor network but also adaptive autonomous interaction between the environment and humans is an important function. In this paper, we propose a reinforcement learning methodology for acquiring suitable interaction for each person's daily behavior. This time, we used vital sensors to detect and classify a user's condition. In an experiment, we show the feasibility of the proposed methodology.
Junya Nakase, Koichi Moriyama, Kiyoshi Kiyokawa, Masayuki Numao, Mayumi Oyama-Higa, Satoshi Kurihara
VR3
2013 Pinch-n-Paste: Direct texture transfer interaction in augmented reality
abstract
Our Pinch-n-Paste allows a user to touch or pinch one part of an object, copy and move its texture, and paste it onto another object, directly with his or her hand, in an augmented reality environment. To transfer texture appropriately from one part of an object to another, two texture images are generated by the Least Square Conformal Map (LSCM) technique. Two regions in the texture images corresponding to source and target areas of interest are then obtained using cross-boundary brushes. Target texel values are sampled from corresponding source texels by Moving Least Squares (MLS), and are finally mapped onto the target object. In this poster, we will describe the basic idea, implementation details, and example interaction results and a preliminary user study.
Atsushi Umakatsu, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
VR3
2013 A content search system considering the activity and context of a mobile user
Mayu Iwata, Hiroki Miyamoto, Takahiro Hara, Daijiro Komaki, Kentaro Shimatani, Tomohiro Mashita, Kiyoshi Kiyokawa, Toshiaki Uemukai, Gen Hattori, Shojiro Nishio, Haruo Takemura
Pers. Ubiquitous Comput.7
2012 Workshop 2: Classifying the AR presentation space
abstract
Already 3D visualization environments provide a large design space not being investigated to the same extent as traditional WIMP-spaces. When using this design space in combination with AR, the design space even further grows. Information can not only be presented in a 3D space, AR also puts virtual information in relation to real objects, locations or events. The different properties of presentation in AR need to be investigated to develop a comprehensive set of dimensions of presentation principles.
Steven K. Feiner, Kiyoshi Kiyokawa, Gudrun Klinker, Marcus Dennis, Charles Woodward
ISMAR2
2012 A waist-mounted ProCam system for remote collaboration
abstract
We propose a waist-mounted projector-camera (ProCam) system for asymmetric remote collaboration. A wearable camera is often used to transmit a worker's situation to a remote instructor, however 3D structure of the worker's environment is not always available and the instructor has a minimal flexibility in changing the camera's viewpoint. A stationary 3D measurement system is also commonly used for remote collaboration, however a narrow measurement area and occlusion from a worker's body can be a severe problem. Our waist-mounted ProCam system reconstructs worker's environment in real-time without occlusion from the worker's body. The remote instructor can give instructions simply by drawing annotations on the reconstructed environment on screen, and they are properly projected in front of the worker. Structured-light based reconstruction, vision-based localization, and visual annotation projection, are processed synchronously with a camera and modified shutter glasses so that both the camera and the worker observe only the information they need. Experimental results show that users prefer our system to a stationary ProCam system.
Shigeki Morishima, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR3
2012 Touch-n-Paste: Direct texture transfer interaction in AR environments
abstract
Our Touch-n-Paste allows a user to touch one part of an object, copy and move its texture, and paste it onto another object, directly with his or her hand, in an augmented reality environment. To transfer texture appropriately from one part of an object to another, two texture images are generated by the Least Square Conformal Map (LSCM) technique. Two regions in the texture images corresponding to source and target areas of interest are then obtained using cross-boundary brushes. Target texel values are sampled from corresponding source texels by Moving Least Squares (MLS), and are finally mapped onto the target object.
Atsushi Umakatsu, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR3
2012 Subjective evaluations on perceptual depth of stereo image and effective field of view of a wide-view head mounted projective display with a semi-transparent retro-reflective screen
abstract
We report two user studies on a wearable hyperboloidal head mounted projective display (HHMPD) with a semi-transparent retro-reflective screen. First experiment revealed that a virtual image is perceived at a similar distance as the real image only when the observation distance is within 2.5m with monocular vision, whereas its threshold is further than 3m with stereo (binocular) vision. Second experiment revealed that users are able to identify visual stimuli in the periphery of the visual field up to ±50 degrees in horizontal, while paying attention to a real object in frontal direction.
Duc Nguyen Van, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR3
2012 Owens Luis - A context-aware multi-modal smart office chair in an ambient environment
abstract
This paper introduces a smart office chair, Owens Luis, whose pronunciation has a meaning of “an encouraging chair (****)” in Japanese. For most of the people, office environments are the place where they spend the longest time while awake. To improve the quality of life (QoL) in the office, Owens Luis monitors an office worker's mental and physiological states such as sleepiness and concentration, and controls the working environment by multi-modal displays including a motion chair, a variable color-temperature LED light and a hypersonic directional speaker.
Kiyoshi Kiyokawa, Masahide Hatanaka, Kazufumi Hosoda, Masashi Okada, Hironori Shigeta, Yasunori Ishihara, Fukuhito Ooshita, Hirotsugu Kakugawa, Satoshi Kurihara, Koichi Moriyama
VR1
2012 A content search system for mobile devices based on user context recognition
abstract
People carry around mobile devices all the time in their daily life and get various information from the Internet in various situations. When searching for information (content) by using mobile devices, users' activities (e.g., walking and standing) and their situations (e.g., commuting in the morning and going out downtown in the evening) often change and this change may affect their degree of concentration on the display of mobile devices and their information needs. Therefore, search systems should provide users with an amount of information suitable for their activities and with a type of information suitable for their situations. In this paper, we present the design and implementation of a content search system considering mobile users' activities and situations, which aims to reduce users' load of operations in content searching. Our system recognizes user's activities and switches between two kinds of content search systems according to the user's activity: the location-based content search system runs when the user is standing, while the menu-based content search system runs when the user is walking. Both systems present information based on the user's situation. We also introduce a user's activity recognition method for mobile devices. This method classifies user's activity into standing, walking, and running using the sensors equipped in mobile devices.
Tomohiro Mashita, Daijiro Komaki, Mayu Iwata, Kentaro Shimatani, Hiroki Miyamoto, Takahiro Hara, Kiyoshi Kiyokawa, Haruo Takemura, Shojiro Nishio
VR7
2012 Human activity recognition for a content search system considering situations of smartphone users
abstract
Smart-phone users can search for information about surrounding facilities or a route to their destination. However, it is difficult to get or search for information while walking because of low legibility. To address this problem, users have to stop walking or enlarge the screen. Our previously proposed system for smart-phone switches the information presentation policies in response to the user's context. In this paper we describe our context recognition mechanism for this system. This mechanism estimates user context from sensors embedded in a smart-phone. We use a Support Vector Machine for the context classification and compare four types of feature values consisting of FFT and 3 types of Wavelet Transforms. Experimental results show that recognition rates are 87.2 % with FFT, 90.9 % with Gabor Wavelet, 91.8 % with Haar Wavelet, and 92.1 % with MexicanHat Wavelet.
Tomohiro Mashita, Kentaro Shimatani, Mayu Iwata, Hiroki Miyamoto, Daijiro Komaki, Takahiro Hara, Kiyoshi Kiyokawa, Haruo Takemura, Shojiro Nishio
VR7
2012 Adaptive interactive device control by using reinforcement learning in ambient information environment
abstract
In ambient information systems, not only extracting human behavior by sensor network but also adaptive autonomous interaction between the environment and humans is an important function. In this paper we propose a reinforcement learning framework to extract suitable interaction for each person from daily behavior. In the experiment, we show the feasibility of the proposed methodology.
Junya Nakase, Koichi Moriyama, Kiyoshi Kiyokawa, Masayuki Numao, Mayumi Oyama-Higa, Satoshi Kurihara
VR3
2012 Implementation of a smart office system in an ambient environment
abstract
We propose a smart office system that recognizes office workers' mental and physiological states to improve their quality of life at office. We integrated our systems into a single smart office environment. In this article we show the implementation of the smart office system and the details of each of its components such as I/O devices.
Hironori Shigeta, Junya Nakase, Yuta Tsunematsu, Kiyoshi Kiyokawa, Masahide Hatanaka, Kazufumi Hosoda, Masashi Okada, Yasunori Ishihara, Fukuhito Ooshita, Hirotsugu Kakugawa, Satoshi Kurihara, Koichi Moriyama
VR4
2012 Pinch-n-paste: direct texture transfer interaction in augmented reality
abstract
Our Pinch-n-Paste allows a user to touch or pinch one part of an object, copy and move its texture, and paste it onto an other object, directly with his or her hand, in an augmented reality environment. To transfer texture appropriately from one part of an object to another, two texture images are generated by the Least Square Conformal Map (LSCM) technique. Two regions in the texture images corresponding to source and target areas of interest are then obtained using cross-boundary brushes. Target texel values are sampled from corresponding source texels by Moving Least Squares (MLS), and are finally mapped onto the target object.
Atsushi Umakatsu, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
VRST3
2012 Message from the Paper Chairs and Guest Editors
abstract
The articles in this special issue contain the full paper proceedings of the IEEE Virtual Reality Conference 2012 (IEEE VR 2012), held March 4-8, 2012 in Orange County, California.
Sabine Coquillart, Steven K. Feiner, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.3
2011 Guest Editors' Introduction: Special Section on the IEEE Virtual Reality Conference (VR)
Kiyoshi Kiyokawa, Gudrun Klinker, Benjamin Lok
IEEE Trans. Vis. Comput. Graph.1
2011 A Wide-View Parallax-Free Eye-Mark Recorder with a Hyperboloidal Half-Silvered Mirror and Appearance-Based Gaze Estimation
abstract
In this paper, we propose a wide-view parallax-free eye-mark recorder with a hyperboloidal half-silvered mirror and a gaze estimation method suitable for the device. Our eye-mark recorder provides a wide field-of-view video recording of the user's exact view by positioning the focal point of the mirror at the user's viewpoint. The vertical angle of view of the prototype is 122 degree (elevation and depression angles are 38 and 84 degree, respectively) and its horizontal view angle is 116 degree (nasal and temporal view angles are 38 and 78 degree, respectively). We implemented and evaluated a gaze estimation method for our eye-mark recorder. We use an appearance-based approach for our eye-mark recorder to support a wide field-of-view. We apply principal component analysis (PCA) and multiple regression analysis (MRA) to determine the relationship between the captured images and their corresponding gaze points. Experimental results verify that our eye-mark recorder successfully captures a wide field-of-view of a user and estimates gaze direction with an angular accuracy of around 2 to 4 degree.
Hiroki Mori, Erika Sumiya, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
IEEE Trans. Vis. Comput. Graph.4
2009 A wide-view parallax-free eye-mark recorder with a hyperboloidal half-silvered mirror
abstract
In this paper, we propose a wide-view parallax-free eye-mark recorder with a hyperboloidal half-silvered mirror. Our eye-mark recorder provides a wide field-of-view (FOV) video recording of the user's exact view by positioning the focal point of the mirror at the user's viewpoint. The vertical view angle of the prototype is 122 [deg] (elevation and depression angles are 38 and 84 [deg], respectively) and its horizontal view angle is 116 [deg] (nasal and temporal view angles are 38 and 78 [deg], respectively). We have implemented and evaluated a gaze estimation method for our eyemark recorder. Experimental results have verified that our eye-mark recorder successfully captures a wide FOV of a user and estimates a rough gaze direction.
Erika Sumiya, Tomohiro Mashita, Kiyoshi Kiyokawa, Haruo Takemura
VRST3
2008 Mutual occlusions on table-top displays in mixed reality applications
abstract
This paper describes an approach to dealing with mutual occlusions between virtual and real objects on a table-top display. Display tables use stereoscopy to make virtual content appear to exist in 3 dimensions on or above a table top. The actual image, however, lies on the physical plane of the display table. Any real physical object introduced above this plane therefore obstructs our view of the display surface and disrupts the illusion of the virtual scene. The occlusions result between real objects and the display surface, not between real objects and virtual objects. For the same reason virtual objects cannot occlude real ones. Our approach uses an additional projector located near the user's head to project those parts of virtual objects that should occlude real ones directly onto the real objects. We describe possible applications and limitations of the approach and its current implementation. Despite its limitations, we believe that the proposed approach can significantly improve interaction quality and performance for mixed reality scenarios.
Daniel Kurz, Kiyoshi Kiyokawa, Haruo Takemura
VRST2
2007 A Wide Field-of-view Head Mounted Projective Display using Hyperbolic Half-silvered Mirrors
abstract
The development of a wide field-of-view (FOV) head mounted display (HMD) has been a technological challenge for decades. Previous HMDs tackled this problem using multiple display units (tiling) or multiple curved mirrors. The former approach tends to be expensive and heavy, whereas the latter approach tends to suffer from image distortion and a small exit pupil. In order to provide a wide FOV image with a large exit pupil, the present paper proposes a novel head mounted projective display (HMPD) using a hyperbolic half-silvered mirror, rather than a conventional planar mirror. The first bench-top prototype has successfully shown wide field-of-view projection capability.
Kiyoshi Kiyokawa
ISMAR1
2007 A 2D-3D integrated tabletop environment for multi-user collaboration
abstract
Abstract This paper proposes a novel tabletop display system for natural communication and flexible information sharing. The proposed system is specifically designed to integrate two‐dimensional (2D) and three‐dimensional (3D) user interfaces by using a multi‐user stereoscopic display, IllusionHole. The proposed system takes awareness into consideration and provides both 2D and 3D information and user interfaces. On the display, a number of standard Windows desktop environments are provided as personal workspaces, as well as a shared workspace with a dedicated graphical user interface. In the personal workspaces, users can simultaneously access existing applications and data, and exchange information between personal and shared workspaces. In this way, the proposed system can seamlessly integrate personal, shared, 2D, and 3D workspaces with conventional user interfaces and effectively support communication and information sharing. To demonstrate the capabilities of the proposed display system, a modeling application was implemented. A preliminary experiment confirmed the effectiveness of this system. Copyright © 2006 John Wiley & Sons, Ltd.
Kousuke Nakashima, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
Comput. Animat. Virtual Worlds3
2006 A 2D-3D integrated interface for mobile robot control using omnidirectional images and 3D geometric models
abstract
This paper proposes a novel visualization and interaction technique for remote surveillance using both 2D and 3D scene data acquired by a mobile robot equipped with an omnidirectional camera and an omnidirectional laser range sensor. In a normal situation, telepresence with an egocentric-view is provided using high resolution omnidirectional live video on a hemispherical screen. As depth information of the remote environment is acquired, additional 3D information can be overlaid onto the 2D video image such as passable area and roughness of the terrain in a manner of video see-through augmented reality. A few functions to interact with the 3D environment through the 2D live video are provided, such as path-drawing and path-preview. Path-drawing function allows to plan a robot's path by simply specifying 3D points on the path on screen. Path- preview function provides a realistic image sequence seen from the planned path using a texture-mapped 3D geometric model in a manner of virtualized reality. In addition, a miniaturized 3D model is overlaid on the screen providing an exocentric view, which is a common technique in virtual reality. In this way, our technique allows an operator to recognize the remote place and navigate the robot intuitively by seamlessly using a variety of mixed reality techniques on a spectrum of Milgram's real-virtual continuum.
Kensaku Saitoh, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR3
2005 A Hybrid Image-Based and Model-Based Telepresence System Using Two-Pass Video Projection onto a 3D Scene Model
abstract
A telepresence system is presented that has the advantage of both model-based and image-based approaches, namely, free viewpoint control and real-time color update with live video. A remote place is presented as a virtual environment by using live video projection captured by a head-worn camera onto the static 3D geometry. The observer can then observe the remote place in cooperation with the remote camera man, and give him a set of 3D instructions by a mouse.
Takefumi Ogawa, Kiyoshi Kiyokawa, Haruo Takemura
ISMAR2
2005 A Study of Depth Visualization Techniques for Virtual Annotations in Augmented Reality
abstract
In this paper, we discuss depth visualization techniques for virtual annotations to alleviate the depth ambiguity problem. We begin by describing the depth ambiguity problem with regard to virtual annotations. Then a number of possible solutions are discussed by introducing a metaphor of monocular depth cues, and the effectiveness and characteristics of the three visualization techniques are shown through the preliminary experiments.
Kengo Uratani, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
VR3
2005 A 2D-3D integrated environment for cooperative work
abstract
This paper proposes a novel tabletop display system for natural communication and flexible information sharing. The proposed system is specifically designed for integration of 2D and 3D user interfaces, using a multi-user stereoscopic display, IllusionHole. The proposed system takes awareness into consideration and provides both 2D and 3D information and user interfaces. On the display, a number of standard Windows desktop environments are provided as personal workspaces, as well as a shared workspace with a dedicated graphical user interface. In personal workspaces, users can simultaneously access existing applications and data, and exchange information between personal and shared workspaces. In this way, the proposed system can seamlessly integrate personal, shared, 2D and 3D workspaces with conventional user interfaces and effectively support communication and information sharing. To demonstrate capabilities of the proposed display system, a modeling application has been implemented. A preliminary experiment confirmed the effectiveness of the system.
Kousuke Nakashima, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
VRST3
2004 Unified Gesture-Based Interaction Techniques for Object Manipulation and Navigation in a Large-Scale Virtual Environment
Yusuke Tomozoe, Takashi Machida, Kiyoshi Kiyokawa, Haruo Takemura
VR3
2003 A High Immersive Tele- Directing System Using CyberDome
Tomoaki Adachi, Takefumi Ogawa, Kiyoshi Kiyokawa, Haruo Takemura
INTERACT3
2003 An Occlusion-Capable Optical See-through Head Mount Display for Supporting Co-located Collaboration
abstract
An ideal augmented reality (AR) display for multi-user co-located collaboration should have following three features: 1) any virtual object should be able to be shown at any arbitrary position, e.g. a user can see a virtual object in front of other users' faces. 2) Correct occlusion of virtual and real objects should be supported. 3) The real world should be naturally and clearly visible, which is important for face-to-face conversation. We have been developing an optical see-through display, ELMO (Enhanced see-through display using an LCD panel for Mutual Occlusion), that satisfies these three requirements. While previous prototype systems were not practical due to their size and weight, we have come up with an improved optics design which has reduced size and is lightweight enough to wear. In this paper, the characteristics of typical multi-user three-dimensional displays are summarized and the design details of the latest optics are then described. Finally, a collaborative AR application employing the new display and its user experience are explained.
Kiyoshi Kiyokawa, Mark Billinghurst, Bruce Donald Campbell, Eric Woods
ISMAR1
2003 Communication Behaviors in Colocated Collaborative AR Interfaces
abstract
The authors present an analysis of communication behavior in face-to-face collaboration using a multi-user augmented reality (AR) interface. 2 experiments were conducted. In the 1st experiment, collaboration with AR technology was compared with more traditional unmediated and screen-based collaboration. In the 2nd experiment, the authors compared collaboration with 3 different AR displays. Several measures were used to analyze communication behavior, and the authors found that users exhibited many of the same behaviors in a collaborative AR interface as in face-to-face unmediated collaboration. However, user communication behavior changed with the type of AR display used. The authors describe implications of these results for the design of collaborative AR interfaces and directions for future research.
Mark Billinghurst, Daniel Belcher, Arnab Gupta, Kiyoshi Kiyokawa
Int. J. Hum. Comput. Interact.4
2002 Communication Behaviors of Co-Located Users in Collaborative AR Interfaces
abstract
We conducted two experiments comparing communication behaviors of co-located users in collaborative augmented reality (AR) interfaces. In the first experiment, we compared optical, stereo- and mono-video, and immersive head mounted displays (HMDs) using a target identification task. It was found that differences in the real world visibility severely affect communication behaviors. The optical see-through case produced the best results with the least extra communication needed. Generally, the more difficult it was to use non-verbal communication cues, the more people resorted to speech cues to compensate. In the second experiment, we compared three different combinations of task and communication spaces using a 2D icon design task with optical see-through HMDs. It was found that the spatial relationship between the task and communication spaces also severely affected communication behaviors. Placing the task space between the subjects produced the most active behaviors in terms of initiatory body languages and utterances with least miscommunications.
Kiyoshi Kiyokawa, Mark Billinghurst, Sean Hayes, Anoop Gupta, Yuki Sannohe, Hirokazu Kato 0001
ISMAR1
2001 An optical see-through display for mutual occlusion with a real-time stereovision system
Kiyoshi Kiyokawa, Yoshinori Kurata, Hiroyuki Ohno
Comput. Graph.1
1998 An augmented reality system using a real-time vision based registration
abstract
Describes a prototype of an augmented reality system using vision-based registration. In order to build an augmented reality system with video see-through image composition, camera parameters for generating virtual objects must be obtained at video-rate. The system estimates camera parameters from four known markers in an image sequence captured by a small CCD camera mounted on a HMD (head mounted display). Virtual objects are overlaid upon a captured image sequence in real-time.
Takashi Okuma, Kiyoshi Kiyokawa, Haruo Takemura, Naokazu Yokoya
ICPR2
1996 VLEGO: a simple two-handed modeling environment based on toy blocks
abstract
This paper describes a case study of building a prototype of an immersive three dimensional (3-D) modeler which supports simple two-handed operations. Designing 3-D objects in a virtual environment has a number of advantages for 3-D geometry creation over designing with traditional computer aided design (CAD) tools. In order to enhance the human-computer interaction in a virtual workspace, two-handed spatial input has been incorporated into a few 3-D designing applications. However, existing 3-D designing tools do not utilize two handed interaction for enhancing the interface sufficiently. Our prototype immersive modeler, VLEGO, employs some features of toy blocks to give flexible two-handed interaction for 3-D design. Features of VLEGO can be summarized as follows: Firstly, VLEGO supports various two-handed operations and hence it makes design environment intuitive and efficient. Secondly, possible location and orientation of primitives are discretely limited so that the user can arrange objects accurately with ease. Finally, the system automatically avoids collisions among primitives and adjusts their positions. As a result, precise design of 3-D objects can be achieved easily by using a set of two-handed operations in intuitive way. This paper describes the design and implementation of VLEGO as well as an experiment for examining the effectiveness of two-handed interaction.
Kiyoshi Kiyokawa, Haruo Takemura, Yoshiaki Katayama, Hidehiko Iwasa, Naokazu Yokoya
VRST1