VLDB 2026 Research / reviewers in the wild / expert
Jason Orlosky
dblp:127/6421
· DBLP profile ↗
33ranked-venue papers
9as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 23 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalable Object Detection in Mixed Reality Using Incremental Re-Training and One-Shot 3D AnnotationabstractWhile object detection can be incredibly useful for a variety of augmented and mixed reality applications, achieving a large number of classifiable objects with high accuracy without extremely large deep learning (DL) or object recognition models is still difficult. More importantly, object recognition frameworks are often rigid in that they don't provide a direct means to add new classes to pretrained models in real time. In this paper, we introduce a novel approach that enables ondemand training of new object classes for consistent detection of in-situ objects for virtual labeling and interaction. By leveraging knowledge of the 3D location of an object in the scene taken from a mixed reality (MR) display's environment mesh, we are able to automate the labeling of subsequent 2D images taken from the frontfacing camera, which requires only a single, initial labeling interaction from an end-user. In addition, we have developed a continual learning approach that allows for on-the-fly retraining of the classifier and provides accurate classification quickly enough for the model to be practically usable in MR applications. We validate this approach by measuring the re-training time required for various object configurations, provide a comparison to other classification strategies, and analyze how the addition of object classes affect detection continuity across 3D scenes. We also demonstrate that labeling interactions work for practical applications in AR that are dependent on object detection, such as language learning, procedural instruction, or manufacturing guidance. Alireza Taheritajar, Jeffrey Benson, Anthony Gibson, Brandon Wilburn, Jieqiong Zhao, Jason Orlosky |
ISMAR | 6 |
| 2025 | User-Centric Locomotion Techniques for Virtual Reality Games: A Survey of User Needs and IssuesabstractVirtual reality (VR) video games that are played on a VR headset are becoming increasingly common in households, and though many games require players to navigate vast virtual spaces, most homes cannot provide a large enough physical space to encompass the entire virtual space. Thus, VR video games that require locomotion often provide users with alternative locomotion techniques. While teleportation or steering is typically used as a standard, new techniques can overcome remaining problems, such as motion sickness. However, a holistic perspective of user needs and issues regarding these techniques in practical situations has not been studied on a broad basis. To address this gap in the literature and contribute to future VR video game development and research, we conducted 16 semi-structured interviews and surveyed 88 participants to help explore issues regarding existing locomotion techniques. Our results revealed preferences related to teleportation versus steering and the postures that users adopt while playing VR video games, along with user needs for locomotion techniques in each posture. Daichi Hirobe, Shizuka Shirai, Jason Orlosky, Mehrasa Alizadeh, Masato Kobayashi 0001, Yuuki Uranishi, Photchara Ratsamee, Haruo Takemura |
IEEE Trans. Games | 3 |
| 2024 | ReAR Indicators: Peripheral Cycling Indicators for Rear-Approaching HazardsabstractDuring cycling activities, cyclists often focus on pedestrians, vehicles or road conditions in front of their bicycle. Because of this forward focus, approaching vehicles from behind can easily be missed, which can result in accidents, injury, or death. Although rear information can viewed with handle-mounted mirrors or monitors, looking down can distract the cyclist from other hazards. Guanghan Zhao, Xiaodan Hu, Jason Orlosky, Kiyoshi Kiyokawa |
AVI | 3 |
| 2024 | GlanXR: A Hands-Free Fast Switching System for Virtual ScreensabstractTo date, virtual and augmented reality technologies enable users to view multiple, large virtual screens in their workspaces. However, users must frequently rotate their heads to shift focus among these screens. This paper presents GlanXR, a fast and robust handsfree approach for screen switching in virtual reality. GlanXR incorporates a peripheral interface that remains fixed within the user’s view, in which screens can be dynamically selected based on the user’s eye-head position beyond an adaptive range. Additionally, the user triggers the switch to the screen chosen by making an opposing head rotation in the direction of the eye-head position to minimize false triggers. We conducted an experiment including a fast-switching scenario and a working simulation scenario with 24 participants to assess the effectiveness of GlanXR as compared to a baseline (taskbar), an expansive multi-screen setup, and a gazebased screen selection method. The results indicate that GlanXR facilitates precise screen-switching, minimizes the necessity for head rotation, and allows users to maintain a neutral head position. Guanghan Zhao, Jason Orlosky, Kiyoshi Kiyokawa, Yuuki Uranishi |
ISMAR | 2 |
| 2024 | EyeShadows: Peripheral Virtual Copies for Rapid Gaze Selection and InteractionabstractIn eye-tracked augmented and virtual reality (AR/VR), instantaneous and accurate hands-free selection of virtual elements is still a significant challenge. Though other methods that involve gaze-coupled head movements or hovering can improve selection times in comparison to methods like gaze-dwell, they are either not instantaneous or have difficulty ensuring that the user’s selection is deliberate. In this paper, we present EyeShadows, an eye gaze-based selection system that takes advantage of peripheral copies (shadows) of items that allow for quick selection and manipulation of an object or corresponding menus. This method is compatible with a variety of different selection tasks and controllable items, avoids the Midas touch problem, does not clutter the virtual environment, and is context sensitive. We have implemented and refined this selection tool for VR and AR, including testing with optical and video see-through (OST/VST) displays. Moreover, we demonstrate that this method can be used for a wide range of AR and VR applications, including manipulation of sliders or analog elements. We test its performance in VR against three other selection techniques, including dwell (baseline), an inertial reticle, and head-coupled selection. Results showed that selection with EyeShadows was significantly faster than dwell (baseline), outperforming in the select and search and select tasks by 29.8% and 15.7%, respectively, though error rates varied between tasks. Jason Orlosky, Chang Liu 0081, Kenya Sakamoto, Ludwig Sidenmark, Adam Mansour |
VR | 1 |
| 2024 | HazARdSnap: Gazed-Based Augmentation Delivery for Safe Information Access While CyclingabstractDuring cycling activities, cyclists often monitor a variety of information such as heart rate, distance, and navigation using a bike-mounted phone or cyclocomputer. In many cases, cyclists also ride on sidewalks or paths that contain pedestrians and other obstructions such as potholes, so monitoring information on a bike-mounted interface can slow the cyclist down or cause accidents and injury. In this article, we present HazARdSnap, an augmented reality-based information delivery approach that improves the ease of access to cycling information and at the same time preserves the user's awareness of hazards. To do so, we implemented real-time outdoor hazard detection using a combination of computer vision and motion and position data from a head mounted display (HMD). We then developed an algorithm that snaps information to detected hazards when they are also viewed so that users can simultaneously view both rendered virtual cycling information and the real-world cues such as depth, position, time to hazard, and speed that are needed to assess and avoid hazards. Results from a study with 24 participants that made use of real-world cycling and virtual hazards showed that both HazARdSnap and forward-fixed augmented reality (AR) user interfaces (UIs) can effectively help cyclists access virtual information without having to look down, which resulted in fewer collisions (51% and 43% reduction compared to baseline, respectively) with virtual hazards. Guanghan Zhao, Jason Orlosky, Joseph L. Gabbard, Kiyoshi Kiyokawa |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Mitigation of VR Sickness During Locomotion With a Motion-Based Dynamic Vision ModulatorabstractIn virtual reality, VR sickness resulting from continuous locomotion via controllers or joysticks is still a significant problem. In this article, we present a set of algorithms to mitigate VR sickness that dynamically modulate the user's field of view by modifying the contrast of the periphery based on movement, color, and depth. In contrast with previous work, this vision modulator is a shader that is triggered by specific motions known to cause VR sickness, such as acceleration, strafing, and linear velocity. Moreover, the algorithm is governed by delta velocity, delta angle, and average color of the view. We ran two experiments with different washout periods to investigate the effectiveness of dynamic modulation on the symptoms of VR sickness, in which we compared this approach against a baseline and pitch-black field-of-view restrictors. Our first experiment made use of a just-noticeable-sickness design, which can be useful for building experiments with a short washout period. Guanghan Zhao, Jason Orlosky, Steven K. Feiner, Photchara Ratsamee, Yuuki Uranishi |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Layerable Apps: Comparing Concurrent and Exclusive Display of Augmented Reality ApplicationsabstractCurrent augmented reality (AR) interfaces are often designed for interacting with one application at a time, significantly limiting a user’s ability to concurrently interact with and switch between multiple applications or modalities that could run in parallel. In this work, we introduce an application model called Layerable Apps, which supports a variety of AR application types while enabling multitasking through concurrent execution, fast application switching, and the ability to layer application views to adjust the degree of augmentation to the user’s preference. We evaluated Layerable Apps through a within-subjects user study (n=44), compared against a traditional single-focus application model on a split-information task involving the simultaneous use of multiple applications. We report the results of our study, where we found differences in quantitative task performance, favoring Layerable mode. We also analyzed app usage patterns, spatial awareness, and overall preferences between both modes as well as between experienced and novice AR users. Brandon Huynh, Abby Wysopal, Vivian Ross, Jason Orlosky, Tobias Höllerer |
ISMAR | 4 |
| 2021 | Genetic crossover in the evolution of time-dependent neural networks
Jason Orlosky, Tim Grabowski |
GECCO | 1 |
| 2021 | UAV Target-Selection: 3D Pointing Interface System for Large-Scale EnvironmentabstractThis paper presents a 3D pointing interface application to signal a UAV’s target in a large-scale environment. This system enables UAVs equipped with a monocular camera to determine which window of a building is selected by a human user in large-scale indoor or outdoor environments. The 3D pointing interface consists of three parts: YOLO, Open- Pose, and ORB-SLAM. YOLO detects the target objects, e.g., windows, OpenPose extracts the user pose, and ORB-SLAM builds a scale-dependent 3D map, a set of 3D sparse feature points. To obtain the visual scale, it performs a calibration step with the user standing in front of the UAV at a certain distance. We detail how we chose the gesture, localize and detect objects, and transform between coordinate systems. The real- world experiment results showed that the 3D pointing interface obtained a 0.73 F1-score average and a 0.58 F1-Score at the maximum distance of 25 meters between UAV and building. Anna Medeiros, Photchara Ratsamee, Jason Orlosky, Yuuki Uranishi, Manabu Higashida, Haruo Takemura |
ICRA | 3 |
| 2020 | OrthoGaze: Gaze-based three-dimensional object manipulation using orthogonal planes
Chang Liu 0081, Alexander Plopski, Jason Orlosky |
Comput. Graph. | 3 |
| 2019 | Revisiting Virtual Reality for Practical Use in Therapy: Patient Satisfaction in Outpatient RehabilitationabstractThough Virtual Reality (VR) has made its way into many commercial applications, it has only begun to gain adoption in applications for rehabilitation and therapy. Though some research exists in this area, we still need to better understand how interactive VR can be integrated into a fast-paced clinical setting. To address this challenge, we present the results of a pilot study with 25 participants, who used an interactive application for conducting upper-limb rehabilitation. Our interface consists of a VR display with two controllers that are used to clean and clear segments of a target virtual window. From observations of participant challenges with the interface, feedback from the occupational therapy staff, and a subjective questionnaire using Likert scales, we examine ways to better integrate VR into a patient's scheduled rehabilitation. Josephine H. Hartney, Sydney N. Rosenthal, Aaron M. Kirkpatrick, J. Maci Skinner, Jason Hughes, Jason Orlosky |
VR | 6 |
| 2019 | Semantic Labeling and Object Registration for Augmented Reality Language LearningabstractWe propose an Augmented Reality vocabulary learning interface in which objects in a user's environment are automatically recognized and labeled in a foreign language. Using AR for language learning in this manner is still impractical for a number of reasons. Scalable object recognition and consistent labeling of objects is still a significant challenge, and interaction with arbitrary physical objects in AR scenes has consequently not been well explored. To help address these challenges, we present a system that utilizes real-time object recognition to perform semantic labeling and object registration in Augmented Reality. We discuss its implementation, our motivations in designing it, and how it can be applied to AR language learning applications. Brandon Huynh, Jason Orlosky, Tobias Höllerer |
VR | 2 |
| 2019 | In-Situ Labeling for Augmented Reality Language LearningabstractAugmented Reality is a promising interaction paradigm for learning applications. It has the potential to improve learning outcomes by merging educational content with spatial cues and semantically relevant objects within a learner's everyday environment. The impact of such an interface could be comparable to the method of loci, a well known memory enhancement technique used by memory champions and polyglots. However, using Augmented Reality in this manner is still impractical for a number of reasons. Scalable object recognition and consistent labeling of objects is a significant challenge, and interaction with arbitrary (unmodeled) physical objects in AR scenes has consequently not been well explored. To help address these challenges, we present a framework for in-situ object labeling and selection in Augmented Reality, with a particular focus on language learning applications. Our framework uses a generalized object recognition model to identify objects in the world in real time, integrates eye tracking to facilitate selection and interaction within the interface, and incorporates a personalized learning model that dynamically adapts to student's growth. We show our current progress in the development of this system, including preliminary tests and benchmarks. We explore challenges with using such a system in practice, and discuss our vision for the future of AR language learning applications. Brandon Huynh, Jason Orlosky, Tobias Höllerer |
VR | 2 |
| 2019 | Evaluation of Pointing Interfaces with an AR Agent for Multi-section Information GuidanceabstractIn educational settings such as art galleries or museums, Augmented Reality (AR) has the potential to provide detailed information about exhibits. However, dealing with items that contain information in multiple sections or areas is still a significant challenge. For example, a large painting may contain many minute details, which requires a system that can explain its broader features rather than just a generic description. To address this challenge, we introduce an AR guidance system that uses an embodied agent to point out items and explain each piece and part of exhibit items in detail. We also designed and tested 3 different pointing interfaces for the embodied agent: gesture only, gesture with a dot laser, and gesture with line laser. To evaluate this interface, we conducted a user experiment simulating painting guidance to test interest and exhibit memory. During the experiment, the agent pointed to various areas of interest in the painting and provided a detailed description to participants. The result shows that the search times for target positions were the fastest with the line laser. However, no particular interface outperformed others in memory recall of exhibit content. Nattaon Techasarntikul, Tomohiro Mashita, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura, Jason Orlosky, Kiyoshi Kiyokawa |
VR | 6 |
| 2019 | Thermal HDR: Applying High Dynamic Range Rendering for Fusion of Thermal Augmentations with Visible LightabstractIn safety applications for fields such as navigation or industrial manufacturing, thermal augmentations have often been used to help users safely navigate low-light environments and detect obstacles. However, one problem with thermal imaging is that it often occludes the environment and coloration that would otherwise be useful, for example traffic sign text or warning labels on industrial equipment. In this paper, we explore a new algorithm that takes advantage of high dynamic range (HDR) rendering techniques in order to more effectively present thermal information. Unlike previous fusion algorithms or conventional blending techniques, our setup makes use of both a thermal camera and HDR frames to synthesize a final overlay. Moreover, we set up a series of two experiments with a simulated heads up display (HUD) to 1) measure reaction times to the sudden appearance of pedestrians, and 2) conduct circuit repair. Results showed that the HDR algorithm was subjectively preferred to other approaches, and that performance on average could match other conventional algorithms on most occasions. Jason Orlosky |
VR | 2 |
| 2019 | A Comparison of Adaptive View Techniques for Exploratory 3D Drone TeleoperationabstractDrone navigation in complex environments poses many problems to teleoperators. Especially in three dimensional (3D) structures such as buildings or tunnels, viewpoints are often limited to the drone’s current camera view, nearby objects can be collision hazards, and frequent occlusion can hinder accurate manipulation. To address these issues, we have developed a novel interface for teleoperation that provides a user with environment-adaptive viewpoints that are automatically configured to improve safety and provide smooth operation. This real-time adaptive viewpoint system takes robot position, orientation, and 3D point-cloud information into account to modify the user’s viewpoint to maximize visibility. Our prototype uses simultaneous localization and mapping (SLAM) based reconstruction with an omnidirectional camera, and we use the resulting models as well as simulations in a series of preliminary experiments testing navigation of various structures. Results suggest that automatic viewpoint generation can outperform first- and third-person view interfaces for virtual teleoperators in terms of ease of control and accuracy of robot operation. John Thomason, Photchara Ratsamee, Jason Orlosky, Kiyoshi Kiyokawa, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2019 | The Influence of Label Design on Search Performance and Noticeability in Wide Field of View Augmented Reality DisplaysabstractIn Augmented Reality (AR), search performance for outdoor tasks is an important metric for evaluating the success of a large number of AR applications. Users must be able to find content quickly, labels and indicators must not be invasive but still clearly noticeable, and the user interface should maximize search performance in a variety of conditions. To address these issues, we have set up a series of experiments to test the influence of virtual characteristics such as color, size, and leader lines on the performance of search tasks and noticeability in both real and simulated environments. We evaluate two primary areas, including 1) the effects of peripheral field of view (FOV) limitations and labeling techniques on target acquisition during outdoor mobile search, and 2) the influence of local characteristics such as color, size, and motion on text labels over dynamic backgrounds. The first experiment showed that limited FOV will severely limit search performance, but that appropriate placement of labels and leaders within the periphery can alleviate this problem without interfering with walking or decreasing user comfort. In the second experiment, we found that different types of motion are more noticeable in optical versus video see-through displays, but that blue coloration is most noticeable in both. Results can aid in designing more effective view management techniques, especially for wider field of view displays. Ernst Kruijff, Jason Orlosky, Naohiro Kishishita, Christina Trepkowski, Kiyoshi Kiyokawa |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | IntelliPupil: Pupillometric Light Modulation for Optical See-Through Head-Mounted DisplaysabstractIn practical use of optical see-through head-mounted displays, users often have to adjust the brightness of virtual content to ensure that it is at the optimal level. Automatic adjustment is still a challenging problem, largely due to the bidirectional nature of the structure of the human eye, complexity of real world lighting, and user perception. Allowing the right amount of light to pass through to the retina requires a constant balance of incoming light from the real world, additional light from the virtual image, pupil contraction, and feedback from the user. While some automatic light adjustment methods exist, none have completely tackled this complex input-output system. As a step towards overcoming this issue, we introduce IntelliPupil, an approach that uses eye tracking to properly modulate augmentation lighting for a variety of lighting conditions and real scenes. We first take the data from a small form factor light sensor and changes in pupil diameter from an eye tracking camera as passive inputs. This data is coupled with user-controlled brightness selections, allowing us to fit a brightness model to user preference using a feed-forward neural network. Using a small amount of training data, both scene luminance and pupil size are used as inputs into the neural network, which can then automatically adjust to a user's personal brightness preferences in real time. Experiments in a high dynamic range AR scenario with varied lighting show that pupil size is just as important as environment light for optimizing brightness and that our system outperforms linear models. Chang Liu 0081, Alexander Plopski, Kiyoshi Kiyokawa, Photchara Ratsamee, Jason Orlosky |
ISMAR | 5 |
| 2017 | VisMerge: Light Adaptive Vision Augmentation via Spectral and Temporal Fusion of Non-visible LightabstractLow light situations pose a significant challenge to individuals working in a variety of different fields such as firefighting, rescue, maintenance and medicine. Tools like flashlights and infrared (IR) cameras have been used to augment light in the past, but they must often be operated manually, provide a field of view that is decoupled from the operator's own view, and utilize color schemes that can occlude content from the original scene. To help address these issues, we present VisMerge, a framework that combines a thermal imaging head mounted display (HMD) and algorithms that temporally and spectrally merge video streams of different light bands into the same field of view. For temporal synchronization, we first develop a variant of the time warping algorithm used in virtual reality (VR), but redesign it to merge video see-through (VST) cameras with different latencies. Next, using computer vision and image compositing we develop five new algorithms designed to merge non-uniform video streams from a standard RGB camera and small form-factor infrared (IR) camera. We then implement six other existing fusion methods, and conduct a series of comparative experiments, including a system level analysis of the augmented reality (AR) time warping algorithm, a pilot experiment to test perceptual consistency across all eleven merging algorithms, and an in-depth experiment on performance testing the top algorithms in a VR (simulated AR) search task. Results showed that we can reduce temporal registration error due to inter-camera latency by an average of 87.04%, that the wavelet and inverse stipple algorithms were perceptually rated the highest, that noise modulation performed best, and that freedom of user movement is significantly increased with visualizations engaged. Jason Orlosky, Peter Kim, Kiyoshi Kiyokawa, Tomohiro Mashita, Photchara Ratsamee, Yuuki Uranishi, Haruo Takemura |
ISMAR | 1 |
| 2017 | Adaptive View Management for Drone Teleoperation in Complex 3D StructuresabstractDrone navigation in complex environments poses many problems to teleoperators. Especially in 3D structures like buildings or tunnels, viewpoints are often limited to the drone's current camera view, nearby objects can be collision hazards, and frequent occlusion can hinder accurate manipulation. To address these issues, we have developed a novel interface for teleoperation that provides a user with environment-adaptive viewpoints that are automatically configured to improve safety and smooth user operation. This real-time adaptive viewpoint system takes robot position, orientation, and 3D pointcloud information into account to modify user-viewpoint to maximize visibility. Our prototype uses simultaneous localization and mapping (SLAM) based reconstruction with an omnidirectional camera and we use resulting models as well as simulations in a series of preliminary experiments testing navigation of various structures. Results suggest that automatic viewpoint generation can outperform first and third-person view interfaces for virtual teleoperators in terms of ease of control and accuracy of robot operation. John Thomason, Photchara Ratsamee, Kiyoshi Kiyokawa, Pakpoom Kriengkomol, Jason Orlosky, Tomohiro Mashita, Yuuki Uranishi, Haruo Takemura |
IUI | 5 |
| 2017 | Monocular focus estimation method for a freely-orienting eye using Purkinje-Sanson imagesabstractWe present a method for focal distance estimation of a freely-orienting eye using Purkinje-Sanson (PS) images, which are reflections of light on the inner structures of the eye. Using an infrared camera with a rigidly-fixed LED, our method creates an estimation model based on 3D gaze and the distance between reflections in the PS images that occur on the corneal surface and anterior surface of the eye lens. The distance between these two reflections changes with focus, so we associate that information to the focal distance on a user. Unlike conventional methods that mainly relies on 2D pupil size which is sensitive to scene lighting and the fourth PS image, our method detects the third PS image which is more representative of accommodation. Our feasibility study on a single user with a focal range from 15–45 cm shows that our method achieves mean and median absolute errors of 3.15 and 1.93 cm for a 10-degree viewing angle. The study shows that our method is also tolerant against environment lighting changes. Yuta Itoh 0001, Jason Orlosky, Kiyoshi Kiyokawa, Toshiyuki Amano, Maki Sugimoto |
VR | 2 |
| 2017 | Emulation of Physician Tasks in Eye-Tracked Virtual Reality for Remote Diagnosis of Neurodegenerative DiseaseabstractFor neurodegenerative conditions like Parkinson's disease, early and accurate diagnosis is still a difficult task. Evaluations can be time consuming, patients must often travel to metropolitan areas or different cities to see experts, and misdiagnosis can result in improper treatment. To date, only a handful of assistive or remote methods exist to help physicians evaluate patients with suspected neurological disease in a convenient and consistent way. In this paper, we present a low-cost VR interface designed to support evaluation and diagnosis of neurodegenerative disease and test its use in a clinical setting. Using a commercially available VR display with an infrared camera integrated into the lens, we have constructed a 3D virtual environment designed to emulate common tasks used to evaluate patients, such as fixating on a point, conducting smooth pursuit of an object, or executing saccades. These virtual tasks are designed to elicit eye movements commonly associated with neurodegenerative disease, such as abnormal saccades, square wave jerks, and ocular tremor. Next, we conducted experiments with 9 patients with a diagnosis of Parkinson's disease and 7 healthy controls to test the system's potential to emulate tasks for clinical diagnosis. We then applied eye tracking algorithms and image enhancement to the eye recordings taken during the experiment and conducted a short follow-up study with two physicians for evaluation. Results showed that our VR interface was able to elicit five common types of movements usable for evaluation, physicians were able to confirm three out of four abnormalities, and visualizations were rated as potentially useful for diagnosis. Jason Orlosky, Yuta Itoh 0001, Maud Ranchet, Kiyoshi Kiyokawa, John Morgan, Hannes Devos |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Automated Spatial Calibration of HMD Systems with Unconstrained Eye-camerasabstractProperly calibrating an optical see-through head-mounted display (OST-HMD) and maintaining a consistent calibration over time can be a very challenging task. Automated methods need an accurate model of both the OST-HMD screen and the user's constantly changing eye-position to correctly project virtual information. While some automated methods exist, they often have restrictions, including fixed eye-cameras that cannot be adjusted for different users.To address this problem, we have developed a method that automatically determines the position of an adjustable eye-tracking camera and its unconstrained position relative to the display. Unlike methods that require a fixed pose between the HMD and eye camera, our framework allows for automatic calibration even after adjustments of the camera to a particular individual's eye and even after the HMD moves on the user's face. Using two sets of IR-LEDs rigidly attached to the camera and OST-HMD frame, we can calculate the correct projection for different eye positions in real time and changes in HMD position within several frames. To verify the accuracy of our method, we conducted two experiments with a commercial HMD by calibrating a number of different eye and camera positions. Ground truth was measured through markers on both the camera and HMD screens, and we achieve a viewing accuracy of 1.66 degrees for the eyes of 5 different experiment participants. Alexander Plopski, Jason Orlosky, Yuta Itoh 0001, Christian Nitschke, Kiyoshi Kiyokawa, Gudrun Klinker |
ISMAR | 2 |
| 2016 | OST Rift: Temporally consistent augmented reality with a consumer optical see-through head-mounted displayabstractWe present an off-the-shelf, low-latency Optical See-through Head-Mounted Displays (OST-HMD) for Augmented Reality (AR). Temporally consistent visualization is crucial for realizing immersive AR experiences. This is challenging since it requires both accurate head-tracking and low-latency rendering of AR content. Building a system which meets both constraints usually requires experts on computer vision/graphics and expensive display hardware. This work demonstrates that such high spatio-temporal fidelity is achievable with commodity hardware available today. We build a custom OST-HMD system that consists of a virtual reality HMD, i.e., the Oculus Rift DK2, and half-mirror optics, and adapt the rendering pipeline in order to integrate the OST-HMD calibration framework. An evaluation with a user-perspective camera shows that the system achieves mean temporal error of <;1 ms (95% reduction of the latency from naive, no-predictive rendering), and median spatial error <;0.3° in the viewing angle with maximum error at most 1.0°. Yuta Itoh 0001, Jason Orlosky, Manuel J. Huber, Kiyoshi Kiyokawa, Gudrun Klinker |
VR | 2 |
| 2015 | Halo Content: Context-aware Viewspace Management for Non-invasive Augmented RealityabstractIn mobile augmented reality, text and content placed in a user's immediate field of view through a head worn display can interfere with day to day activities. In particular, messages, notifications, or navigation instructions overlaid in the central field of view can become a barrier to effective face-to-face meetings and everyday conversation. Many text and view management methods attempt to improve text viewability, but fail to provide a non-invasive personal experience for the user. Jason Orlosky, Kiyoshi Kiyokawa, Takumi Toyama, Daniel Sonntag |
IUI | 1 |
| 2015 | Attention Engagement and Cognitive State Analysis for Augmented Reality Text Display FunctionsabstractHuman eye gaze has recently been used as an effective input interface for wearable displays. In this paper, we propose a gaze-based interaction framework for optical see-through displays. The proposed system can automatically judge whether a user is engaged with virtual content in the display or focused on the real environment and can determine his or her cognitive state. With these analytic capacities, we implement several proactive system functions including adaptive brightness, scrolling, messaging, notification, and highlighting, which would otherwise require manual interaction. The goal is to manage the relationship between virtual and real, creating a more cohesive and seamless experience for the user. We conduct user experiments including attention engagement and cognitive state analysis, such as reading detection and gaze position estimation in a wearable display towards the design of augmented reality text display applications. The results from the experiments show robustness of the attention engagement and cognitive state analysis methods. A majority of the experiment participants (8/12) stated the proactive system functions are beneficial. Takumi Toyama, Daniel Sonntag, Jason Orlosky, Kiyoshi Kiyokawa |
IUI | 3 |
| 2015 | ModulAR: Eye-Controlled Vision Augmentations for Head Mounted DisplaysabstractIn the last few years, the advancement of head mounted display technology and optics has opened up many new possibilities for the field of Augmented Reality. However, many commercial and prototype systems often have a single display modality, fixed field of view, or inflexible form factor. In this paper, we introduce Modular Augmented Reality (ModulAR), a hardware and software framework designed to improve flexibility and hands-free control of video see-through augmented reality displays and augmentative functionality. To accomplish this goal, we introduce the use of integrated eye tracking for on-demand control of vision augmentations such as optical zoom or field of view expansion. Physical modification of the device's configuration can be accomplished on the fly using interchangeable camera-lens modules that provide different types of vision enhancements. We implement and test functionality for several primary configurations using telescopic and fisheye camera-lens systems, though many other customizations are possible. We also implement a number of eye-based interactions in order to engage and control the vision augmentations in real time, and explore different methods for merging streams of augmented vision into the user's normal field of view. In a series of experiments, we conduct an in depth analysis of visual acuity and head and eye movement during search and recognition tasks. Results show that methods with larger field of view that utilize binary on/off and gradual zoom mechanisms outperform snapshot and sub-windowed methods and that type of eye engagement has little effect on performance. Jason Orlosky, Takumi Toyama, Kiyoshi Kiyokawa, Daniel Sonntag |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | A natural interface for multi-focal plane head mounted displays using 3D gazeabstractIn mobile augmented reality (AR), it is important to develop interfaces for wearable displays that not only reduce distraction, but that can be used quickly and in a natural manner. In this paper, we propose a focal-plane based interaction approach with several advantages over traditional methods designed for head mounted displays (HMDs) with only one focal plane. Using a novel prototype that combines a monoscopic multi-focal plane HMD and eye tracker, we facilitate interaction with virtual elements such as text or buttons by measuring eye convergence on objects at different depths. This can prevent virtual information from being unnecessarily overlaid onto real world objects that are at a different range, but in the same line of sight. We then use our prototype in a series of experiments testing the feasibility of interaction. Despite only being presented with monocular depth cues, users have the ability to correctly select virtual icons in near, mid, and far planes in 98.6% of cases. Takumi Toyama, Daniel Sonntag, Jason Orlosky, Kiyoshi Kiyokawa |
AVI | 3 |
| 2014 | Analysing the effects of a wide field of view augmented reality display on search performance in divided attention tasksabstractA wide field of view augmented reality display is a special type of head-worn device that enables users to view augmentations in the peripheral visual field. However, the actual effects of a wide field of view display on the perception of augmentations have not been widely studied. To improve our understanding of this type of display when conducting divided attention search tasks, we conducted an in depth experiment testing two view management methods, in-view and in-situ labelling. With in-view labelling, search target annotations appear on the display border with a corresponding leader line, whereas in-situ annotations appear without a leader line, as if they are affixed to the referenced objects in the environment. Results show that target discovery rates consistently drop with in-view labelling and increase with in-situ labelling as display angle approaches 100 degrees of field of view. Past this point, the performances of the two view management methods begin to converge, suggesting equivalent discovery rates at approximately 130 degrees of field of view. Results also indicate that users exhibited lower discovery rates for targets appearing in peripheral vision, and that there is little impact of field of view on response time and mental workload. Naohiro Kishishita, Kiyoshi Kiyokawa, Jason Orlosky, Tomohiro Mashita, Haruo Takemura, Ernst Kruijff |
ISMAR | 3 |
| 2013 | Towards intelligent view management: A study of manual text placement tendencies in mobile environments using video see-through displaysabstractWhen viewing content in a see-through head mounted display (HMD), displaying readable information is still difficult when text is overlayed onto a changing background or lighted surface. Moving text or content to a more appropriate place on the screen through automation or intelligent algorithms is one viable solution to this kind of issue. However, many of these algorithms fail to act as a human would when placing text in a more appropriate location in real time. In order to improve these text and view management algorithms, we report the results and analysis of an experiment designed to evaluate user tendencies when placing virtual text in the real world through an HMD. In the conducted experiment, 20 users manually overlayed text in real time onto 4 different videos taken from the first-person perspective of a pedestrian. We find that users have a tendency to place overlayed text in locations near the center of the viewing field, gravitating towards a point just below the horizon. Common locations for text overlay such as walls, shaded areas, and pavement are classified and discussed. Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura |
ISMAR | 1 |
| 2013 | Management and manipulation of text in dynamic mixed reality workspacesabstractViewing and interacting with text based content safely and easily while mobile has been an issue with see-through displays for many years. For example, in order to effectively use optical see through Head Mounted Displays (HMDs) in constantly changing dynamic environments, variables like lighting conditions, human or vehicular obstructions in a user's path, and scene variation must be dealt with effectively. My PhD research focuses on answering the following questions: 1) What are appropriate methods to intelligently move digital content such as e-mail, SMS messeges, and news articles, throughout the real world? 2) Once a user stops moving, in what way should dynamics of the current workspace change when migrated to a new static environment? 3) Lastly, how can users manipulate mobile content using the fewest number of interactions possible? My strategy for developing solutions to these problems primarily involves automatic or semi-automatic movement of digital content throughout the real world using camera tracking. I have already developed an intelligent text management system that actively manages movement of text in a user's field of view while mobile [11]. I am optimizing and expanding on this type of management system, developing appropriate interaction methodology, and conducting experiments to verify effectiveness, usability, and safety when used with an HMD in various environments. Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura |
ISMAR | 1 |
| 2013 | Dynamic text management for see-through wearable and heads-up display systemsabstractReading text safely and easily while mobile has been an issue with see-through displays for many years. For example, in order to effectively use optical see through Head Mounted Displays (HMDs) or Heads Up Display (HUD) systems in constantly changing dynamic environments, variables like lighting conditions, human or vehicular obstructions in a user's path, and scene variation must be dealt with effectively. Jason Orlosky, Kiyoshi Kiyokawa, Haruo Takemura |
IUI | 1 |