VLDB 2026 Research / reviewers in the wild / expert
Shohei Mori
dblp:131/4275
· DBLP profile ↗
33ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0003-0540-7312ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 7 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IEEE VR 2026 Message from the Program Chairs and Guest Editors
Lonni Besançon, Bobby Bodenheimer, Daisuke Iwai, Shohei Mori, Tabitha C. Peck, Richard Skarbez |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | Tag-Along Virtual Windows Increase Perceived Resistance and Task Load in Augmented RealityabstractAugmented Reality (AR) can enhance accessibility by anchoring virtual windows to the user's body. Among common approaches, head-following windows help maintain floating virtual windows within the user's field of view. Previous studies have actively explored this new design space to improve user experience and efficiency. In contrast, this study focuses on the perceived resistance of head-following windows in AR, despite their lack of physical mass. We conducted a within-subject experiment with 24 participants, manipulating Follow-Up Delay, Window Size, and UI Type. We measured subjective resistance ratings, NASA-TLX (Raw TLX Scores), and the gaze-head angular offset. The results showed that both a certain level of Follow-Up Delay and the Tag-Along elicited significantly stronger perceived resistance as well as task load. Although Window Size alone did not show a significant effect on resistance ratings, we observed an interaction between the size and UI Type. These findings extend existing pseudo-haptics research by revealing the previously unexplored domain of resistance in head-based interactions with head-following virtual windows. We further provide design implications for head-following windows in AR. Motoki Kagami, Yuta Kataoka, Yutaro Hirao, Monica Perusquía-Hernández, Satoshi Hashiguchi, Hideaki Uchiyama, Kiyoshi Kiyokawa, Shohei Mori |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2026 | Semantic Scene Graphs for Creating a Localization-Ready Internet of ThingsabstractControlling devices connected to the Internet of Things often requires juggling multiple smartphone apps or physical remote controls, creating a fragmented user experience. Augmented Reality (AR) can afford superior control by automatically presenting virtual user interfaces that are spatially aligned with networked devices. However, before such user interfaces can be delivered, physical devices must be localized in the environment. This paper introduces LORIOT (LOcalization-Ready Internet Of Things), a novel end-to-end system that uses a semantic scene graph and a large language model to map the identities of the networked devices to physical objects, given a pre-filtered set of IoT-device candidate nodes. A declarative UI specification enables automatic generation of device control panels for AR and non-AR clients. We evaluate the mapping component on a controlled synthetic-room benchmark of 100 randomly generated rooms. Using device network metadata alone, we achieve a baseline macro-averaged F1 score of 0.80 for digital $\rightarrow$→ physical associations. When device metadata is enriched with physical attributes (mounting location, materials, color, and size), performance improves to 0.88. Moreover, we evaluate the benefit of spatially registered AR control in a within-subject user study ($N{=}20$N=20), comparing in-situ AR panels against conventional non-AR control with smartphone apps or physical remote controls. AR yields significantly faster task completion, lower mental demand, and higher usability. Jan Kolberg, Michael Pabst, Verena Biener, Shohei Mori, Dieter Schmalstieg |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | User-in-the-Loop View Sampling with Error Peaking VisualizationabstractAugmented reality (AR) provides ways to visualize missing view samples for novel view synthesis. Existing approaches present 3D annotations for new view samples and task users with taking images by aligning the AR display. This data collection task is known to be mentally demanding and limits capture areas to pre-defined small areas due to the ideal but restrictive underlying sampling theory. To free users from 3D annotations and limited scene exploration, we propose using locally reconstructed light fields and visualizing errors to be removed by inserting new views. Our results show that the error-peaking visualization is less invasive, reduces disappointment in final results, and is satisfactory with fewer view samples in our mobile view synthesis system. We also show that our approach can contribute to recent radiance field reconstruction for larger scenes, such as 3D Gaussian splatting. Ayaka Yasunaga, Hideo Saito 0001, Shohei Mori |
ICIP | 3 |
| 2025 | IntelliCap: Intelligent Guidance for Consistent View SamplingabstractNovel view synthesis from images, for example, with 3D Gaussian splatting, has made great progress. Rendering fidelity and speed are now ready even for demanding virtual reality applications. However, the problem of assisting humans in collecting the input images for these rendering algorithms has received much less attention. High-quality view synthesis requires uniform and dense view sampling. Unfortunately, these requirements are not easily addressed by human camera operators, who are in a hurry, impatient, or lack understanding of the scene structure and the photographic process. Existing approaches to guide humans during image acquisition concentrate on single objects or neglect view-dependent material characteristics. We propose a novel situated visualization technique for scanning at multiple scales. During the scanning of a scene, our method identifies important objects that need extended image coverage to properly represent view-dependent appearance. To this end, we leverage semantic segmentation and category identification, ranked by a vision-language model. Spherical proxies are generated around highly ranked objects to guide the user during scanning. Our results show superior performance in real scenes compared to conventional view sampling strategies. Ayaka Yasunaga, Hideo Saito 0001, Dieter Schmalstieg, Shohei Mori |
ISMAR | 4 |
| 2025 | Occlusion-Free 4D Gaussians for Open Surgery Videos Using Multi-camera Shadowless Lamps
Yuna Kato, Shohei Mori, Hideo Saito 0001, Yoshifumi Takatsume, Hiroki Kajita, Mariko Isogawa |
MICCAI (10) | 2 |
| 2025 | NeuralPVS: Learned Estimation of Potentially Visible SetsabstractReal-time visibility determination in expansive or dynamically changing environments has long posed a significant challenge in computer graphics. Existing techniques are computationally expensive and often applied as a precomputation step on a static scene. We present NeuralPVS, the first deep-learning approach for visibility computation that efficiently determines from-region visibility in a large scene, running at approximately 100 Hz processing with less than \(1\%\) missing geometry. This approach is possible by using a neural network operating on a froxelized representation of the scene. The network’s performance is achieved by combining sparse convolution with a 3D volume-preserving interleaving for data compression. Moreover, we introduce a novel repulsive visibility loss that can effectively guide the network to converge to the correct data distribution. This loss provides enhanced robustness and generalization to unseen scenes. Our results demonstrate that NeuralPVS outperforms existing visibility methods in terms of both accuracy and efficiency. Thomas Köhler 0006, Jun Lin Qiu, Shohei Mori, Markus Steinberger, Dieter Schmalstieg |
SIGGRAPH Asia | 4 |
| 2025 | Not All WIP Are Perceived Equally: Different Speed Expectations in Seated Walk-in-Place LocomotionabstractGesture-based locomotion enhances immersion in virtual reality (VR), with seated motion being crucial for accessibility and prolonged use. However, existing techniques often apply uniform gesture-to-walking speed mappings, ignoring the fact that different gestures involve varying levels of physical effort and subjective impressions. This mismatch can degrade the user experience. This study investigates how three seated gestures with different physical loads—Tap-in-Place (TIP), Swing-in-Place (SIP), and Grip-in-Place (GIP)—influence users’ expected walking speed. While the evaluations revealed unique experiential trade-offs for each gesture, our primary finding is a consistent perceptual pattern in the expectation of walking speed: Users expected to walk fastest with SIP, followed by GIP, then TIP (SIP > GIP > TIP). These results demonstrate that a one-size-fits-all approach is insufficient and provide empirical recommendations for designing more intuitive seated VR locomotion systems that align walking speed with user perception. Yusuke Kitaura, Keigo Hattori, Fumihiko Nakamura, Yuta Kataoka, Fumihisa Shibata, Asako Kimura, Shohei Mori |
VRST | 7 |
| 2025 | Dense Depth from Event Focal StackabstractWe propose a method for dense depth estimation from an event stream generated when sweeping the focal plane of the driving lens attached to an event camera. In this method, a depth map is inferred from an “event focal stack” composed of the event stream using a convolutional neural network trained with synthesized event focal stacks. The synthesized event stream is created from a focal stack generated by Blender for any arbitrary 3D scene. This allows for training on scenes with diverse structures. Additionally, we explored methods to eliminate the domain gap between real event streams and synthetic event streams. Our method demonstrates superior performance over a depth-from-defocus method in the image domain on synthetic and real datasets. Kenta Horikawa, Mariko Isogawa, Hideo Saito 0001, Shohei Mori |
WACV | 4 |
| 2025 | Radiance Fields in XR: A Survey on How Radiance Fields are Envisioned and Addressed for XR ResearchabstractThe development of radiance fields (RF), such as 3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF), has revolutionized interactive photorealistic view synthesis and presents enormous opportunities for XR research and applications. However, despite the exponential growth of RF research, RF-related contributions to the XR community remain sparse. To better understand this research gap, we performed a systematic survey of current RF literature to analyze (i) how RF is envisioned for XR applications, (ii) how they have already been implemented, and (iii) the remaining research gaps. We collected 365 RF contributions related to XR from computer vision, computer graphics, robotics, multimedia, human-computer interaction, and XR communities, seeking to answer the above research questions. Among the 365 papers, we performed an analysis of 66 papers that already addressed a detailed aspect of RF research for XR. With this survey, we extended and positioned XR-specific RF research topics in the broader RF research field and provide a helpful resource for the XR community to navigate within the rapid development of RF research. Ke Li 0025, Mana Masuda, Susanne Schmidt 0001, Shohei Mori |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Perceived Weight of Mediated Reality SticksabstractMediated reality, where augmented reality (AR) and diminished reality (DR) meet, enables visual modifications to real-world objects. A physical object with a mediated reality visual change retains its original physical properties. However, it is perceived differently from the original when interacted with. We present such a mediated reality object, a stick with different lengths or a stick with a missing portion in the middle, to investigate how users perceive its weight and center of gravity. We conducted two user studies ($N=10$N=10), each of which consisted of two substudies. We found that the length of mediated reality sticks influences the perceived weight. A longer stick is perceived as lighter, and vice versa. The stick with a missing portion tends to be recognized as one continuous stick. Thus, its weight and center of gravity (COG) remain the same. We formulated the relationship between inertia based on the reported COG and perceived weight in the context of dynamic touch. Satoshi Hashiguchi, Yuta Kataoka, Asako Kimura, Shohei Mori |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | DeepDR: Deep Structure-Aware RGB-D Inpainting for Diminished RealityabstractDiminished reality (DR) refers to the removal of real objects from the environment by virtually replacing them with their background. Modern DR frameworks use inpainting to hallucinate unobserved regions. While recent deep learning-based inpainting is promising, the DR use case is complicated by the need to generate coherent structure and 3D geometry (i.e., depth), in particular for advanced applications, such as 3D scene editing. In this paper, we propose Deep DR, a first RGB-D inpainting framework fulfilling all requirements of DR: Plausible image and geometry inpainting with coherent structure, running at real-time frame rates, with minimal temporal artifacts. Our structure-aware generative network allows us to explicitly condition color and depth outputs on the scene semantics, overcoming the difficulty of reconstructing sharp and consistent boundaries in regions with complex backgrounds. Experimental results show that the proposed framework can outperform related work qualitatively and quantitatively. Christina Schwarz-Gsaxner, Shohei Mori, Dieter Schmalstieg, Jan Egger, Gerhard Paar, Werner Bailer, Denis Kalkofen |
3DV | 2 |
| 2024 | Free-Viewpoint Visual Inspection via 3D Gaussian Splatting for Direct Template MatchingabstractMachine vision systems play a pivotal role in streamlining manufacturing processes, notably in quality control through automatic in-line visual inspections. A common practice for inspecting parts, components, and final products is to use a master part benchmark for quality comparison. However, challenges arise when objects enter inspection points in unintended orientations. This misalignment potentially leads to erroneous decisions by automated systems, resulting in additional checkpoints or wastage affecting the production rate. To tackle this issue, we propose a visual inspection pipeline that leverages recent machine learning-based approaches to compare the inspection target and a master part virtually oriented to the same perspective. Specifically, we suggest combining 3D Gaussian Splatting and DUSt3R as a practical solution. Our approach demonstrates its efficacy in real-world scenarios through testing on three mock parts and a real industrial component. Kenta Ito, Shiori Ueda, Shohei Mori, Junichi Sugano, Hideyuki Adachi, Hideo Saito 0001 |
IECON | 3 |
| 2024 | Neural Bokeh: Learning Lens Blur for Computational Videography and Out-of-Focus Mixed RealityabstractWe present Neural Bokeh, a deep learning approach for synthesizing convincing out-of-focus effects with applications in Mixed Reality (MR) image and video compositing. Unlike existing approaches that solely learn the amount of blur for out-of-focus areas, our approach captures the overall characteristic of the bokeh to enable the seamless integration of rendered scene content into real images, ensuring a consistent lens blur over the resulting MR composition. Our method learns spatially varying blur shapes, i.e., bokeh, from a dataset of real images acquired using the physical camera that is used to capture the photograph or video of the MR composition. Accordingly, those learned blur shapes mimic the characteristics of the physical lens. As the run-time and the resulting quality of Neural Bokeh increase with the resolution of input images, we employ low-resolution images for the MR view finding at runtime and high-resolution renderings for compositing with high-resolution photographs or videos in an offline process. We envision a variety of applications, including visual enhancement of image and video compositing containing creative utilization of out-of-focus effects. David Mandl, Shohei Mori, Peter Mohr, Yifan Peng 0001, Tobias Langlotz, Dieter Schmalstieg, Denis Kalkofen |
VR | 2 |
| 2023 | State-Aware Configuration Detection for Augmented Reality Step-by-Step TutorialsabstractPresenting tutorials in augmented reality is a compelling application area, but previous attempts have been limited to objects with only a small numbers of parts. Scaling augmented reality tutorials to complex assemblies of a large number of parts is difficult, because it requires automatically discriminating many similar-looking object configurations, which poses a challenge for current object detection techniques. In this paper, we seek to lift this limitation. Our approach is inspired by the observation that, even though the number of assembly steps may be large, their order is typically highly restricted: Some actions can only be performed after others. To leverage this observation, we enhance a state-of-the-art object detector to predict the current assembly state by conditioning on the previous one, and to learn the constraints on consecutive states. This learned ‘consecutive state prior’ helps the detector disambiguate configurations that are otherwise too similar in terms of visual appearance to be reliably discriminated. Via the state prior, the detector is also able to improve the estimated probabilities that a state detection is correct. We experimentally demonstrate that our technique enhances the detection accuracy for assembly sequences with a large number of steps and on a variety of use cases, including furniture, Lego and origami. Additionally, we demonstrate the use of our algorithm in an interactive augmented reality application. Ana Stanescu 0003, Peter Mohr, Mateusz Kozinski, Shohei Mori, Dieter Schmalstieg, Denis Kalkofen |
ISMAR | 4 |
| 2023 | High-Quality Virtual Single-Viewpoint Surgical Video: Geometric Autocalibration of Multiple Cameras in Surgical Lights
Yuna Kato, Mariko Isogawa, Shohei Mori, Hideo Saito 0001, Hiroki Kajita, Yoshifumi Takatsume |
MICCAI (9) | 3 |
| 2023 | Multi-Layer Scene Representation from Composed Focal StacksabstractMulti-layer images are a powerful scene representation for high-performance rendering in virtual/augmented reality (VR/AR). The major approach to generate such images is to use a deep neural network trained to encode colors and alpha values of depth certainty on each layer using registered multi-view images. A typical network is aimed at using a limited number of nearest views. Therefore, local noises in input images from a user-navigated camera deteriorate the final rendering quality and interfere with coherency over view transitions. We propose to use a focal stack composed of multi-view inputs to diminish such noises. We also provide theoretical analysis for ideal focal stacks to generate multi-layer images. Our results demonstrate the advantages of using focal stacks in coherent rendering, memory footprint, and AR-supported data capturing. We also show three applications of imaging for VR. Reina Ishikawa, Hideo Saito 0001, Denis Kalkofen, Shohei Mori |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Modified Egocentric Viewpoint for Softer Seated Experience in Virtual RealityabstractUsers in a prolonged experience of virtual reality adopt a sitting position according to their task, as they do in the real world. However, inconsistencies in the haptic feedback from a chair they sit on in the real world and that which is expected in the virtual world decrease the feeling of presence. We aimed to change the perceived haptic features of a chair by shifting the position and angle of the users' viewpoints in the virtual reality environment. The targeted features in this study were seat softness and backrest flexibility. To enhance the seat softness, we shifted the virtual viewpoint using an exponential formula soon after a user's bottom contacted the seat surface. The flexibility of the backrest was manipulated by moving the viewpoint, which followed the tilt of the virtual backrest. These shifts make users feel as if their body moves along with the viewpoint; as a result, they would perceive pseudo-softness or flexibility consistently with the body movement. Based on subjective evaluations, we confirmed that the participants perceived the seat as being softer and the backrest as being more flexible than the actual ones. These results demonstrated that only shifting the viewpoint could change the participants' perceptions of the haptic features of their seats, although significant changes created strong discomfort. Miki Matsumuro, Shohei Mori, Yuta Kataoka, Fumiaki Igarashi, Fumihisa Shibata, Asako Kimura |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Good Keyframes to InpaintabstractDiminished Reality (DR) propagates pixels from a keyframe to subsequent frames for real-time inpainting. Keyframe selection has a significant impact on the inpainting quality, but untrained users struggle to identify good keyframes. Automatic selection is not straightforward either, since no previous work has formalized or verified what determines a good keyframe. We propose a novel metric to select good keyframes to inpaint. We examine the heuristics adopted in existing DR inpainting approaches and derive multiple simple criteria measurable from SLAM. To combine these criteria, we empirically analyze their effect on the quality using a novel representative test dataset. Our results demonstrate that the combined metric selects RGBD keyframes leading to high-quality inpainting results more often than a baseline approach in both color and depth domains. Also, we confirmed that our approach has a better ranking ability of distinguishing good and bad keyframes. Compared to random selections, our metric selects keyframes that would lead to higher-quality and more stably converging inpainting results. We present three DR examples, automatic keyframe selection, user navigation, and marker hiding. Shohei Mori, Dieter Schmalstieg, Denis Kalkofen |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Exemplar-Based Inpainting for 6DOF Virtual Reality PhotosabstractMulti-layer images are currently the most prominent scene representation for viewing natural scenes under full-motion parallax in virtual reality. Layers ordered in diopter space contain color and transparency so that a complete image is formed when the layers are composited in a view-dependent manner. Once baked, the same limitations apply to multi-layer images as to conventional single-layer photography, making it challenging to remove obstructive objects or otherwise edit the content. Object removal before baking can benefit from filling disoccluded layers with pixels from background layers. However, if no such background pixels have been observed, an inpainting algorithm must fill the empty spots with fitting synthetic content. We present and study a multi-layer inpainting approach that addresses this problem in two stages: First, a volumetric area of interest specified by the user is classified with respect to whether the background pixels have been observed or not. Second, the unobserved pixels are filled with multi-layer inpainting. We report on experiments using multiple variants of multi-layer inpainting and compare our solution to conventional inpainting methods that consider each layer individually. Shohei Mori, Dieter Schmalstieg, Denis Kalkofen |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | guitARhero: Interactive Augmented Reality Guitar TutorialsabstractThis paper presents guitARhero, an Augmented Reality application for interactively teaching guitar playing to beginners through responsive visualizations overlaid on the guitar neck. We support two types of visual guidance, a highlighting of the frets that need to be pressed and a 3D hand overlay, as well as two display scenarios, one using a desktop magic mirror and one using a video see-through head-mounted display. We conducted a user study with 20 participants to evaluate how well users could follow instructions presented with different guidance and display combinations and compare these to a baseline where users had to follow video instructions. Our study highlights the trade-off between the provided information and visual clarity affecting the user's ability to interpret and follow instructions for fine-grained tasks. We show that the perceived usefulness of instruction integration into an HMD view highly depends on the hardware capabilities and instruction details. Lucchas Ribeiro Skreinig, Denis Kalkofen, Ana Stanescu 0003, Peter Mohr, Frank Heyen, Shohei Mori, Michael Sedlmair, Dieter Schmalstieg, Alexander Plopski |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Exploring Pseudo-Weight in Augmented Reality Extended DisplaysabstractAugmented reality (AR) allows us to wear virtual displays that are registered to our bodies and devices. Such virtually extendable displays, or AR extended displays (AREDs), provide personal display space and are free from physical restrictions. Existing work has explored the new design space to improve user experience and efficiency. Contrary to this direction, we focus on the weight that the user perceives from AREDs, even though they are virtual and have no physical weight. Our user study results show evidence that AREDs can be a source of pseudo-weight, in addition to that of a handheld physical display device. We also systematically evaluate the perceived weight changes depending on the layout and delay in the visualization system. These findings are similar to those in existing pseudo-haptics research. However, we found such behavior in pseudo-weight for a real device and virtual visual stimuli in the air, which differentiates our research from previous work. Shohei Mori, Yuta Kataoka, Satoshi Hashiguchi |
VR | 1 |
| 2022 | Video See-Through Mixed Reality with Focus CuesabstractThis work introduces the first approach to video see-through mixed reality with full support for focus cues. By combining the flexibility to adjust the focus distance found in varifocal designs with the robustness to eye-tracking error found in multifocal designs, our novel display architecture reliably delivers focus cues over a large workspace. In particular, we introduce gaze-contingent layered displays and mixed reality focal stacks, an efficient representation of mixed reality content that lends itself to fast processing for driving layered displays in real time. We thoroughly evaluate this approach by building a complete end-to-end pipeline for capture, render, and display of focus cues in video see-through displays that uses only off-the-shelf hardware and compute components. Christoph Ebner, Shohei Mori, Peter Mohr, Yifan Peng 0001, Dieter Schmalstieg, Gordon Wetzstein, Denis Kalkofen |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Neural Cameras: Learning Camera Characteristics for Coherent Mixed Reality RenderingabstractCoherent rendering is important for generating plausible Mixed Reality presentations of virtual objects within a user’s real-world environment. Besides photo-realistic rendering and correct lighting, visual coherence requires simulating the imaging system that is used to capture the real environment. While existing approaches either focus on a specific camera or a specific component of the imaging system, we introduce Neural Cameras, the first approach that jointly simulates all major components of an arbitrary modern camera using neural networks. Our system allows for adding new cameras to the framework by learning the visual properties from a database of images that has been captured using the physical camera. We present qualitative and quantitative results and discuss future direction for research that emerge from using Neural Cameras. David Mandl, Peter M. Roth, Tobias Langlotz, Christoph Ebner, Shohei Mori, Stefanie Zollmann, Peter Mohr, Denis Kalkofen |
ISMAR | 5 |
| 2021 | Visualization Techniques in Augmented Reality: A Taxonomy, Methods and PatternsabstractIn recent years, the development of Augmented Reality (AR) frameworks made AR application development widely accessible to developers without AR expert background. With this development, new application fields for AR are on the rise. This comes with an increased need for visualization techniques that are suitable for a wide range of application areas. It becomes more important for a wider audience to gain a better understanding of existing AR visualization techniques. In this article we provide a taxonomy of existing works on visualization techniques in AR. The taxonomy aims to give researchers and developers without an in-depth background in Augmented Reality the information to successively apply visualization techniques in Augmented Reality environments. We also describe required components and methods and analyze common patterns. Stefanie Zollmann, Tobias Langlotz, Raphaël Grasset, Wei Hong Lo, Shohei Mori, Holger Regenbrecht |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Mixed Reality Light Fields for Interactive Remote AssistanceabstractRemote assistance represents an important use case for mixed reality. With the rise of handheld and wearable devices, remote assistance has become practical in the wild. However, spontaneous provisioning of remote assistance requires an easy, fast and robust approach for capturing and sharing of unprepared environments. In this work, we make a case for utilizing interactive light fields for remote assistance. We demonstrate the advantages of object representation using light fields over conventional geometric reconstruction. Moreover, we introduce an interaction method for quickly annotating light fields in 3D space without requiring surface geometry to anchor annotations. We present results from a user study demonstrating the effectiveness of our interaction techniques, and we provide feedback on the usability of our overall system. Peter Mohr, Shohei Mori, Tobias Langlotz, Bruce H. Thomas, Dieter Schmalstieg, Denis Kalkofen |
CHI | 2 |
| 2020 | Video-Annotated Augmented Reality Assembly TutorialsabstractWe present a system for generating and visualizing interactive 3D Augmented Reality tutorials based on 2D video input, which allows viewpoint control at runtime. Inspired by assembly planning, we analyze the input video using a 3D CAD model of the object to determine an assembly graph that encodes blocking relationships between parts. Using an assembly graph enables us to detect assembly steps that are otherwise difficult to extract from the video, and generally improves object detection and tracking by providing prior knowledge about movable parts. To avoid information loss, we combine the 3D animation with relevant parts of the 2D video so that we can show detailed manipulations and tool usage that cannot be easily extracted from the video. To further support user orientation, we visually align the 3D animation with the real-world object by using texture information from the input video. We developed a presentation system that uses commonly available hardware to make our results accessible for home use and demonstrate the effectiveness of our approach by comparing it to traditional video tutorials. Masahiro Yamaguchi 0002, Shohei Mori, Peter Mohr, Markus Tatzgern, Ana Stanescu 0003, Hideo Saito 0001, Denis Kalkofen |
UIST | 2 |
| 2020 | InpaintFusion: Incremental RGB-D Inpainting for 3D ScenesabstractState-of-the-art methods for diminished reality propagate pixel information from a keyframe to subsequent frames for real-time inpainting. However, these approaches produce artifacts, if the scene geometry is not sufficiently planar. In this article, we present InpaintFusion, a new real-time method that extends inpainting to non-planar scenes by considering both color and depth information in the inpainting process. We use an RGB-D sensor for simultaneous localization and mapping, in order to both track the camera and obtain a surfel map in addition to RGB images. We use the RGB-D information in a cost function for both the color and the geometric appearance to derive a global optimization for simultaneous inpainting of color and depth. The inpainted depth is merged in a global map by depth fusion. For the final rendering, we project the map model into image space, where we can use it for effects such as relighting and stereo rendering of otherwise hidden structures. We demonstrate the capabilities of our method by comparing it to inpainting results with methods using planar geometric proxies. Shohei Mori, Okan Erat, Wolfgang Broll, Hideo Saito 0001, Dieter Schmalstieg, Denis Kalkofen |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Removing fences from sweep motion videos using global 3D reconstruction and fence-aware light field renderingabstractDiminishing the appearance of a fence in an image is a challenging research area due to the characteristics of fences (thinness, lack of texture, etc.) and the need for occluded background restoration. In this paper, we describe a fence removal method for an image sequence captured by a user making a sweep motion, in which occluded background is potentially observed. To make use of geometric and appearance information such as consecutive images, we use two well-known approaches: structure from motion and light field rendering. Results using real image sequences show that our method can stably segment fences and preserve background details for various fence and background combinations. A new video without the fence, with frame coherence, can be successfully provided. Chanya Lueangwattana, Shohei Mori, Hideo Saito 0001 |
Comput. Vis. Media | 2 |
| 2018 | BrightView: Increasing Perceived Brightness of Optical See-Through Head-Mounted Displays Through Unnoticeable Incident Light ReductionabstractOptical See-Through Head-Mounted Displays (OST-HMDs) lose the visibility of virtual contents under bright environment illumination due to their see-through nature. We demonstrate how a liquid crystal (LC) filter attached to an OST-HMD can be used to dynamically increase the perceived brightness of virtual content without impacting the perceived brightness of the real scene. We present a prototype OST-HMD that continuously adjusts the opacity of the LC filter to attenuate the environment light without users becoming aware of the change. Consequently, virtual content appears to be brighter. The proposed approach is evaluated in psychophysical experiments in three scenes, with 16, 31, and 31 participants, respectively. The participants were asked to compare the magnitude of brightness changes of both real and virtual objects, before and after dimming the LC filter over a period of 5, 10, and 20 seconds. The results showed that the participants felt increases in the brightness of virtual objects while they were less conscious of reductions of the real scene luminance. These results provide evidence for the effectiveness of our display design. Our design can be applied to a wide range of OST-HMDs to improve the brightness and hence realism of virtual content in augmented reality applications. Shohei Mori, Sei Ikeda, Alexander Plopski, Christian Sandor |
VR | 1 |
| 2018 | Perceived weight of a rod under augmented and diminished reality visual effectsabstractWe can use augmented reality (AR) and diminished reality (DR) in combination, in practice. However, to the best of our knowledge, there is no research on the validation of the cross-modal effects in AR and DR. Our research interest here is to investigate how this continuous visual changes between AR and DR would change our weight sensation of an object. In this paper, we built a system that can continuously extend and reduce the amount of visual entity of real objects using AR and DR renderings to confirm that users can perceive things heavier and lighter than they actually are in the same manner as SWI. Different from the existing research where either AR or DR visual effects were used, we validated one of cross-modal effects in the context of both continuous AR and DR visuo-haptic. Regarding the weight sensation, we found that such cross-modal effect can be approximated with a continuous linear relationship between the weight and length of real objects. Our experimental results suggested that the weight sensation is closely related to the positions of the center of gravity (CoG) and perceived CoG positions lie within the object's entity under the examined conditions. Satoshi Hashiguchi, Shohei Mori, Miho Tanaka, Fumihisa Shibata, Asako Kimura |
VRST | 2 |
| 2017 | Diminished hand: A diminished reality-based work area visualizationabstractLive instructors perspective videos are useful to present intuitive visual instructions for trainees in medical and industrial settings. In such videos, the instructors hands often hide the work area. In this demo, we present a diminished hand for visualizing the work area hidden by hands by capturing the work area with multiple cameras. To achieve the diminished reality, we use a light field rendering technique, in which light rays avoid passing through penalty points set in the unstructured light fields reconstructed from the multiple viewpoint images. Shohei Mori, Momoko Maezawa, Naoto Ienaga, Hideo Saito 0001 |
VR | 1 |
| 2011 | Enabling on-set stereoscopic MR-based previsualization for 3D filmmakingabstractPreViz is a computer graphics movie that represents desired scenes in preproduction and is useful for sharing the director's imagination among crews. Therefore, film industry looks upon PreViz as one of the most important processes in filmmaking as the processes become complicated [The Previs Society 2011]. Stereoscopic 3D (S3D) movie production not only changed ways of expressiveness but also changed ways of production and tools. Consequently, it requires more careful planning and PreViz is still important. This paper describes a prototype of S3D MR-PreViz system for HD S3D PreViz shooting using mixed reality (MR) technologies in order to facilitate PreViz making in S3D. This system is designed as an extension of MR-PreViz for classic filmmaking [Ichikari et al. 2010]. Shohei Mori, Ryosuke Ichikari, Fumihisa Shibata, Asako Kimura, Hideyuki Tamura |
SIGGRAPH Asia Sketches | 1 |