Maki Sugimoto

dblp:33/3150 · DBLP profile ↗
← Back
53ranked-venue papers
0as first author
10since 2021 · last 2024
0000-0002-8383-9228ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 43 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 since 2021Artificial intelligence and machine learning · 2
YearPublicationVenuePosition
2024 Facial Gesture Classification with Few-shot Learning Using Limited Calibration Data from Photo-reflective Sensors on Smart Eyewear
abstract
This study investigates smart eyewear for facial gesture classification with low user calibration costs.The smart eyewear is equipped with low-cost, comfortable, energy-efficient photo-reflective sensors, which can detect changes in facial muscle movements.Although the sensor output is useful for facial gesture classification, individual user calibration is considered necessary.Moreover, re-calibration is required whenever the user's wearing position changes.Therefore, reducing the calibration cost is crucial for the wider applicability of the eyewear.To address this issue, we propose a few-shot domain adaptation approach using Convolutional Neural Networks (CNN).We evaluate the accuracy of classifying eight gestures with data augmentation and a supervised contrastive loss.Data augmentation is employed to make the model more robust to noise, while the supervised contrastive loss is introduced to learn user-invariant features.Our approach with data augmentation achieves robust gesture classification, with an average accuracy of 93.46% (SD = 8.34%) for three emotion-related gestures without user-specific data.Furthermore, for user-independent training, we demonstrated that using few-shot learning with pre-trained models and only four repetitions of calibration data per gesture achieved a practical accuracy of 91.38% (SD = 6.03%), showing that a small amount of user-specific data is sufficient for the accurate classification.Also, it works under different wearing conditions, achieving an accuracy of 90.19% (SD = 3.56%).These results illustrate the potential of our method to improve the practicality of smart eyewear for facial expression recognition in cases of limited user data, making it more accessible and user-friendly.
Katsutoshi Masai, Maki Sugimoto, Brian Kenji Iwana
MUM2
2024 Message from the ISMAR 2024 Science and Technology Program Chairs and TVCG Guest Editors
abstract
In this special issue of IEEE Transactions on Visualization and Computer Graphics (TVCG), we are pleased to present the journal papers from the 23rd IEEE International Symposium on Mixed and Augmented Reality (ISMAR 2024), which will be held as a hybrid conference between October 21 and 25, 2024 in the Greater Seattle Area, USA. ISMAR continues the over twenty-year long tradition of IWAR, ISMR, and ISAR, and is the premier conference for Mixed and Augmented Reality in the world.
Ulrich Eck, Maki Sugimoto, Misha Sra, Markus Tatzgern, Jeanine K. Stefanucci, Ian Williams 0001
IEEE Trans. Vis. Comput. Graph.2
2024 Sensory Attenuation With a Virtual Robotic Arm Controlled Using Facial Movements
abstract
When humans generate stimuli voluntarily, they perceive the stimuli more weakly than those produced by others, which is called sensory attenuation (SA). SA has been investigated in various body parts, but it is unclear whether an extended body induces SA. This study investigated the SA of audio stimuli generated by an extended body. SA was assessed using a sound comparison task in a virtual environment. We prepared robotic arms as extended bodies, and the robotic arms were controlled by facial movements. To evaluate the SA of robotic arms, we conducted two experiments. Experiment 1 investigated the SA of the robotic arms under four conditions. The results showed that robotic arms manipulated by voluntary actions attenuated audio stimuli. Experiment 2 investigated the SA of the robotic arm and innate body under five conditions. The results indicated that the innate body and robotic arm induced SA, while there were differences in the sense of agency between the innate body and robotic arm. Analysis of the results indicated three findings regarding the SA of the extended body. First, controlling the robotic arm with voluntary actions in a virtual environment attenuates the audio stimuli. Second, there were differences in the sense of agency related to SA between extended and innate bodies. Third, the SA of the robotic arm was correlated with the sense of body ownership.
Masaaki Fukuoka, Fumihiko Nakamura, Adrien Verhulst, Masahiko Inami, Michiteru Kitazaki, Maki Sugimoto
IEEE Trans. Vis. Comput. Graph.6
2023 Facial Expression Recognition by Photo-Reflective Sensors Considering Time Series and Head Posture
abstract
There is a method to recognize facial expressions of Head-Mounted Display (HMD) wearers by machine learning of reflection intensity information from photo-reflective sensors embedded into the interior of an HMD [1]. This study evaluates whether facial expression recognition accuracy can be improved by using a learning model that considers temporal changes in sensor values. We assessed whether facial expression recognition accuracy could be improved by adding the head posture data acquired from the Inertial Measurement Unit (IMU) in the HMD to the discriminator input and performing time-series learning. The experimental results showed that PRS-based facial expression recognition with time-series data was more accurate than without. The multimodal recognition using the reflection intensity and head posture data was slightly more accurate than the discrimination using only the reflection intensity information. It was especially effective for the learning condition without considering the time-series.
Yuki Nakabayashi, Fumihiko Nakamura, Maki Sugimoto
APCC3
2023 Exploring Enhancements towards Gaze Oriented Parallel Views in Immersive Tasks
abstract
Parallel view is a technique that allows a VR user to see multiple locations at a time. It enables the user to control several remote or virtual body parts while seeing parallel views to solve synchronous tasks. However, these techniques only explored the benefits and drawbacks of a user performing different tasks. In this paper, we explored enhancements on a singular or asynchronous task by utilizing information obtained in parallel views. We developed three prototypes where parallel views are fixed, moving in symmetric order, or following the user's eye gaze. We conducted a user study to compare each prototype against traditional VR (without parallel views) in three types of tasks: object search and interaction tasks in a 1) simple environment and 2) complex environment, and 3) object distances estimation task. We found parallel views improved multi-embodiment while each technique helped different tasks. No parallel view provided a clean interface, thus improving spatial presence, mental effort, and user performance. However, participants' feedback highlighted potential usefulness and a lower physical effort by using parallel views to solve complicated tasks.
Theophilus Teo, Kuniharu Sakurada, Maki Sugimoto
VR3
2023 Analyzing the Effect of Diverse Gaze and Head Direction on Facial Expression Recognition With Photo-Reflective Sensors Embedded in a Head-Mounted Display
abstract
As one of the facial expression recognition techniques for Head-Mounted Display (HMD) users, embedded photo-reflective sensors have been used. In this paper, we investigate how gaze and face directions affect facial expression recognition using the embedded photo-reflective sensors. First, we collected a dataset of five facial expressions (Neutral, Happy, Angry, Sad, Surprised) while looking in diverse directions by moving 1) the eyes and 2) the head. Using the dataset, we analyzed the effect of gaze and face directions by constructing facial expression classifiers in five ways and evaluating the classification accuracy of each classifier. The results revealed that the single classifier that learned the data for all gaze points achieved the highest classification performance. Then, we investigated which facial part was affected by the gaze and face direction. The results showed that the gaze directions affected the upper facial parts, while the face directions affected the lower facial parts. In addition, by removing the bias of facial expression reproducibility, we investigated the pure effect of gaze and face directions in three conditions. The results showed that, in terms of gaze direction, building classifiers for each direction significantly improved the classification accuracy. However, in terms of face directions, there were slight differences between the classifier conditions. Our experimental results implied that multiple classifiers corresponding to multiple gaze and face directions improved facial expression recognition accuracy, but collecting the data of the vertical movement of gaze and face is a practical solution to improving facial expression recognition accuracy.
Fumihiko Nakamura, Masaaki Murakami, Katsuhiro Suzuki, Masaaki Fukuoka, Katsutoshi Masai, Maki Sugimoto
IEEE Trans. Vis. Comput. Graph.6
2022 Consistent Smile Intensity Estimation from Wearable Optical Sensors
abstract
Smiling plays a crucial role in human communication. It is the most frequent expression shown in daily life. Smile analysis usually employs computer vision-based methods that use data sets annotated by experts. However, cameras have space constraints in most realistic scenarios due to occlusions. Wearable electromyography is a promising alternative; however, issue of user comfort is a barrier to long-term use. Other wearable-based methods can detect smiles, but they lack consistency because they use subjective criteria without expert annotation. We investigate a wearable-based method that uses optical sensors for consistent smile intensity estimation while reducing manual annotation cost. First, we use a state-of-art computer vision method (OpenFace) to train a regression model to estimate smile intensity from sensor data. Then, we compare the estimation result to that of OpenFace. We also compared their results to human annotation. The results show that the wearable method has a higher matching coefficient (r=0.67) with human annotated smile intensity than OpenFace (r=0.56). Also, when the sensor data and OpenFace output were fused, the multimodal method produced estimates closer to human annotation (r=0.74). Finally, we investigate how the synchrony of smile dynamics among subjects and their average smile intensity are correlated to assess the potential of wearable smile intensity estimation.
Katsutoshi Masai, Monica Perusquía-Hernández, Maki Sugimoto, Shiro Kumano, Toshitaka Kimura
ACII3
2022 Human Latent Metrics: Perceptual and Cognitive Response Correlates to Distance in GAN Latent Space for Facial Images
abstract
Generative adversarial networks (GANs) generate high-dimensional vector spaces (latent spaces) that can interchangeably represent vectors as images. Advancements have extended their ability to computationally generate images indistinguishable from real images such as faces, and more importantly, to manipulate images using their inherit vector values in the latent space. This interchangeability of latent vectors has the potential to calculate not only the distance in the latent space, but also the human perceptual and cognitive distance toward images, that is, how humans perceive and recognize images. However, it is still unclear how the distance in the latent space correlates with human perception and cognition. Our studies investigated the relationship between latent vectors and human perception or cognition through psycho-visual experiments that manipulates the latent vectors of face images. In the perception study, a change perception task was used to examine whether participants could perceive visual changes in face images before and after moving an arbitrary distance in the latent space. In the cognition study, a face recognition task was utilized to examine whether participants could recognize a face as the same, even after moving an arbitrary distance in the latent space. Our experiments show that the distance between face images in the latent space correlates with human perception and cognition for visual changes in face imagery, which can be modeled with a logistic function. By utilizing our methodology, it will be possible to interchangeably convert between the distance in the latent space and the metric of human perception and cognition, potentially leading to image processing that better reflects human perception and cognition.
Kye Shimizu, Naoto Ienaga, Kazuma Takada, Maki Sugimoto, Shunichi Kasahara
SAP4
2022 Foreword to the Special Section on the International Conference on Artificial Reality and Telexistence & Eurographics Symposium on Virtual Environments (ICAT-EGVE 2020)
Ferran Argelaguet, Ryan P. McMahan, Maki Sugimoto
Comput. Graph.3
2021 Investigating Textual Visual Sound Effects in a Virtual Environment and their impacts on Object Perception and Sound Perception
abstract
In comics, Textual Sound Effects (TE) can describe sounds, but also actions, events, etc. TE could be used in Virtual Environment to efficiently create an easily recognizable scene and add more information to objects at a relatively low design cost. We investigate the impact of TE in a Virtual Environment on objects’ material perception (on category and properties) and on sound perception (on volume [dB] and spatial position). Participants (N=13, repeated measures) categorized metallic and wooden spheres and significantly changed their reaction time depending on the TE congruence with the spheres’ material/sound. They then rated a sphere’s properties (i.e., wetness, warmness, softness, smoothness, and dullness) and significantly changed their rating depending on the TE. When comparing 2 sound volumes, they perceived a sound associated with a shrinking TE as less loud and a sound associated with a growing TE as louder. When locating an audio source location, they located it significantly closer to a TE.
Thibault Fabre, Adrien Verhulst, Alfonso Balandra, Maki Sugimoto, Masahiko Inami
ISMAR4
2020 Face Commands - User-Defined Facial Gestures for Smart Glasses
abstract
We propose the use of face-related gestures involving the movement of the face, eyes, and head for augmented reality (AR). This technique allows us to use computer systems via hands-free, discreet interactions. In this paper, we present an elicitation study to explore the proper use of facial gestures for daily tasks in the context of a smart home. We used Amazon Mechanical Turk to conduct this study (N=37). Based on the proposed gestures, we report usage scenarios and complexity, proposed associations between gestures/tasks, a user-defined gesture set, and insights from the participants. We also conducted a technical feasibility study (N=13) with participants using smart eyewear to consider their uses in daily life. The device has 16 optical sensors and an inertial measurement unit (IMU). We can potentially integrate the system into optical see-through displays or other smart glasses. The results demonstrate that the device can detect eight temporal face-related gestures with a mean F1 score of 0.911 using a convolutional neural network (CNN). We also report the results of user-independent training and a one-hour recording of the experimenter testing two of the gestures.
Katsutoshi Masai, Kai Kunze, Daisuke Sakamoto, Yuta Sugiura, Maki Sugimoto
ISMAR5
2020 Digital Full-Face Mask Display with Expression Recognition using Embedded Photo Reflective Sensor Arrays
abstract
This paper presents a thin digital full-face mask display that can reflect an entire facial expression of a user onto an avatar to support augmented face-to-face communication in real environments. Although camera-based facial expression recognition technology has enabled people to augment their faces with avatars, application was limited to face-to-face communication in virtual environments. To enable digital facial augmentation with an avatar in a real space, we propose a digital face mask display system that integrates a lightweight flexible display with a thin facial expression recognition system. The thin wearable facial expression recognition system was implemented with photo reflective sensor arrays which can measure facial expressions at 40 feature points distributed across an entire face. We investigated a ten-class facial expression identification model based on an SVM training algorithm. The trained model achieved an average accuracy of 79% when identifying the facial expressions of multiple users. User experiments indicated that the proposed thin digital full-face mask display allows the wearer to control the facial expression of the avatar with a fast response rate and create a positive sense of self-agency and self-ownership toward the augmented avatar face.
Yoshinari Takegawa, Yutaka Tokuda, Akino Umezawa, Katsuhiro Suzuki, Katsutoshi Masai, Yuta Sugiura, Maki Sugimoto, Diego Martínez 0001, Sriram Subramanian, Keiji Hirata 0001
ISMAR7
2020 Classification of Spontaneous and Posed Smiles by Photo-reflective Sensors Embedded with Smart Eyewear
abstract
Smile is one of the representative emotional expressions which is observed frequently in daily life and essential for various non-verbal communications. People make spontaneous smiles and intentional ones. It is important to guess properly whether a person is making a smile spontaneously or intentionally to understand the meaning of smiles. In this study, we propose a smile classification system with smart eyewear that equips photo-reflective sensors and examines whether we can distinguish two types of smiles; spontaneous smiles caused by funny videos and posed smiles evoked by instructions. We extract geometric features: reflection intensity distribution of sensors and temporal features in a time axis. By applying for Support Vector Machine, we observed 94.6% as the mean accuracy among 12 participants when we used both geometric and temporal features with user-dependent training. The result suggested that we can distinguish between spontaneous and posed smile by the sensors embedded with the smart eyewear.
Chisa Saito, Katsutoshi Masai, Maki Sugimoto
TEI3
2019 CoSummary: adaptive fast-forwarding for surgical videos by detecting collaborative scenes using hand regions and gaze positions
abstract
This paper presents CoSummary, an adaptive video fast-forwarding technique for browsing surgical videos recorded by wearable cameras. Current wearable technologies allow us to record complex surgical skills, however, an efficient browsing technique for these videos is not well established. In order to assist browsing surgical videos, our study focuses on adaptively changing playback speeds through the learning and detecting collaborative scenes based on surgeon hand placement and gaze information. Our evaluation shows that the proposed method is able to highlight important collaborative scenes and skip less important scenes during surgical procedures. We have also performed a subjective study with surgeons in order to have professional feedback. The results confirmed the effectiveness of the proposed method in comparison to uniform video fast-forwarding.
Irshad Abibouraguimane, Kakeru Hagihara, Keita Higuchi, Yuta Itoh 0001, Yoichi Sato 0001, Tetsu Hayashida, Maki Sugimoto
IUI7
2019 Remapping a Third Arm in Virtual Reality
abstract
This paper presents development on a conceptual method to remap supernumerary limbs using Virtual Reality (VR) as a platform for experimentation. Our VR system allows users to control a third arm through their own limbs such as their head, arms, and feet with the ability to switch between them. To realize and experiment with our remapping method, we used the Oculus Rift in conjunction with OptiTrack to track users in a room-scaled virtual environment. We present some initial findings from a small pilot study and conclude with suggestions for future work.
Adam Drogemuller, Adrien Verhulst, Masahiko Inami, Benjamin Volmer, Maki Sugimoto, Bruce H. Thomas
VR5
2019 Shared Body by Action Integration of Two Persons: Body Ownership, Sense of Agency and Task Performance
abstract
Humans have one own body. However, we can share a body with two different persons in a virtual environment. We have developed a shared body as an avatar that is controlled by two persons' actions. Movements of two subjects were continuously captured, and integrated into the avatar's motion with the ratios of 0:100,25:75, 50:50, 75:25, 100:0. They were not aware of integration ratios, and asked to reach cubes with the right hand. They felt body ownership and sense of agency to the shared avatar more when the responsible ratio was higher. The reaching path of the avatar's hand was shorter in the shared body condition (75:25) than the single body condition (100:0). These results suggest that we have some of body ownership and sense of agency to the shared body, and the task performance is improved by the shared body.
Takayoshi Hagiwara, Maki Sugimoto, Masahiko Inami, Michiteru Kitazaki
VR2
2019 Scrambled Body: A Method to Compare Full Body Illusion and Illusory Body Ownership of Body Parts
abstract
Humans can feel as if a fake body is their own body in the illusory body ownership. The illusion can be induced to a full body by a visual-tactile synchronicity or visual-motor synchronicity. In our previous study, illusory full body ownership of invisible body was elicited by synchronous movements of virtual gloves and socks. In this study, we aimed to investigate whether the spatial relationship is necessary for the full body illusion using a stimulus of scrambled body. In the scrambled body, positions of gloves and socks were scrambled. The results suggest that spatial relationship of body parts is necessary for the full body illusion.
Ryota Kondo, Maki Sugimoto, Masahiko Inami, Michiteru Kitazaki
VR2
2019 Towards Robot Arm Training in Virtual Reality Using Partial Least Squares Regression
abstract
Robot assistance can reduce the user's workload of a task. However, the robot needs to be programmed or trained on how to assist the user. Virtual Reality (VR) can be used to train and validate the actions of the robot in a safer and cheaper environment. In this paper, we examine how a robotic arm can be trained using Coloured Petri Nets (CPN) and Partial Least Squares Regression (PLSR). Based upon these algorithms, we discuss the concept of using the user's acceleration and rotation as a sufficient means to train a robotic arm for a procedural task in VR. We present a work-in-progress system for training robotic limbs using VR as a cost effective and safe medium for experimentation. Additionally, we propose PLSR data that could be considered for training data analysis.
Benjamin Volmer, Adrien Verhulst, Masahiko Inami, Adam Drogemuller, Maki Sugimoto, Bruce H. Thomas
VR5
2018 Intra-/inter-user adaptation framework for wearable gesture sensing device
abstract
The photo reflective sensor (PRS), a tiny distant-measurement module, is a popular electronic component widely used in wearable user-interfaces. An unavoidable issue of such wearable PRS devices in practical use is the need of user-independent training to have high gesture recognition accuracy. Each new user has to re-train a device by providing new training data (we call the inter-user setup). Even worse, re-training is also necessary ideally every time when the same user re-wears the device (we call the intra-user setup). In this paper, we propose a domain adaptation framework to reduce this training cost of users. Specifically, we adapt a pre-trained convolutional neural network (CNN) for both inter-user and intra-user setups to maintain the recognition accuracy high. We demonstrate, with an actual PRS device, that our framework significantly improves the average classification accuracy of the intra-user and inter-user setups up to 87.43% and 80.06% against the baseline (non-adapted) setups with the accuracy 68.96% and 63.26% respectively.
Kosuke Kikui, Yuta Itoh 0001, Makoto Yamada, Yuta Sugiura, Maki Sugimoto
UbiComp5
2018 Illusory Body Ownership Between Different Body Parts: Synchronization of Right Thumb and Right Arm
abstract
Illusory body ownership can be induced by visual-tactile stimulation or visual-motor synchronicity. We aimed to test whether a right thumb could be remapped to a virtual right arm and illusory body ownership of the virtual arm induced through synchronous movements of the right thumb and the virtual right arm. We presented the virtual right arm in synchronization with movements of a participant's right thumb on a head-mounted display (HMD). We found that the participants felt as though their right thumb became the right arm, and that the right arm belonged to their own body.
Ryota Kondo, Maki Sugimoto, Kouta Minamizawa, Masahiko Inami, Michiteru Kitazaki, Yamato Tani
VR2
2018 HySAR: Hybrid Material Rendering by an Optical See-Through Head-Mounted Display with Spatial Augmented Reality Projection
abstract
Spatial augmented reality (SAR) pursues realism in rendering materials and objects. To advance this goal, we propose a hybrid SAR (HySAR) that combines a projector with optical see-through head-mounted displays (OST-HMD). In an ordinary SAR scenario with co-located viewers, the viewers perceive the same virtual material on physical surfaces. In general, the material consists of two components: a view-independent (VI) component such as diffuse reflection, and a view-dependent (VD) component such as specular reflection. The VI component is static over viewpoints, whereas the VD should change for each viewpoint even if a projector can simulate only one viewpoint at one time. In HySAR, a projector only renders the static VI components. In addition, the OST-HMD renders the dynamic VD components according to the viewer's current viewpoint. Unlike conventional SAR, the HySAR concept theoretically allows an unlimited number of co-located viewers to see the correct material over different viewpoints. Furthermore, the combination enhances the total dynamic range, the maximum intensity, and the resolution of perceived materials. With proof-of-concept systems, we demonstrate HySAR both qualitatively and quantitatively with real objects. First, we demonstrate HySAR by rendering synthetic material properties on a real object from different viewpoints. Our quantitative evaluation shows that our system increases the dynamic range by 2.24 times and the maximum intensity by 2.12 times compared to an ordinary SAR system. Second, we replicate the material properties of a real object by SAR and HySAR, and show that HySAR outperforms SAR in rendering VD specular components.
Takumi Hamasaki, Yuta Itoh 0001, Yuichi Hiroi, Daisuke Iwai, Maki Sugimoto
IEEE Trans. Vis. Comput. Graph.5
2017 SofTouch: Turning Soft Objects into Touch Interfaces Using Detachable Photo Sensor Modules
Naomi Furui, Katsuhiro Suzuki, Yuta Sugiura, Maki Sugimoto
ICEC4
2017 DanceDanceThumb: Tablet App for Rehabilitation for Carpal Tunnel Syndrome
Takuro Watanabe, Yuta Sugiura, Natsuki Miyata, Koji Fujita, Akimoto Nimura, Maki Sugimoto
ICEC6
2017 DecoTouch: Turning the Forehead as Input Surface for Head Mounted Display
Koki Yamashita, Yuta Sugiura, Takashi Kikuchi, Maki Sugimoto
ICEC4
2017 EarTouch: turning the ear into an input surface
abstract
In this paper, we propose EarTouch, a new sensing technology for ear-based input for controlling applications by slightly pulling the ear and detecting the deformation by an enhanced earphone device. It is envisioned that EarTouch will enable control of applications such as music players, navigation systems, and calendars as an "eyes-free" interface. As for the operation of EarTouch, the shape deformation of the ear is measured by optical sensors. Deformation of the skin caused by touching the ear with the fingers is recognized by attaching optical sensors to the earphone and measuring the distance from the earphone to the skin inside the ear. EarTouch supports recognition of multiple gestures by applying a support vector machine (SVM). EarTouch was validated through a set of user studies.
Takashi Kikuchi, Yuta Sugiura, Katsutoshi Masai, Maki Sugimoto, Bruce H. Thomas
MobileHCI4
2017 Spatial Calibration of Airborne Ultrasound Tactile Display and Projector-Camera System Using Fur Material
abstract
Airborne Ultrasound Tactile Displays (AUTD) are tactile displays that can generate vibrotactile sensation on human skin. Combining an AUTD with a projector-camera system, it is possible to present synchronous visual and haptic stimuli that is physically aligned in the 3D space. To maintain the synchronous sensation as realistic as possible, an accurate crucial to spatial calibration between the AUTD and the projector-camera system is required. This paper thereby proposes a calibration method for the AUTD system by utilizing fur material to recognize the output focal points of an AUTD in the camera coordinate system. Our method simplifies the calibration procedure with calibration error of 2.63 mm.
Shigo Ko, Yuta Itoh 0001, Yuta Sugiura, Takayuki Hoshi, Maki Sugimoto
TEI5
2017 HySAR: Hybrid material rendering by an optical see-through head-mounted display with spatial augmented reality projection
abstract
We propose a hybrid SAR concept combining a projector and Optical See-Through Head-Mounted Displays (OST-HMD). Our proposed hybrid SAR system utilizes OST-HMD as an extra rendering layer to render a view-dependent property in OST-HMDs according to the viewer's viewpoint. Combined with view-independent components created by a static projector, the viewer can see richer material contents. Unlike conventional SAR systems, our system theoretically allows unlimited number of viewers seeing enhanced contents in the same space while keeping the existing SAR experiences. Furthermore, the system enhances the total dynamic range, the maximum intensity, and the resolution of perceived materials. With a proof-of-concept system that consists of a projector and an OST-HMD, we qualitatively demonstrate that our system successfully creates hybrid rendering on a hemisphere object from five horizontal viewpoints. Our quantitative evaluation also shows that our system increases the dynamic range by 2.1 times and the maximum intensity by 1.9 times compared to an ordinary SAR system.
Yuichi Hiroi, Yuta Itoh 0001, Takumi Hamasaki, Daisuke Iwai, Maki Sugimoto
VR5
2017 Monocular focus estimation method for a freely-orienting eye using Purkinje-Sanson images
abstract
We present a method for focal distance estimation of a freely-orienting eye using Purkinje-Sanson (PS) images, which are reflections of light on the inner structures of the eye. Using an infrared camera with a rigidly-fixed LED, our method creates an estimation model based on 3D gaze and the distance between reflections in the PS images that occur on the corneal surface and anterior surface of the eye lens. The distance between these two reflections changes with focus, so we associate that information to the focal distance on a user. Unlike conventional methods that mainly relies on 2D pupil size which is sensitive to scene lighting and the fourth PS image, our method detects the third PS image which is more representative of accommodation. Our feasibility study on a single user with a focal range from 15–45 cm shows that our method achieves mean and median absolute errors of 3.15 and 1.93 cm for a 10-degree viewing angle. The study shows that our method is also tolerant against environment lighting changes.
Yuta Itoh 0001, Jason Orlosky, Kiyoshi Kiyokawa, Toshiyuki Amano, Maki Sugimoto
VR5
2017 Recognition and mapping of facial expressions to avatar by embedded photo reflective sensors in head mounted display
abstract
We propose a facial expression mapping technology between virtual avatars and Head-Mounted Display (HMD) users. HMD allow people to enjoy an immersive Virtual Reality (VR) experience. A virtual avatar can be a representative of the user in the virtual environment. However, the synchronization of the the virtual avatar's expressions with those of the HMD user is limited. The major problem of wearing an HMD is that a large portion of the user's face is occluded, making facial recognition difficult in an HMD-based virtual environment. To overcome this problem, we propose a facial expression mapping technology using retro-reflective photoelectric sensors. The sensors attached inside the HMD measures the distance between the sensors and the user's face. The distance values of five basic facial expressions (Neutral, Happy, Angry, Surprised, and Sad) are used for training the neural network to estimate the facial expression of a user. We achieved an overall accuracy of 88% in recognizing the facial expressions. Our system can also reproduce facial expression change in real-time through an existing avatar using regression. Consequently, our system enables estimation and reconstruction of facial expressions that correspond to the user's emotional changes.
Katsuhiro Suzuki, Fumihiko Nakamura, Jiu Otsuka, Katsutoshi Masai, Yuta Itoh 0001, Yuta Sugiura, Maki Sugimoto
VR7
2017 CheekInput: turning your cheek into an input surface by embedded optical sensors on a head-mounted display
abstract
In this paper, we propose a novel technology called "CheekInput" with a head-mounted display (HMD) that senses touch gestures by detecting skin deformation. We attached multiple photo-reflective sensors onto the bottom front frame of the HMD. Since these sensors measure the distance between the frame and cheeks, our system is able to detect the deformation of a cheek when the skin surface is touched by fingers. Our system uses a Support Vector Machine to determine the gestures: pushing face up and down, left and right. We combined these 4 directional gestures for each cheek to extend 16 possible gestures. To evaluate the accuracy of the gesture detection, we conducted a user study. The results revealed that CheekInput achieved 80.45 % recognition accuracy when gestures were made by touching both cheeks with both hands, and 74.58 % when by touching both cheeks with one hand.
Koki Yamashita, Takashi Kikuchi, Katsutoshi Masai, Maki Sugimoto, Bruce H. Thomas, Yuta Sugiura
VRST4
2017 Evaluation of Facial Expression Recognition by a Smart Eyewear for Facial Direction Changes, Repeatability, and Positional Drift
abstract
This article presents a novel smart eyewear that recognizes the wearer’s facial expressions in daily scenarios. Our device uses embedded photo-reflective sensors and machine learning to recognize the wearer’s facial expressions. Our approach focuses on skin deformations around the eyes that occur when the wearer changes his or her facial expressions. With small photo-reflective sensors, we measure the distances between the skin surface on the face and the 17 sensors embedded in the eyewear frame. A Support Vector Machine (SVM) algorithm is then applied to the information collected by the sensors. The sensors can cover various facial muscle movements. In addition, they are small and light enough to be integrated into daily-use glasses. Our evaluation of the device shows the robustness to the noises from the wearer’s facial direction changes and the slight changes in the glasses’ position, as well as the reliability of the device’s recognition capacity. The main contributions of our work are as follows: (1) We evaluated the recognition accuracy in daily scenes, showing 92.8% accuracy regardless of facial direction and removal/remount. Our device can recognize facial expressions with 78.1% accuracy for repeatability and 87.7% accuracy in case of its positional drift. (2) We designed and implemented the device by taking usability and social acceptability into account. The device looks like a conventional eyewear so that users can wear it anytime, anywhere. (3) Initial field trials in a daily life setting were undertaken to test the usability of the device. Our work is one of the first attempts to recognize and evaluate a variety of facial expressions with an unobtrusive wearable device.
Katsutoshi Masai, Kai Kunze, Yuta Sugiura, Masa Ogata, Masahiko Inami, Maki Sugimoto
ACM Trans. Interact. Intell. Syst.6
2017 Occlusion Leak Compensation for Optical See-Through Displays Using a Single-Layer Transmissive Spatial Light Modulator
abstract
We propose an occlusion compensation method for optical see-through head-mounted displays (OST-HMDs) equipped with a singlelayer transmissive spatial light modulator (SLM), in particular, a liquid crystal display (LCD). Occlusion is an important depth cue for 3D perception, yet realizing it on OST-HMDs is particularly difficult due to the displays' semitransparent nature. A key component for the occlusion support is the SLM-a device that can selectively interfere with light rays passing through it. For example, an LCD is a transmissive SLM that can block or pass incoming light rays by turning pixels black or transparent. A straightforward solution places an LCD in front of an OST-HMD and drives the LCD to block light rays that could pass through rendered virtual objects at the viewpoint. This simple approach is, however, defective due to the depth mismatch between the LCD panel and the virtual objects, leading to blurred occlusion. This led existing OST-HMDs to employ dedicated hardware such as focus optics and multi-stacked SLMs. Contrary to these viable, yet complex and/or computationally expensive solutions, we return to the single-layer LCD approach for the hardware simplicity while maintaining fine occlusion-we compensate for a degraded occlusion area by overlaying a compensation image. We compute the image based on the HMD parameters and the background scene captured by a scene camera. The evaluation demonstrates that the proposed method reduced the occlusion leak error by 61.4% and the occlusion error by 85.7%.
Yuta Itoh 0001, Takumi Hamasaki, Maki Sugimoto
IEEE Trans. Vis. Comput. Graph.3
2016 Analysis of Multiple Users' Experience in Daily Life Using Wearable Device for Facial Expression Recognition
abstract
In this paper, we present a wearable facial expression recognition system that can analyse and enhance a daily experience. Our aim is to create a mindful experience in daily life by connecting the device with everyday objects and service. To this end, we made two prototypes that supports users to keep right side of emotions: 1) a text chatting system that automatically inserts an emoticon based on his/her facial expressions in the end of a comment a user typed, 2) a plant interface controlled by facial expressions. We also analysed multiple users' facial expressions while they played video games. We confirmed that visualization of sensor data from the device shows the possibility for estimating the transition of different facial expressions.
Katsutoshi Masai, Yuta Itoh 0001, Yuta Sugiura, Maki Sugimoto
ACE4
2016 Facial Expression Recognition in Daily Life by Embedded Photo Reflective Sensors on Smart Eyewear
abstract
This paper presents a novel smart eyewear that uses embedded photo reflective sensors and machine learning to recognize a wearer's facial expressions in daily life. We leverage the skin deformation when wearers change their facial expressions. With small photo reflective sensors, we measure the proximity between the skin surface on a face and the eyewear frame where 17 sensors are integrated. A Support Vector Machine (SVM) algorithm was applied for the sensor information. The sensors can cover various facial muscle movements and can be integrated into everyday glasses. The main contributions of our work are as follows. (1) The eyewear recognizes eight facial expressions (92.8% accuracy for one time use and 78.1% for use on 3 different days). (2) It is designed and implemented considering social acceptability. The device looks like normal eyewear, so users can wear it anytime, anywhere. (3) Initial field trials in daily life were undertaken. Our work is one of the first attempts to recognize and evaluate a variety of facial expressions in the form of an unobtrusive wearable device.
Katsutoshi Masai, Yuta Sugiura, Masa Ogata, Kai Kunze, Masahiko Inami, Maki Sugimoto
IUI6
2016 MARCut: Marker-based Laser Cutting for Personal Fabrication on Existing Objects
abstract
Typical personal fabrication using a laser cutter allows objects to be created from raw material and the engraving of existing objects. Current methods to precisely align an object with the laser is a difficult process due to indirect manipulations. In this paper, we propose a marker-based system as a novel paradigm for direct interactive laser cutting on existing objects. Our system, MARCut, performs the laser cutting based on tangible markers that are applied directly onto the object to express the design. Two types of markers are available; hand constructed Shape Markers that represent the desired geometry, and Command Markers that indicate the operational parameters such as cut, engrave or material.
Takashi Kikuchi, Yuichi Hiroi, Ross Smith 0001, Bruce H. Thomas, Maki Sugimoto
TEI5
2015 VolRec: haptic display of virtual inner volume in consideration of angular moment
abstract
In this paper, we propose a haptic device which displays the virtual inner volume inside using angular moment. The device is developed using a linear motion guide, a weight and a motor which can control the height of the center of the gravity. Also, psychophysical experiment was carried out to evaluate our device. By conducting a regression analysis on the result of the experiment, there was a significant effect on user perception of the inner volume. Based on the result, we created several applications with our proposed method to show the usage.
Ryota Koshiyama, Takashi Kikuchi, Jun Morita, Maki Sugimoto
Advances in Computer Entertainment4
2015 Remote Welding Robot Manipulation Using Multi-view Images
abstract
This paper proposes a remote welding robot manipulation system by using multi-view images. After an operator specifies two-dimensional path on images, the system transforms it into three-dimensional path and displays the movement of the robot by overlaying graphics with images. The accuracy of our system is sufficient to weld objects when combining with a sensor in the robot. The system allows the non-expert operator to weld objects remotely and intuitively, without the need to create a 3D model of a processed object beforehand.
Yuichi Hiroi, Kei Obata, Katsuhiro Suzuki, Naoto Ienaga, Maki Sugimoto, Hideo Saito 0001, Tadashi Takamaru
ISMAR5
2015 Registration and projection method of tumor region projection for breast cancer surgery
abstract
This paper introduces a registration and projection method for directly projecting the tumor region for breast cancer surgery assistance based on the breast procedure of our collaborating doctor. We investigated the steps of the breast cancer procedure of our collaborating doctor and how it can be applied for tumor region projection. We propose a novel way of MRI acquisition so we may correlate the MRI coordinates to the patient in the real world. By calculating the transformation matrix from the MRI coordinates and the coordinates from the markers that is on the patient, we are able to register the acquired MRI data to the patient. Our registration and presentation method of the tumor region was then evaluated by medical doctors.
Motoko Kanegae, Jun Morita, Sho Shimamura, Yuji Uema, Maiko Takahashi, Masahiko Inami, Tetsu Hayashida, Maki Sugimoto
VR8
2015 3D position measurement of planar photo detector using gradient patterns
abstract
We propose a three dimensional position measurement method employing planar photo detectors to calibrate a Spatial Augmented Reality system of unknown geometry. In Spatial Augmented Reality, projectors overlay images onto an object in the physical environment. For this purpose, the alignment of the images and physical objects is required. Traditional camera based 3D position tracking systems, such as multi-camera motion capture systems, detect the positions of optical markers in two-dimensional image plane of each camera device, so those systems require multiple camera devices at known locations to obtain 3D position of the markers. We introduce a detection method of 3D position of a planar photo detector by projecting gradient patterns. The main contribution of our method is to realize an alignment of the projected images with the physical objects and measuring the geometry of the objects simultaneously for Spatial Augmented Reality applications.
Tatsuya Kodera, Maki Sugimoto, Ross Smith 0001, Bruce H. Thomas
VR2
2015 MRI overlay system using optical see-through for marking assistance
abstract
In this paper we propose an augmented reality system that superimposes MRI onto the patient model. We use a half-silvered mirror and a handheld device to superimpose the MRI onto the patient model. By tracking the coordinates of the patient model and the handheld device using optical markers, we are able to transform the images to the correlated position. Voxel data of the MRI are made so that the user is able to view the MRI from many different angles.
Jun Morita, Sho Shimamura, Motoko Kanegae, Yuji Uema, Maiko Takahashi, Masahiko Inami, Tetsu Hayashida, Maki Sugimoto
VR8
2014 RGB-D-T camera system for AR display of temperature change
abstract
The anomalies of power equipment can be founded using temperature changes compared to its normal state. In this paper we present a system for visualizing temperature changes in a scene using a thermal 3D model. Our approach is based on two precomputed 3D models of the target scene achieved with a RGB-D camera coupled with the thermal camera. The first model contains the RGB information, while the second one contains the thermal information. For comparing the status of the temperature between the model and the current time, we accurately estimate the pose of the camera by finding keypoint correspondences between the current view and the RGB 3D model. Knowing the pose of the camera, we are then able to compare the thermal 3D model with the current status of the temperature from any viewpoint.
Kazuki Matsumoto, Wataru Nakagawa, François de Sorbier, Maki Sugimoto, Hideo Saito 0001, Shuji Senda, Takashi Shibata 0001, Akihiko Iketani
ISMAR4
2014 Move-it sticky notes providing active physical feedback through motion
abstract
Post-it notes are a popular paper format that serves a multitude of purposes in our daily lives, as they provide excellent affordances for quick capturing of informal notes, and location-sensitive reminding. In this paper, we present Move-it, a system that combines Post-it notes with a technologically enhanced paperclip to demonstrate how a passive piece of paper can be turned into an "active" medium that conveys information through motion. We present two application examples that investigate the applicability of Move-it sticky notes for ambient information awareness. In comparison to existing notification systems, experimental results show that they reduce negative effects of interruptions on emotional state and performance, and provide unique affordances by combining advantages of physical and digital systems into a novel active paper interface.
Kathrin Probst, Michael Haller, Kentaro Yasu, Maki Sugimoto, Masahiko Inami
TEI4
2014 An AR edutainment system supporting bone anatomy learning
abstract
We present a medical Augmented Reality (AR) edutainment system for bone anatomy learning. This learning environment, called AR bone puzzle, is a metaphor for bone anatomy learning with AR visualization and intuitive interaction. AR bone puzzle uses its user's body as a puzzle frame and computer generated virtual bones as puzzle pieces. Users learn bone anatomy by assembling the virtual bone pieces on their body. Key features of this system are 3D AR visualization and intuitive gesture based user interaction.
Philipp Stefan, Patrick Wucherer, Yuji Oyamada, Alexander Schoch, Motoko Kanegae, Naoki Shimizu, Tatsuya Kodera, Sebastien Cahier, Matthias Weigl, Maki Sugimoto, Pascal Fallavollita, Hideo Saito 0001, Nassir Navab
VR11
2013 PukaPuCam: Enhance Travel Logging Experience through Third-Person View Camera Attached to Balloons
Tsubasa Yamamoto, Yuta Sugiura, Suzanne Low, Koki Toda, Kouta Minamizawa, Maki Sugimoto, Masahiko Inami
Advances in Computer Entertainment6
2013 Virtual Slicer: Development of Interactive Visualizer for Tomographic Medical Images Based on Position and Orientation of Handheld Device
abstract
This paper proposes an interface that helps understanding the correspondence between the patient and medical images. In our proposed method, we have developed an interactive visualizer for tomographic images based on the relative position and orientation of the handheld device and the patient.
Sho Shimamura, Motoko Kanegae, Yuji Uema, Masahiko Inami, Tetsu Hayashida, Hideo Saito 0001, Maki Sugimoto
CW7
2013 Demo chairs
abstract
We are delighted to present the 12th edition of the demonstration program of the IEEE International Symposium on Mixed and Augmented Reality. The ISMAR demonstration program provide hands-on experience to the community of the most recent technical and applied advancement in Augmented Reality.
Ross Smith 0001, Maki Sugimoto, Raphaël Grasset
ISMAR2
2012 Illumination estimation from shadow and incomplete object shape captured by an RGB-D camera
Takuya Ikeda, Yuji Oyamada, Maki Sugimoto, Hideo Saito 0001
ICPR3
2012 Optical camouflage III: Auto-stereoscopic and multiple-view display system using retro-reflective projection technology
abstract
This paper presents a new type of optical camouflage system based on the retro-reflective projection technology. Retro-reflective projection is a method used to create augmented reality that combines the virtual world with the real world. The conventional model of an optical camouflage system consists of a retro-reflective screen, a projection source and a beam splitter. In such a setup, the user needs to observe an object covered with the retro-reflective screen through a single viewpoint. This is called a monocular system. In our new setup, our aim is to construct a system that has multiple viewpoints by applying a novel projection array system. We will describe the method with which this projection array system is achieved using one projection source, the configuration of the system, and the trade-offs of the system. In addition, we will describe an application of our system in a car. The installed system makes the backseat virtually transparent, allowing the driver to see the blind spots at the rear when reversing the car.
Yuji Uema, Naoya Koizumi, Shian Wei Chang, Kouta Minamizawa, Maki Sugimoto, Masahiko Inami
VR5
2011 Finding the Right Way for Interrupting People Improving Their Sitting Posture
Michael Haller, Christoph Richter, Peter Brandl, Sabine Gross, Gerold Schossleitner, Andreas Schrempf, Hideaki Nii, Maki Sugimoto, Masahiko Inami
INTERACT (2)8
2011 Detecting shape deformation of soft objects using directional photoreflectivity measurement
abstract
We present the FuwaFuwa sensor module, a round, hand-size, wireless device for measuring the shape deformations of soft objects such as cushions and plush toys. It can be embedded in typical soft objects in the household without complex installation procedures and without spoiling the softness of the object because it requires no physical connection. Six LEDs in the module emit IR light in six orthogonal directions, and six corresponding photosensors measure the reflected light energy. One can easily convert almost any soft object into a touch-input device that can detect both touch position and surface displacement by embedding multiple FuwaFuwa sensor modules in the object. A variety of example applications illustrate the utility of the FuwaFuwa sensor module. An evaluation of the proposed deformation measurement technique confirms its effectiveness.
Yuta Sugiura, Kakehi Gota, Anusha Withana, Calista Lee, Daisuke Sakamoto, Maki Sugimoto, Masahiko Inami, Takeo Igarashi
UIST6
2010 Myglobe: a navigation service based on cognitive maps
abstract
Myglobe is a user generated navigation service that enables users to share each cognitive map with one another. Cognitive map is a personalized map, shape of which is emphasized according to user's preference and activity in the city. It facilitates users to look back on their own city and have a new understanding by using an application in smart phones and physically interacting with a globe shaped device. In this paper, we present Myglobe service for users to achieve a new city experience with cognitive maps.
Takuo Imbe, Fumitaka Ozaki, Shin Kiyasu, Yusuke Mizukami, Shuichi Ishibashi, Masa Inakage, Naohito Okude, Adrian David Cheok, Masahiko Inami, Maki Sugimoto
TEI10
2007 Remote active tangible interactions
abstract
This paper presents a new form of remote active tangible interactions built with the Display-based Measurement and Control System. A prototype system was constructed to demonstrate the concepts of coupled remote tangible objects on rear projected tabletop displays. A user evaluation measuring social presence for two users performing a furniture placement task was performed, to determine a difference between this new system and a traditional mouse.
Jan Richter, Bruce H. Thomas, Maki Sugimoto, Masahiko Inami
TEI3
2005 Virtual Acceleration with Galvanic Vestibular Stimulation in Virtual Reality Environment
abstract
This study describes the relation between the vection produced by optical flow and that created by galvanic vestibular stimulation. Vection is the illusion of self motion and is most often experienced when an observer views a large screen display containing a translating pattern. This illusion has only limited fidelity and duration unless it is reinforced by confirming vestibular information. Galvanic vestibular stimulation (GVS) can directly produce the sensation of vection.
Taro Maeda, Hideyuki Ando, Maki Sugimoto
VR3