VLDB 2026 Research / reviewers in the wild / expert
Katsutoshi Masai
dblp:161/0128 · also Masai Katsutoshi
· DBLP profile ↗
19ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0001-9768-5314ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 15 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Electrooculography-Based Detection of Refractive Vision ProblemsabstractEarly detection of visual impairments remains a persistent challenge, especially due to the subtle and often unnoticed nature of early-stage symptoms. Recent works have attempted to transition clinical tests to home-based services or develop innovative diagnostic methods, but most approaches remain self-initiated and discrete. In this study, we focused on refractive disorders and explored the feasibility of using electrooculography (EOG) to detect changes in refractive power passively. Thirty-nine participants used optometry trial lenses to simulate different refractive conditions. Participants performed a series of visual tasks while their EOG signals were recorded. We trained classification models to predict simulated refractive power levels relative to baseline visual condition across multiple evaluation settings, including within-subject, temporal generalization, and across-subject scenarios. The findings reveal that refractive power classification models achieve a mean accuracy of $0.950 \pm 0.034$ in within-subject, within-condition scenarios. Within-subject models tested on data from a different time point showed highly variable performance. While some participants achieved promising results, overall accuracy remained low, with a mean of $0.159 \pm 0.285$. We employed three strategies to evaluate the across-subject models. Naive models performed poorly ($0.161 \pm 0.063$) and linear normalization provided limited improvement ($0.175 \pm 0.062$). However, the fine-tuning strategy substantially improved the model's performance ($0.785 \pm 0.123$). EOG signals contain useful information for refractive power classification, particularly in personalized contexts. However, generalizing across time and individuals remains challenging. Overall, this work offers valuable insights for advancing EOG-based systems aimed at passive, real-time monitoring of visual conditions. Xin Wei 0007, Huakun Liu, Yutaro Hirao, Monica Perusquía-Hernández, Katsutoshi Masai, Hideaki Uchiyama, Kiyoshi Kiyokawa |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | MaGEL: A Soft, Transparent Input Device Enabling Deformation Gesture Recognition
Fumika Oguri, Katsutoshi Masai, Yuta Sugiura, Yuichi Itoh |
IUI | 2 |
| 2025 | Mind Your Vision: A Passive Multimodal Framework for Refractive Disorders Measurement Combining Electrooculography and Eye TrackingabstractRefractive errors are among the most common visual impairments globally, yet their diagnosis often relies on active user participation and clinical oversight. This study explores a passive method for estimating refractive power using two eye movement recording techniques: electrooculography (EOG) and video-based eye tracking. Using a publicly available dataset recorded under varying diopter conditions, we trained Long Short-Term Memory (LSTM) models to classify refractive power from unimodal (EOG or video-based eye tracking) and multimodal configurations. In the context of eye movement analysis, EOG captures fine-grained electrical signals, while video-based tracking provides rich features such as pupil dynamics and gaze behavior, making the two modalities complementary. We assess performance in both subject-dependent and subject-independent settings to evaluate model personalization and generalizability across individuals. Results show that the multimodal model consistently outperforms unimodal models, achieving the highest average accuracy in both settings: 96.568% in the subject-dependent scenario and 9.344% in the subject-independent scenario. Statistical comparisons in the subject-dependent setting confirmed that both unimodal and multimodal models significantly exceeded the chance level. Among them, the multimodal model significantly outperformed the EOG and eye-tracking models. The strong performance of subject-dependent models highlights the potential for developing personalized models tailored to the target user for refractive power monitoring. However, generalization remains limited, with classification accuracy only marginally above chance in the subject-independent evaluations. Our findings demonstrate both the potential and current limitations of eye movement data-based refractive error estimation, contributing to the development of continuous, non-invasive screening methods using EOG signals and eye-tracking data. Xin Wei 0007, Huakun Liu, Yutaro Hirao, Monica Perusquía-Hernández, Katsutoshi Masai, Hideaki Uchiyama, Kiyoshi Kiyokawa |
MUM | 5 |
| 2024 | Facial Gesture Classification with Few-shot Learning Using Limited Calibration Data from Photo-reflective Sensors on Smart EyewearabstractThis study investigates smart eyewear for facial gesture classification with low user calibration costs.The smart eyewear is equipped with low-cost, comfortable, energy-efficient photo-reflective sensors, which can detect changes in facial muscle movements.Although the sensor output is useful for facial gesture classification, individual user calibration is considered necessary.Moreover, re-calibration is required whenever the user's wearing position changes.Therefore, reducing the calibration cost is crucial for the wider applicability of the eyewear.To address this issue, we propose a few-shot domain adaptation approach using Convolutional Neural Networks (CNN).We evaluate the accuracy of classifying eight gestures with data augmentation and a supervised contrastive loss.Data augmentation is employed to make the model more robust to noise, while the supervised contrastive loss is introduced to learn user-invariant features.Our approach with data augmentation achieves robust gesture classification, with an average accuracy of 93.46% (SD = 8.34%) for three emotion-related gestures without user-specific data.Furthermore, for user-independent training, we demonstrated that using few-shot learning with pre-trained models and only four repetitions of calibration data per gesture achieved a practical accuracy of 91.38% (SD = 6.03%), showing that a small amount of user-specific data is sufficient for the accurate classification.Also, it works under different wearing conditions, achieving an accuracy of 90.19% (SD = 3.56%).These results illustrate the potential of our method to improve the practicality of smart eyewear for facial expression recognition in cases of limited user data, making it more accessible and user-friendly. Katsutoshi Masai, Maki Sugimoto, Brian Kenji Iwana |
MUM | 1 |
| 2023 | Masktrap: Designing and Identifying Gestures to Transform Mask Strap into an Input InterfaceabstractEmbedding technology into day-to-day wearables and creating smart devices such as smartwatches and smart-glasses has been a growing area of interest. In this paper, we explore the interaction around face masks, a common accessory worn by many to prevent the spread of infectious diseases. Particularly, we propose a method of using the straps of a face mask as an input medium. We identified a set of plausible gestures on mask straps through an elicitation study (N = 20), in which the participants proposed different gestures for a given referent. We then developed a prototype to identify the gestures performed on the mask straps and present the recognition accuracy from a user study with eight participants. Our results show the system achieves 93.07% classification accuracy for 12 gestures. Takumi Yamamoto, Katsutoshi Masai, Anusha Withana, Yuta Sugiura |
IUI | 2 |
| 2023 | Analyzing the Effect of Diverse Gaze and Head Direction on Facial Expression Recognition With Photo-Reflective Sensors Embedded in a Head-Mounted DisplayabstractAs one of the facial expression recognition techniques for Head-Mounted Display (HMD) users, embedded photo-reflective sensors have been used. In this paper, we investigate how gaze and face directions affect facial expression recognition using the embedded photo-reflective sensors. First, we collected a dataset of five facial expressions (Neutral, Happy, Angry, Sad, Surprised) while looking in diverse directions by moving 1) the eyes and 2) the head. Using the dataset, we analyzed the effect of gaze and face directions by constructing facial expression classifiers in five ways and evaluating the classification accuracy of each classifier. The results revealed that the single classifier that learned the data for all gaze points achieved the highest classification performance. Then, we investigated which facial part was affected by the gaze and face direction. The results showed that the gaze directions affected the upper facial parts, while the face directions affected the lower facial parts. In addition, by removing the bias of facial expression reproducibility, we investigated the pure effect of gaze and face directions in three conditions. The results showed that, in terms of gaze direction, building classifiers for each direction significantly improved the classification accuracy. However, in terms of face directions, there were slight differences between the classifier conditions. Our experimental results implied that multiple classifiers corresponding to multiple gaze and face directions improved facial expression recognition accuracy, but collecting the data of the vertical movement of gaze and face is a practical solution to improving facial expression recognition accuracy. Fumihiko Nakamura, Masaaki Murakami, Katsuhiro Suzuki, Masaaki Fukuoka, Katsutoshi Masai, Maki Sugimoto |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Consistent Smile Intensity Estimation from Wearable Optical SensorsabstractSmiling plays a crucial role in human communication. It is the most frequent expression shown in daily life. Smile analysis usually employs computer vision-based methods that use data sets annotated by experts. However, cameras have space constraints in most realistic scenarios due to occlusions. Wearable electromyography is a promising alternative; however, issue of user comfort is a barrier to long-term use. Other wearable-based methods can detect smiles, but they lack consistency because they use subjective criteria without expert annotation. We investigate a wearable-based method that uses optical sensors for consistent smile intensity estimation while reducing manual annotation cost. First, we use a state-of-art computer vision method (OpenFace) to train a regression model to estimate smile intensity from sensor data. Then, we compare the estimation result to that of OpenFace. We also compared their results to human annotation. The results show that the wearable method has a higher matching coefficient (r=0.67) with human annotated smile intensity than OpenFace (r=0.56). Also, when the sensor data and OpenFace output were fused, the multimodal method produced estimates closer to human annotation (r=0.74). Finally, we investigate how the synchrony of smile dynamics among subjects and their average smile intensity are correlated to assess the potential of wearable smile intensity estimation. Katsutoshi Masai, Monica Perusquía-Hernández, Maki Sugimoto, Shiro Kumano, Toshitaka Kimura |
ACII | 1 |
| 2021 | Study of Interviewee's ImpressionMade by Interviewer Wearing Digital Full-face Mask DisplayDuring Recruitment InterviewabstractDuring recruitment interviews, the facial impressions of the interviewers likely affect the nervousness of the interviewees and make it difficult to conduct fair and consistent interview processes. To minimize the difference in facial impressions among interviewers, we have investigated a method to convert an interviewer’s face into an avatar. In this study, we used a digital full-face mask display capable of replacing the wearer’s face with an avatar in the real world. By reproducing the interviewer’s facial expression with an avatar in real time, we investigated the effect of avatar appearance on interviewee’s nervousness during interviews. Two types of avatars (a dignified face and a gentle face) were applied to three male interviewers in different age groups (10s, 20s and 60s). We compared the level of interviewee’s nervousness between before and after augmenting interviewer’s face with avatar. 172 college students were recruited as interviewees to assess the variation in the level of nervousness. Our experimental results show that avatar appearance can elicit more unique and consistent impressions than the interviewer’s real face and reduce the variation in interviewee’s nervousness level across interviewers. Kureha Noguchi, Yoshinari Takegawa, Yutaka Tokuda, Yuta Sugiura, Katsutoshi Masai, Keiji Hirata 0001 |
HAI | 5 |
| 2021 | Online Study Reveals the Multimodal Effects of Discrete Auditory Cues in Moving Target Estimation TaskabstractTasks that require temporal accuracy and precision are common in interactive systems. One such task is the moving target selection task in which visual motion cues are used to estimate timing. Our online study investigates the multimodal effects of discrete auditory cues in such a task in which the input timing has to be estimated from visual targets thrown at different speeds and intervals (moving target estimation). The auditory cues provide timing anticipation of the interception from the interval between the beats of two beeps. This intervention is novel in that it enriches the visual motion information and independently gives timing anticipation. Furthermore, we show that such cues can decrease systematic errors in the moving target estimation task. As the effectiveness of auditory cues is demonstrated in online experiments, we can expect to apply such intervention to interactive systems in various environments. Katsutoshi Masai, Akemi Kobayashi, Toshitaka Kimura |
ICMI | 1 |
| 2020 | Face Commands - User-Defined Facial Gestures for Smart GlassesabstractWe propose the use of face-related gestures involving the movement of the face, eyes, and head for augmented reality (AR). This technique allows us to use computer systems via hands-free, discreet interactions. In this paper, we present an elicitation study to explore the proper use of facial gestures for daily tasks in the context of a smart home. We used Amazon Mechanical Turk to conduct this study (N=37). Based on the proposed gestures, we report usage scenarios and complexity, proposed associations between gestures/tasks, a user-defined gesture set, and insights from the participants. We also conducted a technical feasibility study (N=13) with participants using smart eyewear to consider their uses in daily life. The device has 16 optical sensors and an inertial measurement unit (IMU). We can potentially integrate the system into optical see-through displays or other smart glasses. The results demonstrate that the device can detect eight temporal face-related gestures with a mean F1 score of 0.911 using a convolutional neural network (CNN). We also report the results of user-independent training and a one-hour recording of the experimenter testing two of the gestures. Katsutoshi Masai, Kai Kunze, Daisuke Sakamoto, Yuta Sugiura, Maki Sugimoto |
ISMAR | 1 |
| 2020 | Digital Full-Face Mask Display with Expression Recognition using Embedded Photo Reflective Sensor ArraysabstractThis paper presents a thin digital full-face mask display that can reflect an entire facial expression of a user onto an avatar to support augmented face-to-face communication in real environments. Although camera-based facial expression recognition technology has enabled people to augment their faces with avatars, application was limited to face-to-face communication in virtual environments. To enable digital facial augmentation with an avatar in a real space, we propose a digital face mask display system that integrates a lightweight flexible display with a thin facial expression recognition system. The thin wearable facial expression recognition system was implemented with photo reflective sensor arrays which can measure facial expressions at 40 feature points distributed across an entire face. We investigated a ten-class facial expression identification model based on an SVM training algorithm. The trained model achieved an average accuracy of 79% when identifying the facial expressions of multiple users. User experiments indicated that the proposed thin digital full-face mask display allows the wearer to control the facial expression of the avatar with a fast response rate and create a positive sense of self-agency and self-ownership toward the augmented avatar face. Yoshinari Takegawa, Yutaka Tokuda, Akino Umezawa, Katsuhiro Suzuki, Katsutoshi Masai, Yuta Sugiura, Maki Sugimoto, Diego Martínez 0001, Sriram Subramanian, Keiji Hirata 0001 |
ISMAR | 5 |
| 2020 | Classification of Spontaneous and Posed Smiles by Photo-reflective Sensors Embedded with Smart EyewearabstractSmile is one of the representative emotional expressions which is observed frequently in daily life and essential for various non-verbal communications. People make spontaneous smiles and intentional ones. It is important to guess properly whether a person is making a smile spontaneously or intentionally to understand the meaning of smiles. In this study, we propose a smile classification system with smart eyewear that equips photo-reflective sensors and examines whether we can distinguish two types of smiles; spontaneous smiles caused by funny videos and posed smiles evoked by instructions. We extract geometric features: reflection intensity distribution of sensors and temporal features in a time axis. By applying for Support Vector Machine, we observed 94.6% as the mean accuracy among 12 participants when we used both geometric and temporal features with user-dependent training. The result suggested that we can distinguish between spontaneous and posed smile by the sensors embedded with the smart eyewear. Chisa Saito, Katsutoshi Masai, Maki Sugimoto |
TEI | 2 |
| 2017 | EarTouch: turning the ear into an input surfaceabstractIn this paper, we propose EarTouch, a new sensing technology for ear-based input for controlling applications by slightly pulling the ear and detecting the deformation by an enhanced earphone device. It is envisioned that EarTouch will enable control of applications such as music players, navigation systems, and calendars as an "eyes-free" interface. As for the operation of EarTouch, the shape deformation of the ear is measured by optical sensors. Deformation of the skin caused by touching the ear with the fingers is recognized by attaching optical sensors to the earphone and measuring the distance from the earphone to the skin inside the ear. EarTouch supports recognition of multiple gestures by applying a support vector machine (SVM). EarTouch was validated through a set of user studies. Takashi Kikuchi, Yuta Sugiura, Katsutoshi Masai, Maki Sugimoto, Bruce H. Thomas |
MobileHCI | 3 |
| 2017 | Recognition and mapping of facial expressions to avatar by embedded photo reflective sensors in head mounted displayabstractWe propose a facial expression mapping technology between virtual avatars and Head-Mounted Display (HMD) users. HMD allow people to enjoy an immersive Virtual Reality (VR) experience. A virtual avatar can be a representative of the user in the virtual environment. However, the synchronization of the the virtual avatar's expressions with those of the HMD user is limited. The major problem of wearing an HMD is that a large portion of the user's face is occluded, making facial recognition difficult in an HMD-based virtual environment. To overcome this problem, we propose a facial expression mapping technology using retro-reflective photoelectric sensors. The sensors attached inside the HMD measures the distance between the sensors and the user's face. The distance values of five basic facial expressions (Neutral, Happy, Angry, Surprised, and Sad) are used for training the neural network to estimate the facial expression of a user. We achieved an overall accuracy of 88% in recognizing the facial expressions. Our system can also reproduce facial expression change in real-time through an existing avatar using regression. Consequently, our system enables estimation and reconstruction of facial expressions that correspond to the user's emotional changes. Katsuhiro Suzuki, Fumihiko Nakamura, Jiu Otsuka, Katsutoshi Masai, Yuta Itoh 0001, Yuta Sugiura, Maki Sugimoto |
VR | 4 |
| 2017 | CheekInput: turning your cheek into an input surface by embedded optical sensors on a head-mounted displayabstractIn this paper, we propose a novel technology called "CheekInput" with a head-mounted display (HMD) that senses touch gestures by detecting skin deformation. We attached multiple photo-reflective sensors onto the bottom front frame of the HMD. Since these sensors measure the distance between the frame and cheeks, our system is able to detect the deformation of a cheek when the skin surface is touched by fingers. Our system uses a Support Vector Machine to determine the gestures: pushing face up and down, left and right. We combined these 4 directional gestures for each cheek to extend 16 possible gestures. To evaluate the accuracy of the gesture detection, we conducted a user study. The results revealed that CheekInput achieved 80.45 % recognition accuracy when gestures were made by touching both cheeks with both hands, and 74.58 % when by touching both cheeks with one hand. Koki Yamashita, Takashi Kikuchi, Katsutoshi Masai, Maki Sugimoto, Bruce H. Thomas, Yuta Sugiura |
VRST | 3 |
| 2017 | Evaluation of Facial Expression Recognition by a Smart Eyewear for Facial Direction Changes, Repeatability, and Positional DriftabstractThis article presents a novel smart eyewear that recognizes the wearer’s facial expressions in daily scenarios. Our device uses embedded photo-reflective sensors and machine learning to recognize the wearer’s facial expressions. Our approach focuses on skin deformations around the eyes that occur when the wearer changes his or her facial expressions. With small photo-reflective sensors, we measure the distances between the skin surface on the face and the 17 sensors embedded in the eyewear frame. A Support Vector Machine (SVM) algorithm is then applied to the information collected by the sensors. The sensors can cover various facial muscle movements. In addition, they are small and light enough to be integrated into daily-use glasses. Our evaluation of the device shows the robustness to the noises from the wearer’s facial direction changes and the slight changes in the glasses’ position, as well as the reliability of the device’s recognition capacity. The main contributions of our work are as follows: (1) We evaluated the recognition accuracy in daily scenes, showing 92.8% accuracy regardless of facial direction and removal/remount. Our device can recognize facial expressions with 78.1% accuracy for repeatability and 87.7% accuracy in case of its positional drift. (2) We designed and implemented the device by taking usability and social acceptability into account. The device looks like a conventional eyewear so that users can wear it anytime, anywhere. (3) Initial field trials in a daily life setting were undertaken to test the usability of the device. Our work is one of the first attempts to recognize and evaluate a variety of facial expressions with an unobtrusive wearable device. Katsutoshi Masai, Kai Kunze, Yuta Sugiura, Masa Ogata, Masahiko Inami, Maki Sugimoto |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2016 | Analysis of Multiple Users' Experience in Daily Life Using Wearable Device for Facial Expression RecognitionabstractIn this paper, we present a wearable facial expression recognition system that can analyse and enhance a daily experience. Our aim is to create a mindful experience in daily life by connecting the device with everyday objects and service. To this end, we made two prototypes that supports users to keep right side of emotions: 1) a text chatting system that automatically inserts an emoticon based on his/her facial expressions in the end of a comment a user typed, 2) a plant interface controlled by facial expressions. We also analysed multiple users' facial expressions while they played video games. We confirmed that visualization of sensor data from the device shows the possibility for estimating the transition of different facial expressions. Katsutoshi Masai, Yuta Itoh 0001, Yuta Sugiura, Maki Sugimoto |
ACE | 1 |
| 2016 | Facial Expression Recognition in Daily Life by Embedded Photo Reflective Sensors on Smart EyewearabstractThis paper presents a novel smart eyewear that uses embedded photo reflective sensors and machine learning to recognize a wearer's facial expressions in daily life. We leverage the skin deformation when wearers change their facial expressions. With small photo reflective sensors, we measure the proximity between the skin surface on a face and the eyewear frame where 17 sensors are integrated. A Support Vector Machine (SVM) algorithm was applied for the sensor information. The sensors can cover various facial muscle movements and can be integrated into everyday glasses. The main contributions of our work are as follows. (1) The eyewear recognizes eight facial expressions (92.8% accuracy for one time use and 78.1% for use on 3 different days). (2) It is designed and implemented considering social acceptability. The device looks like normal eyewear, so users can wear it anytime, anywhere. (3) Initial field trials in daily life were undertaken. Our work is one of the first attempts to recognize and evaluate a variety of facial expressions in the form of an unobtrusive wearable device. Katsutoshi Masai, Yuta Sugiura, Masa Ogata, Kai Kunze, Masahiko Inami, Maki Sugimoto |
IUI | 1 |
| 2015 | Quantifying reading habits: counting how many words you readabstractReading is a very common learning activity, a lot of people perform it everyday even while standing in the subway or waiting in the doctors office. However, we know little about our everyday reading habits, quantifying them enables us to get more insights about better language skills, more effective learning and ultimately critical thinking. This paper presents a first contribution towards establishing a reading log, tracking how much reading you are doing at what time. We present an approach capable of estimating the words read by a user, evaluate it in an user independent approach over 3 experiments with 24 users over 5 different devices (e-ink reader, smartphone, tablet, paper, computer screen). We achieve an error rate as low as 5% (using a medical electrooculography system) or 15% (based on eye movements captured by optical eye tracking) over a total of 30 hours of recording. Our method works for both an optical eye tracking and an Electrooculography system. We provide first indications that the method works also on soon commercially available smart glasses. Kai Kunze, Katsutoshi Masai, Masahiko Inami, Ömer Sacakli, Marcus Liwicki, Andreas Dengel 0001, Shoya Ishimaru, Koichi Kise |
UbiComp | 2 |