VLDB 2026 Research / reviewers in the wild / expert
Keiichi Zempo
dblp:190/3105
· DBLP profile ↗
22ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0003-2339-5298ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Whip strike detection using high-sampling-rate audio by evaluating convolutional recurrent neural network configurations and class imbalance strategiesabstractWhip usage in horse racing is regulated in many jurisdictions, with violations typically verified through manual video review—a labor-intensive process unsuitable for real-time enforcement. To automate whip strike detection, this study proposes a convolutional recurrent neural network (CRNN)-based sound event detection (SED) framework designed specifically for impulsive, high-frequency sound events using high-sampling-rate audio data from horse races in Japan. The dataset comprised audio recordings from 24 official horse races with 620 annotated whip strike events. To address the significant class imbalance inherent in the dataset, sample reduction techniques were employed on the majority class. Architectural choices among CRNN models, which combine a convolutional neural network (CNN) for feature extraction and a recurrent neural network (RNN) for temporal analysis, were systematically evaluated, demonstrating that the Convolutional Neural Network 14-Bidirectional Gated Recurrent Unit (CNN14-BiGRU) model achieved the highest F1-score of 69.8%. The results confirm a key insight of this study: capturing impulsive high-frequency sounds is effectively achieved by leveraging high-sampling-rate audio, which provides both accurate high-frequency capture and improved temporal resolution, and by designing CNN architectures to extract richer acoustic features rather than solely enhancing temporal resolution. Furthermore, offline evaluation revealed that the CNN14-BiGRU model achieved detection latencies substantially shorter than real-time audio duration, indicating potential for practical real-time implementation. However, environmental noise and limited dataset size remain significant challenges affecting detection robustness and generalization. This work establishes foundational insights for automatic whip strike detection, guiding future model design and real-time implementations. • CRNN-based system proposed for detecting whip strike sounds in horse racing. • Model performance compared across CNN and RNN architecture combinations. • High-sampling-rate audio improves detection via frequency range and time resolution. Aoi Taguchi, Yuki Fujita, Keiichi Zempo |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Development and Evaluation of an Auditory VR Generative System via Natural Language Interaction to Aid Exposure Therapy for PTSD PatientsabstractPost-Traumatic Stress Disorder (PTSD) is a prevalent disorder triggered by life-threatening trauma, and exposure therapy, which involves confronting traumatic stimuli, has been proven to be highly effective for treating PTSD. However, exposure therapy has not been widely adopted. Virtual Reality (VR) exposure therapy, which has shown comparable effectiveness to that of traditional methods, is therefore advancing. However, this therapy has not been broadly implemented, partly because of the time required to create VR experiences tailored to a patient’s specific trauma. To address this problem, this study proposes a system for exposure therapy that generates auditory VR using a Large Language Model (LLM) for natural language interaction. This system, built on LLM and an audio dataset, generates sounds matching user-provided themes and generates corresponding scenarios and coordinates. An experiment with clinicians using this system to generate auditory stimuli was conducted to assess the usability and therapeutic potential of the generated audio. The results indicated high usability and quality, requiring minimal adjustments for therapeutic applications. Notably, the clinicians generated sounds within the duration of a standard clinical session. However, challenges remain, particularly for complex themes, highlighting the need for further research to enhance usability and verify the system’s clinical feasibility and efficacy. Yuta Yamauchi, Keiko Ino, Masanori Sakaguchi, Keiichi Zempo |
ACM Trans. Comput. Heal. | 4 |
| 2025 | Human-Like Remembering and Forgetting in LLM Agents: An ACT-R-Inspired Memory ArchitectureabstractThis study explores the implementation of human-like memory behavior in language agents by integrating the ACT-R cognitive architecture with large language models (LLMs). We designed a dialogue agent capable of dynamically retrieving and forgetting memories based on context, time, and usage frequency. The proposed system utilizes a vector-based activation mechanism, incorporating temporal decay, semantic similarity, and probabilistic noise to mimic natural memory dynamics. Simulation experiments confirmed the model’s ability to reproduce memory reinforcement through repeated topics, as well as stochastic variability in memory retrieval, reflecting human memory behavior. We also identified optimal parameters for balancing contextual sensitivity and memory stability. This work contributes to the development of more human-compatible AI dialogue systems by modeling memory not as mere storage, but as a strategic, context-sensitive process. Yudai Honda, Yuki Fujita, Keiichi Zempo, Shogo Fukushima |
HAI | 3 |
| 2025 | Enhancing Pedestrian Situation Awareness Through Auditory Augmented Reality: Effects of Frequency Shift on Vehicle Looming Perception
Yuichi Mashiba, Keitaro Tokunaga, Naoto Wakatsuki, Hiroaki Yano, Keiichi Zempo |
INTERACT (1) | 5 |
| 2025 | Background Sound Tempo Modulation Can Influence Scene-Specific Memory in Virtual RealityabstractSustaining user memory in digital environments such as virtual reality (VR) is a significant challenge. We show that temporary tempo modulations in background music (BGM) can selectively and naturally enhance memory in VR. In a user study (N = 20), decreasing the BGM tempo by approximately 21% significantly improved recall of events. These findings point to a new acoustic design approach that adapts scene by scene to narrative pacing and importance while maintaining a natural user experience. Hyuma Auchi, Shogo Fukushima, Yuki Fujita, Keiichi Zempo |
VRST | 4 |
| 2025 | Inducing the Presence of Rear Objects through Changes in Auditory Stimuli Using a Virtual Sound-Absorbing WallabstractWith the growing presence of personal mobility vehicles and autonomous mobile robots in public spaces, pedestrian collisions have become an increasing concern, particularly because these devices operate quietly and often approach unnoticed. Traditional notification methods typically emphasize urgency through explicit and intense notifications, which often lead to psychological burden. To address this issue, this study proposes a method aimed at reducing sensory stimuli by using a virtual sound-absorbing wall. This approach leverages the human "obstacle sense", which is the ability to perceive objects based on changes in ambient sound blocked by their presence. A prototype system was developed in a virtual reality (VR) environment using a head-mounted display (HMD), where a virtual sound-absorbing wall appears in the direction of objects approaching from behind to enhance their perceived presence. The proposed method was compared to conventional notification methods in a VR environment, and its psychological impact was evaluated. Additionally, the impact of varying the angular width of the virtual sound-absorbing wall on detection rates was investigated. Results indicated that although the detection rate was lower than that of the conventional method, the proposed approach significantly reduced psychological burden. With the widest absorption angle condition, accuracy exceeding 90% was achieved in detecting the presence of the virtual sound-absorbing wall from directions other than directly behind. Keitaro Tokunaga, Yuichi Mashiba, Naoto Wakatsuki, Keiichi Zempo, Hiroaki Yano |
VRST | 4 |
| 2023 | Study on a Visible Light Communication and Positioning System Utilizing an Optical Diffusion Filter and Rolling Shutter SensorabstractThe demand for indoor positioning technology is currently increasing in various industries. Although positioning methods using sensors and functions installed on smartphones have been proposed in previous studies, no definitive indoor positioning technology has yet been established. Therefore, this paper presents a positioning system that performs visible light communication using a rolling shutter sensor installed in smartphones and an optical diffusion filter that diffuses light in one direction. The system calculates the camera position from information obtained through communication and the LED position in the captured image. The performance of the positioning method using four LEDs was evaluated in both simulations and experiments. The obtained results suggest that the proposed positioning system achieved a positioning error of less than 100 mm. The proposed method can become a viable alternative to provide seamless transition from outdoor to indoor positioning using current technologies such as global navigation satellite systems. Reo Okawara, Tadashi Ebihara, Naoto Wakatsuki, Keiichi Zempo, Koichi Mizutani |
IPIN | 4 |
| 2023 | Effects of symmetrical avatar arm movements on the sense of ownership of both hands inverted in a mirrorabstractIn this study, we investigated whether visual and tactile symmetrical stimuli in the arm affect the mirrored avatar’s sense of ownership of the hand. In the experiment, we tested the user, avatar, and mirrored avatar’s sense of ownership of their hand by catching a ball multiple times in a virtual space. The results suggest that non-ambidextrous individuals were significantly more likely to recognize the mirror-image avatar’s hand as their own in scenarios that included symmetrical visual and tactile stimuli than in scenarios without such stimuli. This experiment may contribute to a better understanding of the relationship between left-right sensation and body ownership sensation. It may contribute to the pursuit of immersive experiences in virtual environments. Toko Fujita, Moeki Horii, Luis David Torres Mailleux, Masatatsu Miyagi, Yukihiko Okada, Keiichi Zempo |
VRST | 6 |
| 2023 | Effect of voice imitation using voice conversion by avatar on customer service in Virtual EnvironmentsabstractWe investigate the impact of voice imitation on rapport building in customer service in a virtual environment(VE). We simulated a VE store and conducted a within-subject experiment in three customer service scenarios with 16 participants. The imitation condition used a voice generated by a machine learning model to reproduce the participant’s linguistic content, with the voice identity set midway between the participant and the salesperson. In a group of men, voice imitation significantly improved their impression of salespeople. The findings of this study can be used to design interpersonal services in VE with the help of salesperson avatar voice design. Hiiro Okano, Yukihiko Okada, Naoto Wakatsuki, Keiichi Zempo |
VRST | 4 |
| 2023 | Whispering salesperson: perceptual illusion of interpersonal distance and ventriloquism effect in service of virtual environment by use of whisper voiceabstractThe opportunities to receive and provide services on virtual environment (VE) continue to increase. In this study, we focused on the voice in VE, and investigated the effects of whisper voice, which is used in intimate relationships, on interpersonal distance between salesperson and customer and the effective range of ventriloquism effect. The experimental results showed two results when the salesperson used whisper voice. The interpersonal distance between them was significantly smaller than that of the normal voice. The effective range of the ventriloquism effect was significantly larger on the front side and significantly smaller on the backside compared to normal voice. Mizuki Yabutani, Azusa Yamazaki, Naoto Wakatsuki, Yukihiko Okada, Keiichi Zempo |
VRST | 5 |
| 2022 | Inspection of unexpected defective products by semi-supervised learning based on a probability density function in high-yield food productionabstractIn this research, we propose a method for evaluating images to be inspected using only good images and assuming defective images are unavailable in anticipation of quality inspection applications in food processing facilities and other factories. We propose a discriminator based on a CNN that can evaluate the degree of deviation from good images by treating only good images as training data and assuming a fitness probability distribution. The results show that the proposed discriminator can detect defective products even without prior training data on defective products. The detection accuracy depends on the inspected object and the threshold that defines the deviation, which is comparable to previous studies that require defective images. With adjustable detection thresholds and automatic categorization of defective products, the proposed method is expected to be flexible enough to incorporate the knowledge of shop-floor workers on the production line. Masahiro Nakahara, Yuichi Mashiba, Ryusuke Miyamoto, Yuki Fujita, Hisashi Ishida, Keiichi Zempo |
IEEE Big Data | 6 |
| 2022 | Avatar Voice Morphing to Match Subjective and Objective Self Voice PerceptionabstractWe investigated the effect of morphing the avatar’s voice from the user’s voice on its impressions. We also investigated whether the image of morphing differed between those who liked and disliked their voice. The experiment was conducted by morphing the acoustic parameters such as fundamental frequency, spectral envelope, and aperiodic component based on the acoustic signals recorded by the participants themselves, and investigating their impressions of an avatar speaking with that voice. The result showed that those who liked their voice were most impressed by their original voice, while those who disliked it were more impressed by the morphed voice. This suggests that people who dislike their voice tend to seek their ideal in the avatar’s voice. Hiiro Okano, Keisuke Mizuno, Haruna Miyakawa, Keiichi Zempo |
VRST | 4 |
| 2021 | How Much is the Noise Level be Reduced? - Speech Recognition Threshold in Noise Environments Using a Parametric Speaker -
Noko Kuratomo, Tadashi Ebihara, Naoto Wakatsuki, Koichi Mizutani, Keiichi Zempo |
INTERACT (2) | 5 |
| 2021 | Effects of Audio-Visual Information on Interpersonal Distance in Real and Virtual Environments
Azusa Yamazaki, Naoto Wakatsuki, Koichi Mizutani, Yukihiko Okada, Keiichi Zempo |
INTERACT (5) | 5 |
| 2021 | ALiSE: Non-wearable AR display through the looking glass, and what looks solid thereabstractWith the Augmented Reality mirror display method using a half-mirror, there is a difference in the focal length between the mirror image and the AR image. Therefore, the observer perceives a mismatch in depth perception, which impairs usability. In this study, we developed an optical-reflection AR display, ALiSE (Augment Layer interweaved Semi-reflecting Existence), which enhances the depth perception experience of AR images by adding a gap zone with the same depth as the target depth between the display and the half-mirror. We conducted an experiment to view 3D objects and achieve virtual fitting using the existing AR with video synthesis and the proposed ALiSE method. As a result of the questionnaire survey, although the comfort of wearing virtual objects was below existing methods, we confirmed that the presence and solidity were superior with the proposed method to other approaches. This is an attempt to create a sense of the stereoscopic effect despite the 2D projection, as the object to be projected is simultaneously reflected in the mirror along with the observer themselves. Hiroki Uchida, Takayuki Kawamura, Keito Kamimura, Keiichi Zempo |
VRST | 4 |
| 2021 | Visual Transition of Avatars Improving Speech Comprehension in Noisy VR EnvironmentsabstractIn order to construct a comfortable communication in the VR space, it is important to improve the speech comprehension in environmental noise. Although there have been many reports on the interaction between vision and acoustic, few studies using noisy VR spaces. In this study, sixteen Japanese male and female were tested to listen to a some sentence in a VR space with environmental noise, to evaluate the effect of the visual stimulus to the avatar speech comprehension against the environmental noise, with using the up-and-down method. The results showed that the cocktail party effect was also observed in the VR avatars, and the cocktail party effect continued even if the avatar vanished visually. In addition, it was suggested that the cocktail party effect was enhanced if the lip of the avatar synchronized correctly. Takayoshi Yamada, Azusa Yamazaki, Haruna Miyakawa, Yuichi Mashiba, Keiichi Zempo |
VRST | 5 |
| 2020 | Real-time reassurance monitoring shopping basket in retail store: poster abstractabstractIn this study, we propose a sensor network to estimate service reassurance in a store by estimating customer stress level based on a pulse wave. We developed a basket that can calculate an RR interval (RRI) with sufficient accuracy to estimate customer reassurance from the measured pulse wave under in-store conditions. Two processes were applied to the measurement values based on the assumption that certain circumstances may interfere with measuring pulse waves. Furthermore, the authors conducted pulse wave measurement experiments under simulated store conditions. The results showed RRIs' values with an error rate of 4.9% compared to the true value. Azusa Yamazaki, Ryo Akatsu, Yukihiko Okada, Keiichi Zempo |
SenSys | 4 |
| 2019 | Shopping Baskets for On-line Beacon Sensor Network in Retail StoreabstractAlthough a significant volume of research about sensing networks are conducted to measure environmental distribution changes mainly in outdoor used GPS positioning techniques, there are few indoor positioning systems collecting information other than position. In this research, we developed the sensor network platform for customer behavior analysis in retail stores. These performances are confirmed through some experiments, and the increase or decrease of the goods, which are put into the baskets, is detected through the sensor network of IR signals. Yuji Suda, Taiga Arai, Takahiro Yoshizawa, Yuki Fujita, Keiichi Zempo, Yukihiko Okada |
CCNC | 5 |
| 2019 | Integrating a Binaural Beat into the Soundscape for the Alleviation of Feelings
Noko Kuratomo, Yuichi Mashiba, Keiichi Zempo, Koichi Mizutani, Naoto Wakatsuki |
INTERACT (4) | 3 |
| 2018 | Ultrasound from Cyber Map: Intuitive Guidance Method using Binaural Parametric LoudspeakerabstractThe people with visual impairment are facing the problem of the difficulty of reaching their desired destination. However, to have people who are with visual impairment updated with the necessary information about the location that they are traversing in order to secure safety and independent mobility, they need to depend on non-visual information to get the information they need.This paper proposes a system that automatically presents the sound image directly from arrival point without depending on the installation position of the sound source.This system works via presenting audible cues in the direction of the destination of the user without depending on the loudspeaker location. The success probability of direction guiding to the fixed destination for the proposed method was calculated and the results has revealed success probability of about 90%. As a result, this system presents the direction of the destination intuitively and easily by presenting the sound image of the guidance sound only for the person who needs information. Ryunosuke Iwaoka, Masataka Moriga, Hisham Elser Bilal Salih, Keiichi Zempo, Koichi Mizutani, Naoto Wakatsuki |
ISS | 4 |
| 2017 | The Alteration of Gustatory Sense by Virtual Chromatic Transition of Food Items
Yuto Sugita, Keiichi Zempo, Koichi Mizutani, Naoto Wakatsuki |
ICEC | 2 |
| 2016 | Effect of parameters of phase-modulated M-sequence signal on direction-of-arrival and localization errorabstractIn this paper, signals for acoustic beacons with M sequence are evaluated with various parameters for self-localization of mobile robot. The effect of (i) baud rate, (ii) carrier frequency and (iii) order of M sequence on DOA estimation error and localization error are evaluated by computer simulation. As a result, localization is possible with same class or more precision compared to the result from sound beacon with three chirp signals within 5 kHz-18 kHz by choosing the appropriate parameters, which are (i)baud rate is more than around 8 kHz and (ii)carrier frequency is 10 kHz. Particularly, (i) baud rate and (ii) carrier frequency have influence on the localization error. However, the localization error do not change into even a further baud rate in this condition if baud rate is higher than carrier frequency. In this simulation, no influence is confirmed from (iii)order of M sequence. The result indicates that M sequence can be used as signal of acoustic beacons instead of chirp, which occupies frequency band. Satoki Ogiso, Takuji Kawagishi, Koichi Mizutani, Naoto Wakatsuki, Keiichi Zempo |
IPIN | 5 |