Toshiharu Horiuchi

dblp:13/10147 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0002-1501-8532ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Three-Dimensional Sound Wave Propagation Reproduction by CE-FDTD Simulation Applying Actual Radiation Characteristics
abstract
The compact explicit finite difference time domain (CE-FDTD) method has been proposed as a non-steady-state analysis method and used for sound field simulation. For sound field simulation, it is necessary to apply the radiation characteristics to reproduce actual sound wave propagation. Radiation characteristics can be applied by setting sound pressure to dense grid arrangement in the CE-FDTD method. However, conventional techniques use a sparse array of microphones and are considered insufficient. Furthermore, the technique of applying captured acoustic signals for dense grid arrangement in the CE-FDTD method has not been considered. In this paper, we propose a hardware and software system that captures the radiation characteristics for a dense grid arrangement and applies them in the CE-FDTD method while controlling the sound wave propagation with non-propagation region. As a result of the proposed method, the average differences of sound pressure, propagation time, nd center frequency from the measured values are 1.8 dB, 0.04 ms, and 200 Hz, respectively, which is more accurate than the conventional techniques. It is shown that this system is useful for improving the accuracy of sound wave propagation reproduction with actual radiation characteristics.
Shota Okubo, Toshiharu Horiuchi
ICASSP2
2023 Directional Sound Source Representation Using Paired Microphone Array with Different Characteristics Suitable for Volumetric Video Capture
abstract
In this research, we propose a directional sound source representation technique for 3D contents such as volumetric video in metaverse and digital twin. Our proposed technique enables us to have a novel 3D audio-visual experience which is derived from immersive audio presentation expressing the radiation characteristics of sound source. To realize such an experience, we configure the spaced placement of paired microphone array and capture sound source signals completely without obstacles for volumetric video capture. Then, we synthesize the directional sound source signal using our technique which conducts signal processing to capture sound signals based on the positional and directional information of an object relative to a user. We developed and demonstrated a VR application using this technique to evaluate the change of sound with the object or user's movement in accordance with visual rendering. In our user study, we received lots of positive feedback for a novel audio-visual experience.
Shota Okubo, Tomoaki Konno, Toshiharu Horiuchi, Tatsuya Kobayashi
MMAsia3
2022 Sync Sofa: Sofa-type Side-by-side Communication Experience Based on Multimodal Expression
abstract
Lifestyle changes and digitalization have reduced opportunities for face-to-face, intimate communication, which is an indispensable activity for human beings. We have realized a method that allows intimate communication between two persons even though they are faraway from each other by integrating multimodal technologies. Sync Sofa is a new sofa-type communication tool. It senses the partner with a camera, microphones, and accelerometers. The sensed data are cross-modally integrated and then presented to the user through a life-size display, multichannel loudspeakers, and multichannel vibrotactile actuators. We have received positive feedback from many users who have experienced Sync Sofa, such as, "I felt like the other person was really sitting next to me".
Yuki Tajima, Shota Okubo, Tomoaki Konno, Toshiharu Horiuchi, Tatsuya Kobayashi
ACM Multimedia4
2021 Sync Glass: Virtual Pouring and Toasting Experience with Multimodal Presentation
abstract
One of the challenges of non-face-to-face communication is the absence of the haptic dimension. To solve this, a haptic communication system via the Internet has been proposed. The system has to be designed in such a way that it does not create discomfort during general use. The "Sync Glass" that we have developed transmits and presents the feeling of pouring a drink and making a toast accompanied by haptic, sound and visual effects. The device is designed to resemble a glass cup and, moreover, each action, including drinking and making a toast is performed in the customary way, making its use more acceptable to users. In the internal user demonstrations we performed, the experience has been reviewed with participants saying that "the feeling of pouring is so realistic", "so enjoyable!", and similar affirmative statements.
Yuki Tajima, Toshiharu Horiuchi, Gen Hattori
ACM Multimedia2
2021 Facial Action Unit-based Deep Learning Framework for Spotting Macro- and Micro-expressions in Long Video Sequences
abstract
In this paper, we utilize facial action units (AUs) detection to construct an end-to-end deep learning framework for the macro- and micro-expressions spotting task in long video sequences. The proposed framework focuses on individual components of facial muscle movement rather than processing the whole image, which eliminates the influence of image change caused by noises, such as body or head movement. Compared with existing models deploying deep learning methods with classical Convolutional Neural Network (CNN) models, the proposed framework utilizes Gated Recurrent Unit (GRU) or Long Short-term Memory (LSTM) or our proposed Concat-CNN models to learn the characteristic correlation between AUs of distinctive frames. The Concat-CNN uses three convolutional kernels with different sizes to observe features of different duration and emphasizes both local and global mutation features by changing dimensionality (max-pooling size) of the output space. Our proposal achieves state-of-the-art performance from the aspect of overall F1-scores: 0.2019 on CAS(ME)2-cropped, 0.2736 on SAMM Long Video, and 0.2118 on CAS(ME)2, which not only outperforms the baseline but is also ranked the 3rd of FME challenge 2021 for combined datasets of CAS(ME)2-cropped and SAMM-LV.
Zhiguang Zhou, Megumi Komiya, Koki Kishimoto, Keisuke Nonaka, Toshiharu Horiuchi, Satoshi Komorita, Gen Hattori, Sei Naito, Yasuhiro Takishima
ACM Multimedia8
2019 OtonoVR: Arbitrarily Angled Audio-visual VR Experience Using Selective Synthesis Sound Field Technique
abstract
We present an arbitrarily angled audio-visual VR experience app called OtonoVR for 360-degree panoramic videos using our selective synthesis sound field technique. This technique can synthesize two-channel stereo sound with scaled stereo width having an arbitrary angle range from 0 to 360 degrees centering on an arbitrary direction from multi-channel surround sound based on spectral modification. In the app, users can enjoy arbitrarily angled videos that they choose themselves by manipulating the touchscreen, and the stereo sound changes in terms of its spatial synchronization depending on the view. The app has been released for iOS and has been officially endorsed by Japanese idol groups.
Toshiharu Horiuchi, Sumaru Niida, Yasuhiro Takishima
ACM Multimedia1
2016 Typing Tutor: Individualized Tutoring in Text Entry for Older Adults Based on Input Stumble Detection
abstract
Many older adults are interested in smartphones. However most of them encounter difficulties in self-instruction and need support. Text entry, which is essential for various applications, is one of the most difficult operations to master. In this paper, we propose Typing Tutor, an individualized tutoring system for text entry that detects input stumbles and provides instructions. By conducting two user studies, we clarify the common difficulties that novice older adults experience and how skill level is related to input stumbles. Based on these studies, we develop Typing Tutor to support learning how to enter text on a smartphone. A two-week evaluation experiment with novice older adults (65+) showed that Typing Tutor was effective in improving their text entry proficiency, especially in the initial stage of use.
Toshiyuki Hagiya, Toshiharu Horiuchi, Tomonori Yazaki
CHI2
2014 Close/distant talker discrimination based on kurtosis of linear prediction residual signals
abstract
Desired/undesired speech discrimination is as important as speech/non-speech discrimination to achieve useful applications such as speech interfaces and teleconferencing systems. Conventional methods of voice activity detection (VAD) utilize the directional information of sound sources to distinguish desired from undesired speech. However, these methods have to utilize multiple microphones to estimate the directions of sound sources. Here, we propose a new method to discriminate desired from undesired speech with a single microphone. We assumed that the desired talkers would be close to the microphone, and the proposed method could distinguish close/distant-talking speech from observed signals based on the kurtosis of the linear prediction (LP) residual signals. The experimental results revealed that the proposed method could distinguish close-talking speech from distant-talking speech within a 10% equal error rate (EER) in ordinary reverberant environments with less processing time.
Kohei Hayashida, Masato Nakayama, Takanobu Nishiura, Yoichi Yamashita, Toshiharu Horiuchi, Tsuneo Kato
ICASSP5
2012 Acoustic-based passive pointing system for distant screens
abstract
This paper presents a passive pointing system for a distant screen based on an acoustic position estimation technology in conjunction with a gravity sensor on a smartphone. The system is designed to interact with a distant large screen such as a television set at home or digital signage in public. The system consists of a screen, two loudspeakers set around the screen, and a smartphone as a pointing device having a microphone and a gravity sensor inside. The position of the pointer is theoretically determined by the position and direction. This smartphone-based system approximates the position and direction by the two-dimensional position of the microphone horizontally and the pitch angle from the gravity sensor vertically. In this paper, we report experiments to evaluate the performance of the system. The loudspeakers of the system radiate burst signal from 18 to 24 kHz. The position of the smartphone is estimated at a frame rate of 15 Hz with a latency of 0.4 second. The accuracy of the pointer was measured as an angle error below 10 degrees for 100% of all frames. We confirmed that it has enough accuracy to point to each region which is divided area in the screen for applications such as quiz or questionnaire on digital signage.
Toshiharu Horiuchi, Shinya Takayama, Tsuneo Kato
ICASSP1
2012 Interactive music video application for smartphones based on free-viewpoint video and audio rendering
abstract
This paper presents a novel interactive music video application for smartphones based on free-viewpoint video technology in conjunction with three-dimensional positional audio technology. A user can enjoy a music video from a moving viewpoint that the user can manipulate by the touch screen, with the positional audio through the headphone. A user can even manipulate the positions of the performers on the stage as well as the viewpoint. The application, consisting of our audio rendering engine for multiple AAC ADTS files and our video rendering engine for multiple H.264 ES files, runs on a smartphone in stand-alone mode. The application has been released as official content from a music label for Android and iOS.
Toshiharu Horiuchi, Hiroshi Sankoh, Tsuneo Kato, Sei Naito
ACM Multimedia1
2001 Out-of-head sound localization using adaptive inverse filter
abstract
It is well known that the transfer functions for out-of-head sound localization differ with the listener. Our goal is to find a method that easily realizes excellent sound localization without measuring the individual transfer functions. This paper proposes an out-of-head sound localization system that uses adaptive inverse filtering without any measuring preprocesses. This system estimates the inverse ear canal transfer function (ECTF), and can adaptively obtain the transfer function to fit the listener in real-time. This paper also proposes an adaptive inverse filtering method for out-of-head sound localization that differs from the filtered-x method. In addition, the relationship between convergence time and initial value is assessed. It is clarified that the proposed method is more effective, in terms of convergence, if the initial value is the average of many listeners' impulse responses.
Toshiharu Horiuchi, Haruhide Hokari, Shoji Shimada
ICASSP1