Takahiro Fukumori

dblp:59/8315 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-4317-9704ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2024 Environmental Sound Synthesis from Vocal Imitations and Sound Event Labels
abstract
One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as rhythm and pitch, using vocal imitations, which cannot be expressed by conventional input information, such as sound event labels, images, or texts, in an environmental sound synthesis model. In this paper, we propose a framework for environmental sound synthesis from vocal imitations and sound event labels based on a framework of a vector quantized encoder and the Tacotron2 decoder. Using vocal imitations is expected to control the pitch and rhythm of the synthesized sound, which only sound event labels cannot control. Our objective and subjective experimental results show that vocal imitations effectively control the pitch and rhythm of synthesized sounds.
Yuki Okamoto, Keisuke Imoto, Shinnosuke Takamichi, Ryotaro Nagase, Takahiro Fukumori, Yoichi Yamashita
ICASSP5
2024 RISC: A Corpus for Shout Type Classification and Shout Intensity Prediction
abstract
The detection of shouted speech is crucial in audio surveillance and monitoring. Although it is desirable for a security system to be able to identify emergencies, existing corpora provide only a binary label (i.e., shouted or normal) for each speech sample, making it difficult to predict the shout intensity. Furthermore, most corpora comprise only utterances typical of hazardous situations, meaning that classifiers cannot learn to discriminate such utterances from shouts typical of less hazardous situations such as cheers. Thus, this paper presents a novel research source, the RItsumeikan Shout Corpus (RISC), which contains wide variety types of shouted speech samples collected in recording experiments. Each shouted speech sample in RISC has a shout type and is also assigned shout intensity ratings via a crowdsourcing service. We also present a comprehensive performance comparison among deep learning approaches for speech type classification tasks and a shout intensity prediction task. The results show that feature learning based on the spectral and cepstral domains achieves high performance, no matter which network architecture is used. The results also demonstrate that shout type classification and intensity prediction are still challenging tasks, and RISC is expected to contribute to further development in this research area.
Takahiro Fukumori, Taito Ishida, Yoichi Yamashita
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Speech Emotion Recognition by Estimating Emotional Label Sequences with Phoneme Class Attribute
abstract
In recent years, much research has been into speech emotion recognition (SER) using deep learning to predict emotions conveyed by speech. We studied the method that detected the emotion for the whole utterance using the frame-based SER, which estimates emotions in each frame rather than in a whole utterance. One of the problems with this method is that the emotional label sequence, which is used in training the frame-based SER, does not sufficiently consider phonemic characteristics. To solve this problem, we propose new methods of recognizing the emotion for the whole utterance using frame-based SER that considers the phoneme class attribute such as vowels, voiced consonants, unvoiced consonants, and other symbols in training. As a result, we found that the proposed methods significantly improve the performance of the result for the whole utterance compared to conventional methods.
Ryotaro Nagase, Takahiro Fukumori, Yoichi Yamashita
INTERSPEECH2
2022 Sound Event Detection Guided by Semantic Contexts of Scenes
abstract
Some studies have revealed that contexts of scenes (e.g., "home," "office," and "cooking") are advantageous for sound event detection (SED). Mobile devices and sensing technologies give useful information on scenes for SED without the use of acoustic signals. However, conventional methods can employ pre-defined contexts in inference stages but not undefined contexts. This is because one-hot representations of pre-defined scenes are exploited as prior contexts for such conventional methods. To alleviate this problem, we propose scene-informed SED where pre-defined scene-agnostic contexts are available for more accurate SED. In the proposed method, pre-trained large-scale language models are utilized, which enables SED models to employ unseen semantic contexts of scenes in inference stages. Moreover, we investigated the extent to which the semantic representation of scene contexts is useful for SED. Experimental results performed with TUT Sound Events 2016/2017 and TUT Acoustic Scenes 2016/2017 datasets show that the proposed method improves micro and macro F-scores by 4.34 and 3.13 percentage points compared with conventional Conformer- and CNN– BiGRU-based SED, respectively.
Noriyuki Tonami, Keisuke Imoto, Ryotaro Nagase, Yuki Okamoto, Takahiro Fukumori, Yoichi Yamashita
ICASSP5
2021 Sound Event Detection Based on Curriculum Learning Considering Learning Difficulty of Events
abstract
In conventional sound event detection (SED) models, two types of events, namely, those that are present and those that do not occur in an acoustic scene, are regarded as the same type of the events. The conventional SED methods cannot effectively exploit the difference between the two types of events. The all time frames of sound events that do not occur in an acoustic scene are easily regarded as inactive in the scene, that is, the events are easy-to-train. The time frames of the events that are present in a scene must be classified as active in addition to inactive in the acoustic scene, that is, the events are difficult-to-train. To take advantage of the training difficulty, we apply curriculum learning into SED, where models are trained from easy- to difficult-to-train events. To utilize the curriculum learning, we propose a new objective function for SED, wherein the events are trained from easy-to difficult-to-train events. Experimental results show that the F-score of the proposed method is improved by 10.09 percentage points compared with that of the conventional binary cross entropy-based SED.
Noriyuki Tonami, Keisuke Imoto, Yuki Okamoto, Takahiro Fukumori, Yoichi Yamashita
ICASSP4
2021 Deep Spectral-Cepstral Fusion for Shouted and Normal Speech Classification
Takahiro Fukumori
Interspeech1
2018 Single-Channel Speech Enhancement With Phase Reconstruction Based on Phase Distortion Averaging
abstract
Speech enhancement has been widely investigated for several decades, but by modifying only the amplitude spectrum of a speech signal, ignoring the phase spectrum, which has been regarded as an unimportant feature. However, it was recently reported that the phase spectrum plays an important role in speech quality and intelligibility. In this paper, we propose a phase reconstruction method based on harmonic enhancement using the fundamental frequency and phase distortion feature. This feature is known to show fluctuations in the phase spectrum with respect to time and frequency. We estimate the speech phase spectrum by considering the relationship between harmonic phase spectra. Experimental evaluations indicate that the proposed phase reconstruction method improves speech quality in various noisy environments.
Yukoh Wakabayashi, Takahiro Fukumori, Masato Nakayama, Takanobu Nishiura, Yoichi Yamashita
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Phase reconstruction method based on time-frequency domain harmonic structure for speech enhancement
abstract
Speech enhancement in noisy environments has been widely investigated by modifying only the amplitude spectrum of the speech signal, while the phase spectrum, which is regarded as an unimportant feature, is ignored. However, it has recently been reported that the phase spectrum plays an important role in the intelligibility and quality of speech. We propose a speech-enhancement method with phase reconstruction, which estimates inartificial phase spectrum by using the time-frequency feature called phase distortion, though a conventional phase reconstruction estimates artificial one. The objective experimental results indicate improvement in speech quality with and efficiency of the proposed method.
Yukoh Wakabayashi, Takahiro Fukumori, Masato Nakayama, Takanobu Nishiura, Yoichi Yamashita
ICASSP2
2012 Digital Archive for Japanese Intangible Cultural Heritage Based on Reproduction of High-Fidelity Sound Field in Yamahoko Parade of Gion Festival
abstract
We digitally archived festival music signals (“Ohayashi”) in the Yamahoko parades of Gion festival in Kyoto, Japan. Besides, the festival music, which consists of Japanese traditional drums, flutes, bells, ambient noise and in-float driving noise are needed to reproduce the authentic atmosphere of this festival. To reproduce a high-quality sound field, we recorded the festival music in the presence of ambient noise and float noise by using multi-channel recording. We reproduced the sound field of one of the parades with the recorded sound sources. We employed point-source loudspeakers for reproducing Japanese traditional drums, flutes, and bells with omni-directional radiation characteristics. After that, we built a web-based system linked to a map of the parade route that would produce an acoustic sound field with realistic sensations at different points along the route. This system reproduces the Ohayashi at particular positions along the route when the user clicks circular buttons on the map.
Takahiro Fukumori, Takanobu Nishiura, Yoichi Yamashita
SNPD1
2010 Performance estimation of reverberant speech recognition based on reverberant criteria RSR-dn with acoustic parameters
Takahiro Fukumori, Masanori Morise, Takanobu Nishiura
INTERSPEECH1