Naoya Takahashi

dblp:19/8442 · DBLP profile ↗
← Back
30ranked-venue papers
13as first author
18since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 11 first-author · 14 since 2021Artificial intelligence and machine learning · 15 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 SilentCipher: Deep Audio Watermarking
Mayank Kumar Singh, Naoya Takahashi, Wei-Hsiang Liao 0001, Yuki Mitsufuji
INTERSPEECH2
2023 Diffroll: Diffusion-Based Generative Music Transcription with Unsupervised Pretraining Capability
abstract
In this paper we propose a novel generative approach, DiffRoll, to tackle automatic music transcription (AMT). Instead of treating AMT as a discriminative task in which the model is trained to convert spectrograms into piano rolls, we think of it as a conditional generative task where we train our model to generate realistic looking piano rolls from pure Gaussian noise conditioned on spectrograms. This new AMT formulation enables DiffRoll to transcribe, generate and even inpaint music. Due to the classifier-free nature, DiffRoll is also able to be trained on unpaired datasets where only piano rolls are available. Our experiments show that DiffRoll outperforms its discriminative counterpart by 19 percentage points (ppt.) and our ablation studies also indicate that it outperforms similar existing methods by 4.8 ppt.Source code and demonstration are available at https://sony.github.io/DiffRoll/.
Kin Wai Cheuk, Ryosuke Sawata, Toshimitsu Uesaka, Naoki Murata, Naoya Takahashi, Shusuke Takahashi, Dorien Herremans, Yuki Mitsufuji
ICASSP5
2023 Nonparallel Emotional Voice Conversion for Unseen Speaker-Emotion Pairs Using Dual Domain Adversarial Network & Virtual Domain Pairing
abstract
Primary goal of an emotional voice conversion (EVC) system is to convert the emotion of a given speech signal from one style to another style without modifying the linguistic content of the signal. Most of the state-of-the-art approaches convert emotions for seen speaker-emotion combinations only. In this paper, we tackle the problem of converting the emotion of speakers whose only neutral data are present during the time of training and testing (i.e., unseen speaker-emotion combinations). To this end, we extend a recently proposed StartGANv2-VC architecture by utilizing dual encoders for learning the speaker and emotion style embeddings separately along with dual domain source classifiers. For achieving the conversion to unseen speaker-emotion combinations, we propose a Virtual Domain Pairing (VDP) training strategy, which virtually incorporates the speaker-emotion pairs that are not present in the real data without compromising the min-max game of a discriminator and generator in adversarial training. We evaluate the proposed method using a Hindi emotional database.
Nirmesh J. Shah, Mayank Kumar Singh, Naoya Takahashi, Naoyuki Onoe
ICASSP3
2023 Hierarchical Diffusion Models for Singing Voice Neural Vocoder
abstract
Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness, and pronunciations. In this work, we propose a hierarchical diffusion model for singing voice neural vocoders. The proposed method consists of multiple diffusion models operating in different sampling rates; the model at the lowest sampling rate focuses on generating accurate low-frequency components such as pitch, and other models progressively generate the waveform at higher sampling rates on the basis of the data at the lower sampling rate and acoustic features. Experimental results show that the proposed method produces high-quality singing voices for multiple singers, outperforming state-of-the-art neural vocoders with a similar range of computational costs.
Naoya Takahashi, Mayank Kumar Singh, Yuki Mitsufuji
ICASSP1
2023 CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos
Hao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley, Taylor Berg-Kirkpatrick
ICLR2
2023 Multi-Person Tracking Method Robust to Dynamic Viewport Changes for AR apps
abstract
Augmented reality (AR) devices have gained a lot of attention in recent years due to their ability to enhance people’s abilities through 2D/3D spatial sensing and recognition functions. RGB cameras are most often used as sensors in this type of spatial recognition, and a particularly important task is the detection and tracking of objects and people in physical space. However, the camera positions and orientations on AR devices such as smartphones and smart glasses, frequently change due to the user wearing them on their head, leading to non-linear and complex motion in the video frames and reducing the accuracy of tracking people. To address this issue, the proposed method combines person re-identification based on deep metric learning with trajectory prediction to estimate the person’s sequential positions in 3D space around the camera. The experimental result shows 95.45% accuracy with our dataset.
Naoya Takahashi, Tatsuya Amano, Hirozumi Yamaguchi
IE1
2023 Iteratively Improving Speech Recognition and Voice Conversion
Mayank Kumar Singh, Naoya Takahashi, Naoyuki Onoe
INTERSPEECH2
2023 STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events
abstract
While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from the feet of a walker. This paper proposes an audio-visual sound event localization and detection (SELD) task, which uses multichannel audio and video information to estimate the temporal activation and DOA of target sound events. Audio-visual SELD systems can detect and localize sound events using signals from a microphone array and audio-visual correspondence. We also introduce an audio-visual dataset, Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23), which consists of multichannel audio data recorded with a microphone array, video data, and spatiotemporal annotation of sound events. Sound scenes in STARSS23 are recorded with instructions, which guide recording participants to ensure adequate activity and occurrences of sound events. STARSS23 also serves human-annotated temporal activation labels and human-confirmed DOA labels, which are based on tracking results of a motion capture system. Our benchmark results demonstrate the benefits of using visual object positions in audio-visual SELD tasks. The data is available at https://zenodo.org/record/7880637.
Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel Krause 0001, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji
NeurIPS9
2022 Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection
abstract
Recording and annotating real sound events for a sound event localization and detection (SELD) task is time consuming, and data augmentation techniques are often favored when the amount of data is limited. However, how to augment the spatial information in a dataset, including unlabeled directional interference events, remains an open research question. Furthermore, directional interference events make it difficult to accurately extract spatial characteristics from target sound events. To address this problem, we propose an impulse response simulation framework (IRS) that augments spatial characteristics using simulated room impulse responses (RIR). RIRs corresponding to a microphone array assumed to be placed in various rooms are accurately simulated, and the source signals of the target sound events are extracted from a mixture. The simulated RIRs are then convolved with the extracted source signals to obtain an augmented multi-channel training dataset. Evaluation results obtained using the TAU-NIGENS Spatial Sound Events 2021 dataset show that the IRS contributes to improving the overall SELD performance. Additionally, we conducted an ablation study to discuss the contribution and need for each component within the IRS.
Yuichiro Koyama, Kazuhide Shigemi, Masafumi Takahashi, Kazuki Shimada, Naoya Takahashi, Emiru Tsunoo, Shusuke Takahashi, Yuki Mitsufuji
ICASSP5
2022 Multi-ACCDOA: Localizing And Detecting Overlapping Sounds From The Same Class With Auxiliary Duplicating Permutation Invariant Training
abstract
Sound event localization and detection (SELD) involves identifying the direction-of-arrival (DOA) and the event class. The SELD methods with a class-wise output format make the model predict activities of all sound event classes and corresponding locations. The class-wise methods can output activity-coupled Cartesian DOA (ACCDOA) vectors, which enable us to solve a SELD task with a single target using a single network. However, there is still a challenge in detecting the same event class from multiple locations. To overcome this problem while maintaining the advantages of the class-wise format, we extended ACCDOA to a multi one and proposed auxiliary duplicating permutation invariant training (ADPIT). The multi-ACCDOA format (a class- and track-wise output format) enables the model to solve the cases with overlaps from the same class. The class-wise ADPIT scheme enables each track of the multi-ACCDOA format to learn with the same target as the single-ACCDOA format. In evaluations with the DCASE 2021 Task 3 dataset, the model trained with the multi-ACCDOA format and with the class-wise ADPIT detects overlapping events from the same class while maintaining its performance in the other cases. Also, the proposed method performed comparably to state-of-the-art SELD methods with fewer parameters.
Kazuki Shimada, Yuichiro Koyama, Shusuke Takahashi, Naoya Takahashi, Emiru Tsunoo, Yuki Mitsufuji
ICASSP4
2022 Amicable Examples for Informed Source Separation
abstract
This paper deals with the problem of informed source separation (ISS), where the sources are accessible during the so-called encoding stage. Previous works computed side-information during the encoding stage and source separation models were designed to utilize the side-information to improve the separation performance. In contrast, in this work, we improve the performance of a pre-trained separation model that does not use any side-information. To this end, we propose to adopt an adversarial attack for the opposite purpose, i.e., rather than computing the perturbation to degrade the separation, we compute an imperceptible perturbation called amicable noise to improve the separation. Experimental results show that the proposed approach selectively improves the performance of the targeted separation model by 2.23 dB on average and is robust to signal compression. Moreover, we propose multi-model multi-purpose learning that control the effect of the perturbation on different models individually.
Naoya Takahashi, Yuki Mitsufuji
ICASSP1
2022 Amicable Examples for Informed Source Separation
abstract
This paper deals with the problem of informed source separation (ISS), where the sources are accessible during the so-called encoding stage. Previous works computed side-information during the encoding stage and source separation models were designed to utilize the side-information to improve the separation performance. In contrast, in this work, we improve the performance of a pre-trained separation model that does not use any side-information. To this end, we propose to adopt an adversarial attack for the opposite purpose, i.e., rather than computing the perturbation to degrade the separation, we compute an imperceptible perturbation called amicable noise to improve the separation. Experimental results show that the proposed approach selectively improves the performance of the targeted separation model by 2.23 dB on average and is robust to signal compression. Moreover, we propose multi-model multi-purpose learning that control the effect of the perturbation on different models individually.
Naoya Takahashi, Yuki Mitsufuji
ICASSP1
2022 Leveraging Symmetrical Convolutional Transformer Networks for Speech to Singing Voice Style Transfer
abstract
In this paper, we propose a model to perform style transfer of speech to singing voice. Contrary to the previous signal processing-based methods, which require high-quality singing templates or phoneme synchronization, we explore a data-driven approach for the problem of converting natural speech to singing voice. We develop a novel neural network architecture, called SymNet, which models the alignment of the input speech with the target melody while preserving the speaker identity and naturalness. The proposed SymNet model is comprised of symmetrical stack of three types of layers - convolutional, transformer, and self-attention layers. The paper also explores novel data augmentation and generative loss annealing methods to facilitate the model training. Experiments are performed on the NUS and NHSS datasets which consist of parallel data of speech and singing voice. In these experiments, we show that the proposed SymNet model improves the objective reconstruction quality significantly over the previously published methods and baseline architectures. Further, a subjective listening test confirms the improved quality of the audio obtained using the proposed approach (absolute improvement of 0.37 in mean opinion score measure over the baseline system).
Shrutina Agarwal, Naoya Takahashi, Sriram Ganapathy
INTERSPEECH2
2021 Densely Connected Multi-Dilated Convolutional Networks for Dense Prediction Tasks
abstract
Tasks that involve high-resolution dense prediction require a modeling of both local and global patterns in a large input field. Although the local and global structures often depend on each other and their simultaneous modeling is important, many convolutional neural network (CNN)-based approaches interchange representations in different resolutions only a few times. In this paper, we claim the importance of a dense simultaneous modeling of multiresolution representation and propose a novel CNN architecture called densely connected multidilated DenseNet (D3Net). D3Net involves a novel multidilated convolution that has different dilation factors in a single layer to model different resolutions simultaneously. By combining the multidilated convolution with the DenseNet architecture, D3Net incorporates multiresolution learning with an exponentially growing receptive field in almost all layers, while avoiding the aliasing problem that occurs when we naively incorporate the dilated convolution in DenseNet. Experiments on the image semantic segmentation task using Cityscapes and the audio source separation task using MUSDB18 show that the proposed method has superior performance over stateof-the-art methods.
Naoya Takahashi, Yuki Mitsufuji
CVPR1
2021 End-to-End Lyrics Recognition with Voice to Singing Style Transfer
abstract
Automatic transcription of monophonic/polyphonic music is a challenging task due to the lack of availability of large amounts of transcribed data. In this paper, we propose a data augmentation method that converts natural speech to singing voice based on vocoder based speech synthesizer. This approach, called voice to singing (V2S), performs the voice style conversion by modulating the F0 contour of the natural speech with that of a singing voice. The V2S model based style transfer can generate good quality singing voice thereby enabling the conversion of large corpora of natural speech to singing voice that is useful in building an E2E lyrics transcription system. In our experiments on monophonic singing voice data, the V2S style transfer provides a significant gain (relative improvements of 21 %) for the E2E lyrics transcription system. We also discuss additional components like transfer learning and lyrics based language modeling to improve the performance of the lyrics transcription system.
Sakya Basak, Shrutina Agarwal, Sriram Ganapathy, Naoya Takahashi
ICASSP4
2021 Accdoa: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization And Detection
abstract
Neural-network (NN)-based methods show high performance in sound event localization and detection (SELD). Conventional NN-based methods use two branches for a sound event detection (SED) target and a direction-of-arrival (DOA) target. The two-branch representation with a single network has to decide how to balance the two objectives during optimization. Using two networks dedicated to each task increases system complexity and network size. To address these problems, we propose an activity-coupled Cartesian DOA (ACCDOA) representation, which assigns a sound event activity to the length of a corresponding Cartesian DOA vector. The ACCDOA representation enables us to solve a SELD task with a single target and has two advantages: avoiding the necessity of balancing the objectives and model size increase. In experimental evaluations with the DCASE 2020 Task 3 dataset, the ACCDOA representation outperformed the two-branch representation in SELD metrics with a smaller network size. The ACCDOA-based SELD system also performed better than state-of-the-art SELD systems in terms of localization and location-dependent detection.
Kazuki Shimada, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Yuki Mitsufuji
ICASSP3
2021 Adversarial Attacks on Audio Source Separation
abstract
Despite the excellent performance of neural-network-based audio source separation methods and their wide range of applications, their robustness against intentional attacks has been largely neglected. In this work, we reformulate various adversarial attack methods for the audio source separation problem and intensively investigate them under different attack conditions and target models. We further propose a simple yet effective regularization method to obtain imperceptible adversarial noise while maximizing the impact on separation quality with low computational complexity. Experimental results show that it is possible to largely degrade the separation quality by adding imperceptibly small noise when the noise is crafted for the target model. We also show the robustness of source separation models against a black-box attack. This study provides potentially useful insights for developing content protection methods against the abuse of separated signals and improving the separation performance and robustness.
Naoya Takahashi, Shota Inoue, Yuki Mitsufuji
ICASSP1
2021 Hierarchical disentangled representation learning for singing voice conversion
abstract
Conventional singing voice conversion (SVC) methods often suffer from operating in high-resolution audio owing to a high dimensionality of data. In this paper, we propose a hierarchical representation learning that enables the learning of disentangled representations with multiple resolutions independently. With the learned disentangled representations, the proposed method progressively performs SVC from low to high resolutions. Experimental results show that the proposed method outperforms baselines that operate with a single resolution in terms of mean opinion score (MOS), similarity score, and pitch accuracy.
Naoya Takahashi, Mayank Kumar Singh, Yuki Mitsufuji
IJCNN1
2020 Improving Voice Separation by Incorporating End-To-End Speech Recognition
abstract
Despite recent advances in voice separation methods, many challenges remain in realistic scenarios such as noisy recording and the limits of available data. In this work, we propose to explicitly incorporate the phonetic and linguistic nature of speech by taking a transfer learning approach using an end-to-end automatic speech recognition (E2EASR) system. The voice separation is conditioned on deep features extracted from E2EASR to cover the long-term dependence of phonetic aspects. Experimental results on speech separation and enhancement task on the AVSpeech dataset show that the proposed method significantly improves the signal-to-distortion ratio over the baseline model and even outperforms an audio visual model, that utilizes visual information of lip movements.
Naoya Takahashi, Mayank Kumar Singh, Sakya Basak, Sudarsanam Parthasaarathy, Sriram Ganapathy, Yuki Mitsufuji
ICASSP1
2019 A Teaching Assistant Robot Design Tool Based on Knowledge Chunks Reuse
abstract
To reduce the cost to develop service robot applications, we have been developing PRINTEPS, an AI and service robot application development platform for end users. PRINTEPS provides a scenario editor to describe the workflows for robots' actions and human-robot interactions. The scenario editor enables the developer to create workflows while reusing the entire workflows or individual operation processes; however, since end users such as domain experts usually do not know functions and properties of robots, it is difficult to identify and modify the reusable parts to suit the purposes of reuse when reusing the entire workflows. In contrast, it is also difficult to work out different combinations of operation processes when reusing the individual operation processes. To solve these issues, this study proposes a teaching assistant (TA) robot design tool based on knowledge chunks reuse. Two public elementary school teachers created workflows for TA robots using the proposed tool, and each teacher conducted a lesson with TA robots once. Through questionnaires given to the teachers and pupils, the proposed tool and TA robots were evaluated to confirm their usefulness.
Takeshi Morita 0001, Naoya Takahashi, Mizuki Kosuda, Takahira Yamaguchi
COMPSAC (2)2
2019 A Knowledge Chunk Reuse Support Tool based on Heterogeneous Ontologies
abstract
To develop service robot applications, it is necessary to acquire domain expert knowledge and develop the applications based on the knowledge. However, since, currently, many of these applications have been developed by engineers using the middleware for robots, the domain expert knowledge is embedded in the codes and is difficult to reuse. Therefore, it is considered necessary to have a tool that supports the development of the applications based on machine-readable knowledge of domain experts. We also believe that the machine-readable knowledge can be reused not only for service robots but also for novices in the domain. To address the problems, this paper proposes a knowledge chunk (KC) reuse support tool based on heterogeneous ontologies. In this study, the parts of the reusable workflow, indexes required for a search, and a movie recording of robots movement based on the parts of the workflow are collectively known as a KC. Using the framework of case-based reasoning, the proposed tool accumulates parts of reusable workflows as case examples based on heterogeneous ontologies and facilitates search and reuse of KCs. It promotes domain expert knowledge acquisition and supports novices to learn the knowledge. As a case study, we have applied the proposed tool to teaching assistant (TA) robots. Two public elementary school teachers created workflows for TA robots using the proposed tool, and each teacher conducted a lesson with TA robots once. Through questionnaires given to the teacher, the proposed tool and TA robot application were evaluated to confirm their usefulness.
Takeshi Morita 0001, Naoya Takahashi, Mizuki Kosuda, Takahira Yamaguchi
KEOD2
2019 Recursive Speech Separation for Unknown Number of Speakers
abstract
In this paper we propose a method of single-channel speakerindependent multi-speaker speech separation for an unknown number of speakers.As opposed to previous works, in which the number of speakers is assumed to be known in advance and speech separation models are specific for the number of speakers, our proposed method can be applied to cases with different numbers of speakers using a single model by recursively separating a speaker.To make the separation model recursively applicable, we propose one-and-rest permutation invariant training (OR-PIT).Evaluation on WSJ0-2mix and WSJ0-3mix datasets show that our proposed method achieves state-ofthe-art results for two-and three-speaker mixtures with a single model.Moreover, the same model can separate four-speaker mixture, which was never seen during the training.We further propose the detection of the number of speakers in a mixture during recursive separation and show that this approach can more accurately estimate the number of speakers than detection in advance by using a deep neural network based classifier.
Naoya Takahashi, Sudarsanam Parthasaarathy, Nabarun Goswami, Yuki Mitsufuji
INTERSPEECH1
2018 PhaseNet: Discretized Phase Modeling with Deep Neural Networks for Audio Source Separation
abstract
Previous research on audio source separation based on deep neural networks (DNNs) mainly focuses on estimating the magnitude spectrum of target sources and typically, phase of the mixture signal is combined with the estimated magnitude spectra in an ad-hoc way. Although recovering target phase is assumed to be important for the improvement of separation quality, it can be difficult to handle the periodic nature of the phase with the regression approach. Unwrapping phase is one way to eliminate the phase discontinuity, however, it increases the range of value along with the times of unwrapping, making it difficult for DNNs to model. To overcome this difficulty, we propose to treat the phase estimation problem as a classification problem by discretizing phase values and assigning class indices to them. Experimental results show that our classification based approach 1) successfully recovers the phase of the target source in the discretized domain, 2) improves signal-to distortion ratio (SDR) over the regression-based approach in both speech enhancement task and music source separation (MSS) task, and 3) outperforms state-of-the-art MSS.
Naoya Takahashi, Purvi Agrawal, Nabarun Goswami, Yuki Mitsufuji
INTERSPEECH1
2018 AENet: Learning Deep Audio Features for Video Analysis
abstract
We propose a new deep network for audio event recognition, called AENet. In contrast to speech, sounds coming from audio events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an extended time period due to the lack of clear subword units that are present in speech. In order to incorporate this long-time frequency structure of audio events, we introduce a convolutional neural network (CNN) operating on a large temporal input. In contrast to previous works, this allows us to train an audio event detection system end to end. The combination of our network architecture and a novel data augmentation outperforms previous methods for audio event detection by 16%. Furthermore, we perform transfer learning and show that our model learned generic audio features, similar to the way CNNs learn generic features on vision tasks. In video analysis, combining visual features and traditional audio features, such as mel frequency cepstral coefficients, typically only leads to marginal improvements. Instead, combining visual features with our AENet features, which can be computed efficiently on a GPU, leads to significant performance improvements on action recognition and video highlight detection. In video highlight detection, our audio features improve the performance by more than 8% over visual features alone.
Naoya Takahashi, Michael Gygli, Luc Van Gool
IEEE Trans. Multim.1
2017 Improving music source separation based on deep neural networks through data augmentation and network blending
abstract
This paper deals with the separation of music into individual instrument tracks which is known to be a challenging problem. We describe two different deep neural network architectures for this task, a feed-forward and a recurrent one, and show that each of them yields themselves state-of-the art results on the SiSEC DSD100 dataset. For the recurrent network, we use data augmentation during training and show that even simple separation networks are prone to overfitting if no data augmentation is used. Furthermore, we propose a blending of both neural network systems where we linearly combine their raw outputs and then perform a multi-channel Wiener filter post-processing. This blending scheme yields the best results that have been reported to-date on the SiSEC DSD100 dataset.
Stefan Uhlich, Marcello Porcu, Franck Giron, Michael Enenkl, Thomas Kemp, Naoya Takahashi, Yuki Mitsufuji
ICASSP6
2017 Implementation of Teacher-Robot Collaboration Lesson Application in PRINTEPS
abstract
PRINTEPS is currently being developed as a total intelligent application, which has sub systems for knowledge-based reasoning, speech dialogue, image sensing, motion planning, and machine learning, in order to support end users on easily developing intelligent applications for human-machine collaboration. In this paper, a lesson application for collaborative teaching among a robot, laptop PC, sensor, teachers and students was developed with PRINTEPS. The implementation lesson was performed in a science class for six grade elementary students.
Shunsuke Akashiba, Chihiro Nishimoto, Naoya Takahashi, Takeshi Morita 0001, Reiji Kukihara, Misae Kuwayama, Takahira Yamaguchi
KES3
2017 Development of applications for teaching assistant robots with teachers in PRINTEPS
abstract
PRINTEPS is currently being developed as a total intelligent application, which has sub systems for knowledge-based reasoning, speech dialogue, image sensing, motion planning, and machine learning, in order to support end users on easily developing intelligent applications for human-machine collaboration. In this paper, a lesson application for collaborative teaching among a robot, laptop PC, sensor, teachers and students was developed with PRINTEPS. The implementation lesson was performed in a science class for six grade elementary students.
Shunsuke Akashiba, Chihiro Nishimoto, Naoya Takahashi, Takeshi Morita 0001, Reiji Kukihara, Misae Kuwayama, Takahira Yamaguchi
WI3
2016 Deep Convolutional Neural Networks and Data Augmentation for Acoustic Event Recognition
Naoya Takahashi, Michael Gygli, Beat Pfister, Luc Van Gool
INTERSPEECH1
2016 Automatic Pronunciation Generation by Utilizing a Semi-Supervised Deep Neural Networks
abstract
Phonemic or phonetic sub-word units are the most commonly used atomic elements to represent speech signals in modern ASRs.However they are not the optimal choice due to several reasons such as: large amount of effort required to handcraft a pronunciation dictionary, pronunciation variations, human mistakes and under-resourced dialects and languages.Here, we propose a data-driven pronunciation estimation and acoustic modeling method which only takes the orthographic transcription to jointly estimate a set of sub-word units and a reliable dictionary.Experimental results show that the proposed method which is based on semi-supervised training of a deep neural network largely outperforms phoneme based continuous speech recognition on the TIMIT dataset.
Naoya Takahashi, Tofigh Naghibi, Beat Pfister
INTERSPEECH1
2010 Fluorescent pipettes for optically targeted patch-clamp recordings
Daisuke Ishikawa, Naoya Takahashi, Takuya Sasaki, Atsushi Usami, Norio Matsuki, Yuji Ikegaya
Neural Networks2