EDBT 2026 Demo / reviewers in the wild / expert
Taichi Asami
dblp:50/5246
· DBLP profile ↗
41ranked-venue papers
10as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 9 first-author · 7 since 2021Artificial intelligence and machine learning · 31 · 8 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Microphone array geometry-independent multi-talker distant ASR: NTT system for DASR task of the CHiME-8 challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato 0002, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki |
Comput. Speech Lang. | 17 |
| 2024 | What Do Self-Supervised Speech and Speaker Models Learn? New Findings from a Cross Model Layer-Wise AnalysisabstractSelf-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representations. In contrast, speaker SSL models, exemplified by DINO-based models, adopt utterance-level training objectives primarily for speaker representation. Understanding how these models represent information is essential for refining model efficiency and effectiveness. Unlike the various analyses of speech SSL, there has been limited investigation into what information speaker SSL captures and how its representation differs from speech SSL or other fully-supervised speaker models. This paper addresses these fundamental questions. We explore the capacity to capture various speech properties by applying SUPERB evaluation probing tasks to speech and speaker SSL models. We also examine which layers are predominantly utilized for each task to identify differences in how speech is represented. Furthermore, we conduct direct comparisons to measure the similarities between layers within and across models. Our analysis unveils that 1) the capacity to represent content information is somewhat unrelated to enhanced speaker representation, 2) specific layers of speech SSL models would be partly specialized in capturing linguistic information, and 3) speaker SSL models tend to disregard linguistic information but exhibit more sophisticated speaker representation. Takanori Ashihara, Marc Delcroix, Takafumi Moriya, Kohei Matsuura, Taichi Asami, Yusuke Ijima |
ICASSP | 5 |
| 2024 | Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
Takafumi Moriya, Takanori Ashihara, Masato Mimura, Hiroshi Sato 0002, Kohei Matsuura, Ryo Masumura, Taichi Asami |
INTERSPEECH | 7 |
| 2023 | An Improved Approximation Algorithm for Wage Determination and Online Task Allocation in Crowd-SourcingabstractCrowd-sourcing has attracted much attention due to its growing importance to society, and numerous studies have been conducted on task allocation and wage determination. Recent works have focused on optimizing task allocation and workers' wages, simultaneously. However, existing methods do not provide good solutions for real-world crowd-sourcing platforms due to the low approximation ratio or myopic problem settings. We tackle an optimization problem for wage determination and online task allocation in crowd-sourcing and propose a fast 1-1/(k+3)^(1/2)-approximation algorithm, where k is the minimum of tasks' budgets (numbers of possible assignments). This approximation ratio is greater than or equal to the existing method. The proposed method reduces the tackled problem to a non-convex multi-period continuous optimization problem by approximating the objective function. Then, the method transforms the reduced problem into a minimum convex cost flow problem, which is a well-known combinatorial optimization problem, and solves it by the capacity scaling algorithm. Synthetic experiments and simulation experiments using real crowd-sourcing data show that the proposed method solves the problem faster and outputs higher objective values than existing methods. Yuya Hikima, Yasunori Akagi, Hideaki Kim, Taichi Asami |
AAAI | 4 |
| 2023 | SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
Takanori Ashihara, Takafumi Moriya, Kohei Matsuura, Tomohiro Tanaka, Yusuke Ijima, Taichi Asami, Marc Delcroix, Yukinori Honma |
INTERSPEECH | 6 |
| 2023 | What are differences? Comparing DNN and Human by Their Performance and Characteristics in Speaker Age Estimation
Yuki Kitagishi, Naohiro Tawara, Atsunori Ogawa, Ryo Masumura, Taichi Asami |
INTERSPEECH | 5 |
| 2023 | Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data
Takafumi Moriya, Hiroshi Sato 0002, Tsubasa Ochiai, Marc Delcroix, Takanori Ashihara, Kohei Matsuura, Tomohiro Tanaka, Ryo Masumura, Atsunori Ogawa, Taichi Asami |
INTERSPEECH | 10 |
| 2022 | Fast Bayesian Estimation of Point Process Intensity as Function of CovariatesabstractIn this paper, we tackle the Bayesian estimation of point process intensity as a function of covariates. We propose a novel augmentation of permanental process called augmented permanental process, a doubly-stochastic point process that uses a Gaussian process on covariate space to describe the Bayesian a priori uncertainty present in the square root of intensity, and derive a fast Bayesian estimation algorithm that scales linearly with data size without relying on either domain discretization or Markov Chain Monte Carlo computation. The proposed algorithm is based on a non-trivial finding that the representer theorem, one of the most desirable mathematical property for machine learning problems, holds for the augmented permanental process, which provides us with many significant computational advantages. We evaluate our algorithm on synthetic and real-world data, and show that it outperforms state-of-the-art methods in terms of predictive accuracy while being substantially faster than a conventional Bayesian method. Hideaki Kim, Taichi Asami, Hiroyuki Toda |
NeurIPS | 2 |
| 2021 | Streaming End-to-End Speech Recognition for Hybrid RNN-T/Attention Architecture
Takafumi Moriya, Tomohiro Tanaka, Takanori Ashihara, Tsubasa Ochiai, Hiroshi Sato 0002, Atsushi Ando, Ryo Masumura, Marc Delcroix, Taichi Asami |
Interspeech | 9 |
| 2019 | Recurrent out-of-vocabulary word detection based on distribution of features
Taichi Asami, Ryo Masumura, Yushi Aono, Koichi Shinoda |
Comput. Speech Lang. | 1 |
| 2018 | Neural Confnet Classification: Fully Neural Network Based Spoken Utterance Classification Using Word Confusion NetworksabstractThis paper describes neural ConfNet classification, a novel fully neural network based spoken utterance classification method that uses word confusion networks (ConfNets). Our motivation is to establish a spoken utterance classification method that can precisely understand natural language and robustly handle automatic speech recognition (ASR) errors. Remarkable progress has been made in neural networks for accurate modeling, however, most previous methods could not handle ASR errors since they were developed for reference transcriptions. Therefore, in our work we utilized ConfNets, which are compact and efficient graph representations of ASR hypotheses. Our idea is to regard the ConfNet as a sequence of bag-of-weighted-arcs and introduce a mechanism that converts the bag-of-weighted-arcs into a continuous representation called a modified weighted sum representation. This enables us to flexibly connect ConfNets to arbitrary model structures developed for reference transcriptions. We demonstrate the effectiveness of the neural ConfNet classification in dialogue act, extended named entity, and question type classification tasks. Ryo Masumura, Yusuke Ijima, Taichi Asami, Hirokazu Masataki, Ryuichiro Higashinaka |
ICASSP | 3 |
| 2017 | Domain adaptation of DNN acoustic models using knowledge distillationabstractConstructing deep neural network (DNN) acoustic models from limited training data is an important issue for the development of automatic speech recognition (ASR) applications that will be used in various application-specific acoustic environments. To this end, domain adaptation techniques that train a domain-matched model without overfitting by lever-aging pre-constructed source models are widely used. In this paper, we propose a novel domain adaptation method for DNN acoustic models based on the knowledge distillation framework. Knowledge distillation transfers the knowledge of a teacher model to a student model and offers better generalizability of the student model by controlling the shape of posterior probability distribution of the teacher model, which was originally proposed for model compression. We apply this framework to model adaptation. Our domain adaptation method avoids overfitting of the adapted model trained on limited data by transferring the knowledge of the source model to the adapted model by distillation. Experiments show that the proposed method can effectively avoid the overfitting of convolutional neural network based acoustic models and yield lower error rates than conventional adaptation methods. Taichi Asami, Ryo Masumura, Yoshikazu Yamaguchi, Hirokazu Masataki, Yushi Aono |
ICASSP | 1 |
| 2017 | Cross-modal transfer with neural word vectors for image feature learningabstractNeural word vector (NWV) such as word2vec is a powerful text representation tool that can encode extensive semantic information into compact vectors. This ability poses an interesting question in relation to image processing research - Can we learn better semantic image features from NWVs? We empirically explore this question in the context of semantic content-based image retrieval (CBIR). In this paper, we consider cross-modal transfer learning (CMT) to improve initial convolutional neural network (CNN) image features by using NWVs. We first show that NWVs can improve semantic CBIR performance compared to classical word vectors, even if it is with simple CMT models, i.e., canonical correlation analysis (CCA). Next, inspired by a characteristic property of NWVs, we propose a new CMT model and demonstrate that it can improve CBIR performance even further. Go Irie, Taichi Asami, Shuhei Tarashima, Takayuki Kurozumi, Tetsuya Kinebuchi |
ICASSP | 2 |
| 2017 | Parallel phonetically aware DNNs and LSTM-RNNS for frame-by-frame discriminative modeling of spoken language identificationabstractParallel phonetically aware deep neural networks (PPA-DNNs) and long short-term memory recurrent neural networks (PPA-LSTM-RNNs) to enhance frame-by-frame discriminative modeling of spoken language identification are proposed. This idea is inspired by traditional systems based on parallel phoneme recognition followed by language modeling (PPRLM). The proposed methods utilize multiple senone bottleneck features individually extracted from language-dependent senone-based DNNs in a frame-by-frame manner. The multiple senone bottleneck features can yield phonetic awareness to frame-by-frame DNNs and LSTM-RNNs without losing compatibility to real time applications. In experiments, three senone-based DNNs are introduced in order to extract senone bottleneck features, and both single use and parallel use of them are examined. Furthermore, we also examine a combination of PPA-DNNs and PPA-LSTM-RNNs. The proposed method's effectiveness is investigated by comparison with a simple speech aware modeling and traditional systems based on PPRLM. Ryo Masumura, Taichi Asami, Hirokazu Masataki, Yushi Aono |
ICASSP | 2 |
| 2017 | Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic modelsabstractAdapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount of speech data such as a single utterance. However, the auxiliary features are usually computed in batch mode, which causes some inevitable latency. In this paper we explore an extension of the auxiliary feature-based adaptation to online processing. We employ auxiliary features obtained from bottleneck speaker vectors and extend their computation to online processing using cumulative moving averaging. We test our proposed approach for deep CNN-based acoustic models, using context adaptive networks to exploit the auxiliary features. Experimental results on the CHiME-3 task demonstrate that the proposed approach can realize online speaker adaptation. Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Taichi Asami, Shigeru Katagiri, Tomohiro Nakatani |
ICASSP | 5 |
| 2017 | Prosody Aware Word-Level Encoder Based on BLSTM-RNNs for DNN-Based Speech Synthesis
Yusuke Ijima, Nobukatsu Hojo, Ryo Masumura, Taichi Asami |
INTERSPEECH | 4 |
| 2017 | Online End-of-Turn Detection from Speech Based on Stacked Time-Asynchronous Sequential Networks
Ryo Masumura, Taichi Asami, Hirokazu Masataki, Ryo Ishii, Ryuichiro Higashinaka |
INTERSPEECH | 2 |
| 2016 | Recurrent Out-of-Vocabulary Word Detection Using Distribution of FeaturesabstractThe repeated use of out-of-vocabulary (OOV) words in a spo-\nken document seriously degrades a speech recognizer’s perfor-\nmance. This paper provides a novel method for accurately de-\ntecting such recurrent OOV words. Standard OOV word de-\ntection methods classify each word segment into in-vocabulary\n(IV) or OOV. This word-by-word classification tends to be af-\nfected by sudden vocal irregularities in spontaneous speech,\ntriggering false alarms. To avoid this sensitivity to the irreg-\nularities, our proposal focuses on consistency of the repeated\noccurrence of OOV words. The proposed method preliminar-\nily detects recurrent segments, segments that contain the same\nword, in a spoken document by open vocabulary spoken term\ndiscovery using a phoneme recognizer. If the recurrent seg-\nments are OOV words, features for OOV detection in those\nsegments should exhibit consistency. We capture this consis-\ntency by using the mean and variance (distribution) of features\n(DOF) derived from the recurrent segments, and use the DOF\nfor IV/OOV classification. Experiments illustrate that the pro-\nposed method’s use of the DOF significantly improves its per-\nformance in recurrent OOV word detection.\nIndex Terms: speech recognition, OOV word detection, recur-\nrent OOV words, distribution of features Taichi Asami, Ryo Masumura, Yushi Aono, Koichi Shinoda |
INTERSPEECH | 1 |
| 2016 | Objective Evaluation Using Association Between Dimensions Within Spectral Features for Statistical Parametric Speech Synthesis
Yusuke Ijima, Taichi Asami, Hideyuki Mizuno |
INTERSPEECH | 2 |
| 2016 | Language Identification Based on Generative Modeling of Posteriorgram Sequences Extracted from Frame-by-Frame DNNs and LSTM-RNNs
Ryo Masumura, Taichi Asami, Hirokazu Masataki, Yushi Aono, Sumitaka Sakauchi |
INTERSPEECH | 2 |
| 2015 | Hierarchical Latent Words Language Models for Robust Modeling to Out-Of Domain TasksabstractThis paper focuses on language modeling with adequate robustness to support different domain tasks.To this end, we propose a hierarchical latent word language model (h-LWLM).The proposed model can be regarded as a generalized form of the standard LWLMs.The key advance is introducing a multiple latent variable space with hierarchical structure.The structure can flexibly take account of linguistic phenomena not present in the training data.This paper details the definition as well as a training method based on layer-wise inference and a practical usage in natural language processing tasks with an approximation technique.Experiments on speech recognition show the effectiveness of h-LWLM in out-of domain tasks. Ryo Masumura, Taichi Asami, Takanobu Oba, Hirokazu Masataki, Sumitaka Sakauchi, Akinori Ito |
EMNLP | 2 |
| 2015 | Agreement and disagreement utterance detection in conversational speech by extracting and integrating local features
Atsushi Ando, Taichi Asami, Manabu Okamoto, Hirokazu Masataki, Sumitaka Sakauchi |
INTERSPEECH | 2 |
| 2015 | Training data selection for acoustic modeling via submodular optimization of joint kullback-leibler divergence
Taichi Asami, Ryo Masumura, Hirokazu Masataki, Manabu Okamoto, Sumitaka Sakauchi |
INTERSPEECH | 1 |
| 2015 | Combinations of various language model technologies including data expansion and adaptation in spontaneous speech recognition
Ryo Masumura, Taichi Asami, Takanobu Oba, Hirokazu Masataki, Sumitaka Sakauchi, Akinori Ito |
INTERSPEECH | 2 |
| 2015 | Latent words recurrent neural network language models
Ryo Masumura, Taichi Asami, Takanobu Oba, Hirokazu Masataki, Sumitaka Sakauchi, Akinori Ito |
INTERSPEECH | 2 |
| 2014 | Read and spontaneous speech classification based on variance of GMM supervectors
Taichi Asami, Ryo Masumura, Hirokazu Masataki, Sumitaka Sakauchi |
INTERSPEECH | 1 |
| 2014 | Mixture of latent words language models for domain adaptation
Ryo Masumura, Taichi Asami, Takanobu Oba, Hirokazu Masataki, Sumitaka Sakauchi |
INTERSPEECH | 2 |
| 2014 | Efficient data selection for speech recognition based on prior confidence estimation using speech and monophone models
Satoshi Kobashikawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
Comput. Speech Lang. | 2 |
| 2013 | Unsupervised confidence calibration using examples of recognized words and their contexts
Taichi Asami, Satoshi Kobashikawa, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi |
INTERSPEECH | 1 |
| 2013 | Fast unsupervised adaptation based on efficient statistics accumulation using frame independent confidence within monophone states
Satoshi Kobashikawa, Atsunori Ogawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
Comput. Speech Lang. | 3 |
| 2013 | Normalizing Complex Functional Expressions in Japanese Predicates: Linguistically-Directed Rule-Based Paraphrasing and Its ApplicationabstractThe growing need for text mining systems, such as opinion mining, requires a deep semantic understanding of the target language. In order to accomplish this, extracting the semantic information of functional expressions plays a crucial role, because functional expressions such aswould like toandcan’tare key expressions to detecting customers’ needs and wants. However, in Japanese, functional expressions appear in the form of suffixes, and two different types of functional expressions are merged into one predicate: one influences the factual meaning of the predicate while the other is merely used for discourse purposes. This triggers an increase in surface forms, which hinders information extraction systems. In this article, we present a novel normalization technique that paraphrases complex functional expressions into simplified forms that retain only the crucial meaning of the predicate. We construct paraphrasing rules based on linguistic theories in syntax and semantics. The results of experiments indicate that our system achieves a high accuracy of 79.7%, while it reduces the differences in functional expressions by up to 66.7%. The results also show an improvement in the performance of predicate extraction, providing encouraging evidence of the usability of paraphrasing as a means of normalizing different language expressions. Tomoko Izumi, Kenji Imamura, Taichi Asami, Kuniko Saito, Gen-ichiro Kikui, Satoshi Sato |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2012 | Speech Data Clustering Based on Phoneme Error Trend for Unsupervised Acoustic Model Adaptation
Taichi Asami, Satoshi Kobashikawa, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi |
INTERSPEECH | 1 |
| 2012 | Efficient Beam Width Control to Suppress Excessive Speech Recognition Computation Time Based on Prior Score Range Normalization
Satoshi Kobashikawa, Takaaki Hori, Yoshikazu Yamaguchi, Taichi Asami, Hirokazu Masataki, Satoshi Takahashi |
INTERSPEECH | 4 |
| 2012 | Efficient prior and incremental beam width control to suppress excessive speech recognition time based on score range estimationabstractThis paper proposes a technique that efficiently controls the beam width to yield practical computation times when auto-transcribing massive volumes of speeches. We focus on the fact that a lot of time is wasted by recognizing poor quality speeches that will yield, with inordinate slowness, erroneous transcriptions and provide no useful results. To stabilize the time regardless of quality, our proposal controls the beam width based on prolonged score spread against the target speech; it formulates the score range within the width and maximizes computation efficiency by regulating the range relevant to the hypotheses' survival rate. The proposed technique can control the width rapidly by using just monophones prior to decoding. It also restricts the width in decoding by using the processing speed and remaining data time to better handle stubborn speeches. Experiments with several SNRs and actual call-center speeches confirm a reduction in computation time while matching the accuracy of existing techniques. Satoshi Kobashikawa, Takaaki Hori, Yoshikazu Yamaguchi, Taichi Asami, Hirokazu Masataki, Satoshi Takahashi |
SLT | 4 |
| 2011 | Extracting call-reason segments from contact center dialogs by using automatically acquired boundary expressionsabstractTo improve the performance of call-reason analysis at contact centers, we introduce a novel method to extract call-reason segments from dialogs. It is based on the following two characteristics of contact center conversations; 1) customers state their requests at the beginning of the calls, 2) agents tend to use typical phrases at the end of the call-reason segments. Our proposal acquires these typical phrases from stored speech data automatically and extracts the call-reason segment precisely by detecting the typical phrases. Experiments show that it significantly improves the performance of call-reason information retrieval since it allows the search scope to be limited to the call-reason segments of calls. Takaaki Fukutomi, Satoshi Kobashikawa, Taichi Asami, Tsubasa Shinozaki, Hirokazu Masataki, Satoshi Takahashi |
ICASSP | 3 |
| 2011 | Spoken Document Confidence Estimation Using Contextual Coherence
Taichi Asami, Narichika Nomoto, Satoshi Kobashikawa, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
INTERSPEECH | 1 |
| 2010 | Efficient data selection for speech recognition based on prior confidence estimation using speech and context independent models
Satoshi Kobashikawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
INTERSPEECH | 2 |
| 2010 | Efficient data selection for spoken document retrieval based on prior confidence estimation using speech and context independent modelsabstractThis paper proposes an efficient speech sample selection technique that can identify those samples that will be well recognized. Conventional confidence measures can identify well-recognized speech samples, but they require speech recognition to estimate confidence scores. Speech samples with low confidence should not undergo recognition since they yield speech documents that will eventually be rejected. The proposed technique can select the samples that will justify the application of speech recognition. It is based on rapid prior confidence estimation by using speech and context independent models to calculate acoustic likelihood values on a frame-by-frame basis. Tests show that the proposed confidence estimation technique is over 50 times faster than the conventional posterior confidence measure while maintaining equivalent data selection performance for speech recognition and spoken document retrieval. Satoshi Kobashikawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
SLT | 2 |
| 2006 | A Stream-Weight and Threshold Estimation Method Using Adaboost for Multi-Stream Speaker VerificationabstractThis paper proposes an automatic stream-weight and threshold estimation method for noise-robust speaker verification using multistream HMMs integrating segmental and prosodic information. The proposed method simultaneously optimizes stream-weights and a decision threshold by combining the linear discriminant analysis (LDA) and Adaboost techniques. Experiments were conducted using Japanese connected digit speech contaminated by white noise with various SNRs. In this experiment, a target ratio of false acceptance rate (FAR) and false rejection rate (FRR) was set by 1:1 so as to adjust them to approach an equal error rate (EER). Experimental results show that the proposed method effectively estimates stream-weights and thresholds so that FARs and FRRs are adjusted to EERs in most of the SNR conditions. Taichi Asami, Koji Iwano, Sadaoki Furui |
ICASSP (5) | 1 |
| 2005 | Stream-weight optimization by LDA and adaboost for multi-stream speaker verificationabstractThis paper proposes an automatic stream-weight optimization method for noise-robust speaker verification using multi-stream HMMs integrating spectral and prosodic information. The paper first shows the effectiveness of the multi-stream technique in our speaker verification framework. Next, a stream-weight adaptation method combining the linear discriminant analysis (LDA) and Adaboost techniques is proposed. Experiments were conducted using four-connected-digit utterances of Japanese contaminated by white noise with various SNRs. Experimental results show that 1) the verification performance was improved in all SNR conditions by using stream weights estimated by the LDA and 2) the performance is further improved by using the Adaboost in 10 - 30dB SNR conditions. Taichi Asami, Koji Iwano, Sadaoki Furui |
INTERSPEECH | 1 |
| 2004 | Noise-robust speaker verification using F0 featuresabstractThis paper proposes a noise-robust speaker verification method augmented by fundamental frequency (F0).The paper first describes a noise-robust F0 extraction method using the Hough transform.Then, it proposes a robust speaker verification method using multi-stream HMMs which fuse the extracted F0 and cepstral features.Experiments are conducted using fourconnected-digit utterances of Japanese by 37 male speakers recorded at five sessions over a half year period.The utterances are contaminated with white noise at various SNR levels.Experimental results show that the F0 features improve the verification performance in all SNR conditions. Koji Iwano, Taichi Asami, Sadaoki Furui |
INTERSPEECH | 2 |