VLDB 2026 Research / reviewers in the wild / expert
Takeshi Yamada
dblp:19/5706
· DBLP profile ↗
66ranked-venue papers
13as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 11 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10Applied, interdisciplinary, general and emerging computing · 5Theory of computation · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task LossabstractArray processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone location that is available during training but not at inference. However, this training objective may not be optimal for a specific array processing back-end, such as beamforming. An alternative approach is to use a training objective considering the array-processing back-end, such as a loss on the beamformer output. This approach may generate signals optimal for beamforming but not physically grounded. To combine the advantages of both approaches, this paper proposes a multi-task loss for NN-VME that combines both VM-level and beamformer-level losses. We evaluate the proposed multi-task NN-VME on multi-talker underdetermined conditions and show that it achieves a 33.1 % relative WER improvement compared to using only real microphones and 10.8 % compared to using a prior NN-VME approach. Hanako Segawa, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Shoko Araki, Takeshi Yamada, Shoji Makino |
ICASSP | 7 |
| 2023 | Toward Enhancing Software Developers' Persuasion and Engagement in GamificationabstractGamification as a method to enhance the engagement of engineers to software development, has recently received attention. Surveillance is regarded as one of the drawbacks of gamification. We assume that some participants do not like this aspect of gamification, and it can weaken the persuasion and engagement of gamification. To enhance persuasion and engagement, we propose a new approach that utilizes costly apology (i.e., apology which is accompanied with some expenses) and the anchoring effect, and to care emotions of such participants. We assume that apology is needed for participants who are not in much favor of gamification and that a costly apology is more effective than a simple apology. We used a graphical leaderboard as costly apology. The anchoring effect means that the judgment is affected by the information given before the judgment. We used a simple leaderboard as the anchoring information. We preliminarily evaluated the proposed approach through a case study. Masateru Tsunoda, Hidetsugu Suto, Takeshi Yamada |
APSEC | 3 |
| 2022 | Nonparametric Relational Models with SuperrectangulationabstractThis paper addresses the question, ”What is the smallest object that contains all rectangular partitions with n or fewer blocks?” and shows its application to relational data analysis using a new strategy we call super Bayes as an alternative to Bayesian nonparametric (BNP) methods. Conventionally, standard BNP methods have combined the Aldous-Hoover-Kallenberg representation with parsimonious stochastic processes on rectangular partitioning to construct BNP relational models. As a result, conventional methods face the great difficulty of searching for a parsimonious random rectangular partition that fits the observed data well in Bayesian inference. As a way to essentially avoid such a problem, we propose a strategy to combine an extremely redundant rectangular partition as a deterministic (non-probabilistic) object. Specifically, we introduce a special kind of rectangular partitioning, which we call superrectangulation, that contains all possible rectangular partitions. Delightfully, this strategy completely eliminates the difficult task of searching around for random rectangular partitions, since the superrectangulation is deterministically fixed in inference. Experiments on predictive performance in relational data analysis show that the super Bayesian model provides a more stable analysis than the existing BNP models, which are less likely to be trapped in bad local optima. Masahiro Nakano, Ryo Nishikimi, Yasuhiro Fujiwara, Akisato Kimura, Takeshi Yamada, Naonori Ueda |
AISTATS | 5 |
| 2022 | Performance Improvement of Speech Emotion Recognition by Neutral Speech Detection Using Autoencoder and Intermediate Representation
Jennifer Santoso, Takeshi Yamada, Kenkichi Ishizuka, Taiichi Hashimoto, Shoji Makino |
INTERSPEECH | 2 |
| 2021 | Bayesian nonparametric model for arbitrary cubic partitioningabstractIn this paper, we propose a continuous-time Markov process for cubic partitioning models of three-dimensional (3D) arrays and its application to Bayesian nonparametric relational data analysis of 3D array data. Relational data analysis is a topic that has been actively studied in the field of Bayesian nonparametrics, and in particular, models for analyzing 3D arrays have attracted much attention in recent years. In particular, the cubic partitioning model is very popular due to its practical usefulness, and various models such as the infinite relational model and the Mondrian process have been proposed. However, these conventional models have the disadvantage that they are limited to a certain class of cubic partitions, and there is a need for a model that can represent a broader class of arbitrary cubic partitions, which has long been an open issue in this field. In this study, we propose a stochastic process that can represent arbitrary cubic partitions of 3D arrays as a continuous-time Markov process. Furthermore, by combining it with the Aldous-Hoover-Kallenberg representation theorem, we construct an infinitely exchangeable 3D relational model and apply it to real data to show its application to relational data analysis. Experiments show that the proposed model improves the prediction performance by expanding the class of representable cubic partitioning. Masahiro Nakano, Yasuhiro Fujiwara, Akisato Kimura, Takeshi Yamada, Naonori Ueda |
ACML | 4 |
| 2021 | Speech Emotion Recognition Based on Attention Weight Correction Using Word-Level Confidence Measure
Jennifer Santoso, Takeshi Yamada, Shoji Makino, Kenkichi Ishizuka, Takekatsu Hiramura |
Interspeech | 2 |
| 2021 | Permuton-induced Chinese Restaurant ProcessabstractThis paper proposes the permuton-induced Chinese restaurant process (PCRP), a stochastic process on rectangular partitioning of a matrix. This distribution is suitable for use as a prior distribution in Bayesian nonparametric relational model to find hidden clusters in matrices and network data. Our main contribution is to introduce the notion of permutons into the well-known Chinese restaurant process (CRP) for sequence partitioning: a permuton is a probability measure on $[0,1]\times [0,1]$ and can be regarded as a geometric interpretation of the scaling limit of permutations. Specifically, we extend the model that the table order of CRPs has a random geometric arrangement on $[0,1]\times [0,1]$ drawn from the permuton. By analogy with the relationship between the stick-breaking process (SBP) and CRP for the infinite mixture model of a sequence, this model can be regarded as a multi-dimensional extension of CRP paired with the block-breaking process (BBP), which has been recently proposed as a multi-dimensional extension of SBP. While BBP always has an infinite number of redundant intermediate variables, PCRP can be composed of varying size intermediate variables in a data-driven manner depending on the size and quality of the observation data. Experiments show that PCRP can improve the prediction performance in relational data analysis by reducing the local optima and slow mixing problems compared with the conventional BNP models because the local transitions of PCRP in Markov chain Monte Carlo inference are more flexible than the previous models. Masahiro Nakano, Yasuhiro Fujiwara, Akisato Kimura, Takeshi Yamada, Naonori Ueda |
NeurIPS | 4 |
| 2021 | Shortest reconfiguration of sliding tokens on subclasses of interval graphs
Takeshi Yamada, Ryuhei Uehara |
Theor. Comput. Sci. | 1 |
| 2020 | Baxter Permutation ProcessabstractIn this paper, a Bayesian nonparametric (BNP) model for Baxter permutations (BPs), termed BP process (BPP) is proposed and applied to relational data analysis. The BPs are a well-studied class of permutations, and it has been demonstrated that there is one-to-one correspondence between BPs and several interesting objects including floorplan partitioning (FP), which constitutes a subset of rectangular partitioning (RP). Accordingly, the BPP can be used as an FP model. We combine the BPP with a multi-dimensional extension of the stick-breaking process called the {\it block-breaking process} to fill the gap between FP and RP, and obtain a stochastic process on arbitrary RPs. Compared with conventional BNP models for arbitrary RPs, the proposed model is simpler and has a high affinity with Bayesian inference. Masahiro Nakano, Akisato Kimura, Takeshi Yamada, Naonori Ueda |
NeurIPS | 3 |
| 2019 | Time-frequency-bin-wise Switching of Minimum Variance Distortionless Response Beamformer for Underdetermined SituationsabstractIn this paper, we present a speech enhancement method using two microphones in underdetermined situations. Time-frequency (TF) binary masking is a conventional method of enhancing speech in underdetermined situations by appropriately multiplying each TF component by zero or one. Extending this method, we previously proposed a new method called the time-frequency-bin-wise switching (TFS) beamformer. In this method, we switch multiple preconstructed beamformers in each TF bin, each of which suppresses a particular interferer. However, this method requires the pre-estimation of beamformer filter coefficients using the target-active period and interferer-wise-active periods as the prior information. In this paper, to overcome this limitation, we formulate the switching and construction of spatial filters as a joint optimization problem, which can be understood from two viewpoints: the clustering of the most dominant interferer signal in each TF bin and the construction of a minimum variance distortionless response beamformer using such bins. In an experiment, we confirmed that the proposed method was superior to conventional TF masking and fixed beamforming during speech enhancement regardless of the direction of interferers. Kouei Yamaoka, Nobutaka Ono, Shoji Makino, Takeshi Yamada |
ICASSP | 4 |
| 2018 | New Survival Prediction System for Terminal Patients based on Machine Learning
Tatsuki Hirozawa, Takeshi Yamada, Hayato Ohwada |
BIBM | 2 |
| 2015 | Cross-Domain Matching for Bag-of-Words Data via Kernel Embeddings of Latent DistributionsabstractWe propose a kernel-based method for finding matching between instances across different domains, such as multilingual documents and images with annotations. Each instance is assumed to be represented as a multiset of features, e.g., a bag-of-words representation for documents. The major difficulty in finding cross-domain relationships is that the similarity between instances in different domains cannot be directly measured. To overcome this difficulty, the proposed method embeds all the features of different domains in a shared latent space, and regards each instance as a distribution of its own features in the shared latent space. To represent the distributions efficiently and nonparametrically, we employ the framework of the kernel embeddings of distributions. The embedding is estimated so as to minimize the difference between distributions of paired instances while keeping unpaired instances apart. In our experiments, we show that the proposed method can achieve high performance on finding correspondence between multi-lingual Wikipedia articles, between documents and tags, and between images and tags. Yuya Yoshikawa, Tomoharu Iwata, Hiroshi Sawada, Takeshi Yamada |
NIPS | 4 |
| 2015 | Linear-time algorithm for sliding tokens on trees
Erik D. Demaine, Martin L. Demaine, Eli Fox-Epstein, Duc A. Hoang 0001, Takehiro Ito, Hirotaka Ono 0001, Yota Otachi, Ryuhei Uehara, Takeshi Yamada |
Theor. Comput. Sci. | 9 |
| 2014 | Rectangular Tiling ProcessabstractThis paper proposes a novel stochastic process that represents the arbitrary rectangular partitioning of an infinite-dimensional matrix as the conditional projective limit. Rectangular partitioning is used in relational data analysis, and is classified into three types: regular grid, hierarchical, and arbitrary. Conventionally, a variety of probabilistic models have been advanced for the first two, including the product of Chinese restaurant processes and the Mondrian process. However, existing models for arbitrary partitioning are too complicated to permit the analysis of the statistical behaviors of models, which places very severe capability limits on relational data analysis. In this paper, we propose a new probabilistic model of arbitrary partitioning called the rectangular tiling process (RTP). Our model has a sound mathematical base in projective systems and infinite extension of conditional probabilities, and is capable of representing partitions of infinite elements as found in ordinary Bayesian nonparametric models. Masahiro Nakano, Katsuhiko Ishiguro, Akisato Kimura, Takeshi Yamada, Naonori Ueda |
ICML | 4 |
| 2014 | Polynomial-Time Algorithm for Sliding Tokens on Trees
Erik D. Demaine, Martin L. Demaine, Eli Fox-Epstein, Duc A. Hoang 0001, Takehiro Ito, Hirotaka Ono 0001, Yota Otachi, Ryuhei Uehara, Takeshi Yamada |
ISAAC | 9 |
| 2013 | Modeling Noisy Annotated Data with Application to Social AnnotationabstractWe propose a probabilistic topic model for analyzing and extracting content-related annotations from noisy annotated discrete data such as webpages stored using social bookmarking services. With these services, because users can attach annotations freely, some annotations do not describe the semantics of the content, thus they are noisy, i.e., not content related. The extraction of content-related annotations can be used as a prepossessing step in machine learning tasks such as text classification and image recognition, or can improve information retrieval performance. The proposed model is a generative model for content and annotations, in which the annotations are assumed to originate either from topics that generated the content or from a general distribution unrelated to the content. We demonstrate the effectiveness of the proposed method by using synthetic data and real social annotation data for text and images. Tomoharu Iwata, Takeshi Yamada, Naonori Ueda |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | New analytical update rule for TDOA inference for underdetermined BSS in noisy environmentsabstractIn this paper, we propose a new technique for sparseness-based underdetermined BSS that is based on the clustering of the frequency-dependent time difference of arrival (TDOA) information and that can cope with diffused noise environments. Such a method with an EM algorithm has already been proposed, however, it required a time-consuming exhaust search for TDOA inference. To remove the need for such an exhaust search, we propose a new technique by focusing on a stereo case. We derive an update rule for analytical TDOA estimation. This update rule eliminates the need for the exhaustive TDOA search, and therefore reduces the computational load. We show experimental results for separation performance and calculation time in comparison with those obtained with the conventional approach. Our reported results validate our proposed method, that is, our proposed method achieves high performance without a high computational cost. Takuro Maruyama, Shoko Araki, Tomohiro Nakatani, Shigeki Miyabe, Takeshi Yamada, Shoji Makino, Atsushi Nakamura |
ICASSP | 5 |
| 2012 | Probabilistic Speaker Diarization With Bag-of-Words Representations of Speaker Angle InformationabstractSpeaker diarization determines “who spoke when” from the recorded conversations of an unknown number of people. In general, we have no a priori information about the number, the locations, or even the characteristics of the speakers. Additionally, speakers' speech utterances vary dynamically because of turn-taking during the conversations. These conditions make the speaker-clustering task extremely difficult. The problem becomes even harder if online (incremental) processing is required. In this paper, we formulate the speaker-clustering problem as the clustering of the sequential audio features generated by an unknown number of latent mixture components (speakers). We employ a probabilistic model that assumes time-sensitive speaker mixtures at every time frame, which, surprisingly, suits the diarization scenario. We combine the time-varying probabilistic model with direction of arrival (DOA) information calculated from a microphone array in a bag-of-words (BoW)-style feature representation. The proposed system effectively estimates the number and locations of the speakers in an online manner based on the standard Bayes inference scheme. Experiments confirm that the proposed model can successfully infer the number and features of speakers and yield better or comparable speaker diarization results compared with conventional methods in several datasets. Katsuhiko Ishiguro, Takeshi Yamada, Shoko Araki, Tomohiro Nakatani, Hiroshi Sawada |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Sequential Modeling of Topic Dynamics with Multiple TimescalesabstractWe propose an online topic model for sequentially analyzing the time evolution of topics in document collections. Topics naturally evolve with multiple timescales. For example, some words may be used consistently over one hundred years, while other words emerge and disappear over periods of a few days. Thus, in the proposed model, current topic-specific distributions over words are assumed to be generated based on the multiscale word distributions of the previous epoch. Considering both the long- and short-timescale dependency yields a more robust model. We derive efficient online inference procedures based on a stochastic EM algorithm, in which the model is sequentially updated using newly obtained data; this means that past data are not required to make the inference. We demonstrate the effectiveness of the proposed method in terms of predictive performance and computational efficiency by examining collections of real documents with timestamps. Tomoharu Iwata, Takeshi Yamada, Yasushi Sakurai, Naonori Ueda |
ACM Trans. Knowl. Discov. Data | 2 |
| 2011 | Improving Classifier Performance Using Data with Different TaxonomiesabstractWe propose a framework for improving classifier performance by effectively using auxiliary samples. The auxiliary samples are labeled not in terms of the target taxonomy according to which we wish to classify samples, but according to classification schemes or taxonomies that are different from the target taxonomy. Our method finds a classifier by minimizing a weighted error over the target and auxiliary samples. The weights are defined so that the weighted error approximates the expected error when samples are classified into the target taxonomy. Experiments using synthetic and text data show that our method significantly improves the classifier performance in most cases compared to conventional data augmentation methods. Tomoharu Iwata, Toshiyuki Tanaka 0003, Takeshi Yamada, Naonori Ueda |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2010 | Statistical Learning-based Approach for Automatic Generation System of Multiple-choice Cloze Questions
Tomoko Kojiri, Takuya Goto, Toyohide Watanabe, Tomoharu Iwata, Takeshi Yamada |
ICCE | 5 |
| 2010 | Performance estimation of noisy speech recognition considering recognition task complexityabstractTo ensure a satisfactory QoE (Quality of Experience) and facilitate system design in speech recognition services, it is essential to establish a method that can be used to efficiently investigate recognition performance in different noise environments. Previously, we proposed a performance estimation method using a spectral distortion measure. However, there is the problem that recognition task complexity affects the relationship between the recognition performance and the distortion value. To solve this problem, this paper proposes a novel performance estimation method considering the recognition task complexity. We confirmed that the proposed method gives accurate estimates of the recognition performance for various recognition tasks by an experiment using noisy speech data recorded in a real room. Index Terms: performance estimation, noisy speech recognition, recognition task difficulty Takeshi Yamada, Tomohiro Nakajima, Nobuhiko Kitawaki, Shoji Makino |
INTERSPEECH | 1 |
| 2010 | Cepstral smoothing of separated signals for underdetermined speech separationabstractMusical noise is a typical problem with blind source separation using a time-frequency mask. Recently, the cepstral smoothing of spectral masks (CSM) was proposed. Based on the idea of smoothing in the cepstral domain, this paper proposes the cepstral smoothing of separated signals (CSS) on the assumption that a cepstral representation better reflects the characteristics of speech signals than those of masks (or filter gains). We also report a comparative evaluation study of CSM and CSS with other musical noise reduction methods. Our experimental results show that CSM is effective for musical noise reduction, but the target speech was relatively distorted. On the other hand, our proposed CSS produced less distorted target signals with the same musical noise reduction as CSM. Yumi Ansa, Shoko Araki, Shoji Makino, Tomohiro Nakatani, Takeshi Yamada, Atsushi Nakamura, Nobuhiko Kitawaki |
ISCAS | 5 |
| 2010 | Online multiscale dynamic topic modelsabstractWe propose an online topic model for sequentially analyzing the time evolution of topics in document collections. Topics naturally evolve with multiple timescales. For example, some words may be used consistently over one hundred years, while other words emerge and disappear over periods of a few days. Thus, in the proposed model, current topic-specific distributions over words are assumed to be generated based on the multiscale word distributions of the previous epoch. Considering both the long-timescale dependency as well as the short-timescale dependency yields a more robust model. We derive efficient online inference procedures based on a stochastic EM algorithm, in which the model is sequentially updated using newly obtained data; this means that past data are not required to make the inference. We demonstrate the effectiveness of the proposed method in terms of predictive performance and computational efficiency by examining collections of real documents with timestamps. Tomoharu Iwata, Takeshi Yamada, Yasushi Sakurai, Naonori Ueda |
KDD | 2 |
| 2009 | Bayesian Unsupervised Word Segmentation with Nested Pitman-Yor Language Modeling
Daichi Mochihashi, Takeshi Yamada, Naonori Ueda |
ACL/IJCNLP | 2 |
| 2009 | Topic Tracking Model for Analyzing Consumer Purchase Behavior
Tomoharu Iwata, Shinji Watanabe 0001, Takeshi Yamada, Naonori Ueda |
IJCAI | 3 |
| 2009 | Modeling Social Annotation Data with Content Relevance using a Topic ModelabstractWe propose a probabilistic topic model for analyzing and extracting content-related annotations from noisy annotated discrete data such as web pages stored in social bookmarking services. In these services, since users can attach annotations freely, some annotations do not describe the semantics of the content, thus they are noisy, i.e. not content-related. The extraction of content-related annotations can be used as a preprocessing step in machine learning tasks such as text classification and image recognition, or can improve information retrieval performance. The proposed model is a generative model for content and annotations, in which the annotations are assumed to originate either from topics that generated the content or from a general distribution unrelated to the content. We demonstrate the effectiveness of the proposed method by using synthetic data and real social annotation data for text and images. Tomoharu Iwata, Takeshi Yamada, Naonori Ueda |
NIPS | 2 |
| 2009 | Personalized recommendation based on the personal innovator degreeabstractThis paper proposes a novel Collaborative Filtering scheme; it focuses on the dynamics and precedence of user preference to recommend items that match the latest preference of the target user. In predicting which items this user will purchase in the near future, the proposed algorithm identifies purchase history logs of users who have similar preferences and a high degree of purchase precedence (i.e., purchasing the same items earlier) relative to the target user. We call this metric the Personal Innovator Degree (PID). Experiments using real online sales data show that the proposed method outperforms existing methods. Noriaki Kawamae, Hitoshi Sakano, Takeshi Yamada |
RecSys | 3 |
| 2008 | Simultaneous clustering and tracking unknown number of objectsabstractIn this paper, we present a novel on-line probabilistic generative model that simultaneously deals with both the clustering and the tracking of an unknown number of moving objects. The proposed model assumes that i) time series data are composed of a time-varying number of objects and that ii) each object is governed by a mixture of an unknown number of different patterns of dynamics. The problem of learning patterns of dynamics is formulated as the clustering of tracked objects based on a nonparametric Bayesian model with conjugate priors, and this clustering in turn improves the tracking. We present a particle filter for posterior estimation of simultaneous clustering and tracking. Through experiments with synthetic and real movie data, we confirmed that the proposed model successfully learned the hidden cluster patterns and obtained better tracking results than conventional models without clustering. Katsuhiko Ishiguro, Takeshi Yamada, Naonori Ueda |
CVPR | 2 |
| 2008 | CENSREC-4: development of evaluation framework for distant-talking speech recognition under reverberant environmentsabstractIn this paper, we newly introduce a collection of databases and evaluation tools called CENSREC-4, which is an evaluation framework for distant-talking speech under hands-free conditions. Distant-talking speech recognition is crucial for a handsfree speech interface. Therefore, we measured room impulse responses to investigate reverberant speech recognition in various environments. The data contained in CENSREC-4 are connected digit utterances, as in CENSREC-1. Two subsets are included in the data: basic data sets and extra data sets. The basic data sets are used for the evaluation environment for the room impulse response-convolved speech data. The extra data sets consist of simulated and recorded data. An evaluation framework is only provided for the basic data sets as evaluation tools. The results of evaluation experiments proved that CENSREC-4 is an effective database for evaluating the new dereverberation method because the traditional dereverberation process had difficulty sufficiently improving the recognition performance. Index Terms: Various environments, Impulse response, Convolution, Real recorded data, Evaluation framework Masato Nakayama, Takanobu Nishiura, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Tetsuji Ogawa, Shigeki Matsuda, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2008 | Probabilistic latent semantic visualization: topic model for visualizing documentsabstractWe propose a visualization method based on a topic model for discrete data such as documents. Unlike conventional visualization methods based on pairwise distances such as multi-dimensional scaling, we consider a mapping from the visualization space into the space of documents as a generative process of documents. In the model, both documents and topics are assumed to have latent coordinates in a two- or three-dimensional Euclidean space, or visualization space. The topic proportions of a document are determined by the distances between the document and the topics in the visualization space, and each word is drawn from one of the topics according to its topic proportions. A visualization, i.e. latent coordinates of documents, can be obtained by fitting the model to a given set of documents using the EM algorithm, resulting in documents with similar topics being embedded close together. We demonstrate the effectiveness of the proposed model by visualizing document and movie data sets, and quantitatively compare it with conventional visualization methods. Tomoharu Iwata, Takeshi Yamada, Naonori Ueda |
KDD | 2 |
| 2008 | English Grammar Learning System Based on Knowledge Network of Fill-in-the-Blank Exercises
Takuya Goto, Tomoko Kojiri, Toyohide Watanabe, Takeshi Yamada, Tomoharu Iwata |
KES (3) | 4 |
| 2008 | Evaluation Framework for Distant-talking Speech Recognition under Reverberant Environments: newest Part of the CENSREC Series -
Takanobu Nishiura, Masato Nakayama, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
LREC | 6 |
| 2008 | Recommendation Algorithm for Learning Materials That Maximizes Expected Test Scores
Tomoharu Iwata, Tomoko Kojiri, Takeshi Yamada, Toyohide Watanabe |
PRICAI | 3 |
| 2008 | Topigraphy: visualization for large-scale tag cloudsabstractThis paper proposes a new method for displaying large-scale tag clouds. We use a topographical image that helps users to grasp the relationship among tags intuitively as a background to the tag clouds. We apply this interface to a blog navigation system and show that the proposed method enables users to find the desired tags easily even if the tag clouds are very large, 5,000 and above tags. Our approach is also effective for understanding the overall structure of a large amount of tagged documents. Ko Fujimura, Shigeru Fujimura 0002, Tatsushi Matsubayashi, Takeshi Yamada, Hidenori Okuda |
WWW | 4 |
| 2008 | Recommendation Method for Improving Customer Lifetime ValueabstractIt is important for online stores to improve customer lifetime value (LTV) if they are to increase their profits. Conventional recommendation methods suggest items that best coincide with user's interests to maximize the purchase probability, and this does not necessarily help improve LTV. We present a novel recommendation method that maximizes the probability of the LTV being improved, which can apply to both measured and subscription services. Our method finds frequent purchase patterns among high-LTV users and recommends items for a new user that simulate the found patterns. Using survival analysis techniques, we efficiently find the patterns from log data. Furthermore, we infer a user's interests from the purchase history based on maximum entropy models and use the interests to improve recommendation. Since a higher LTV is the result of greater user satisfaction, our method benefits users as well as online stores. We evaluate our method using two sets of real log data for measured and subscription services. Tomoharu Iwata, Kazumi Saito, Takeshi Yamada |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2007 | Development of VAD evaluation framework CENSREC-1-C and investigation of relationship between VAD and speech recognition performanceabstractVoice activity detection (VAD) plays an important role in speech processing including speech recognition, speech enhancement, and speech coding in noisy environments. We developed an evaluation framework for VAD in such environments, called corpus and environment for noisy speech recognition 1 concatenated (CENSREC-1-C). This framework consists of noisy continuous digit utterances and evaluation tools for VAD results. By adoptiong two evaluation measures, one for frame-level detection performance and the other for utterance-level detection performance, we provide the evaluation results of a power-based VAD method as a baseline. When using VAD in speech recognizer, the detected speech segments are extended to avoid the loss of speech frames and the pause segments are then absorbed by a pause model. We investigate the balance of an explicit segmentation by VAD and an implicit segmentation by a pause model using an experimental simulation of segment extension and show that a small extension improves speech recognition. Norihide Kitaoka, Kazumasa Yamamoto, Tomohiro Kusamizu, Seiichi Nakagawa, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Takanobu Nishiura, Masato Nakayama, Yuki Denda, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
ASRU | 5 |
| 2007 | The world of mushrooms: human-computer interaction prototype systems for ambient intelligenceabstractOur new research project called concentrates on the creation of new lifestyles through research on communication science and intelligence integration. It is premised on the creation of such virtual communication partners as fairies and goblins that can be constantly at our side. We call these virtual communication partners mushrooms.To show the essence of ambient intelligence, we developed two multimodal prototype systems: mushrooms that watch, listen, and answer questions and a Quizmaster Mushroom. These two systems work in real time using speech, sound, dialogue, and vision technologies.We performed preliminary experiments with the Quizmaster Mushroom. The results showed that the system can transmit knowledge to users while they are playing the quizzes.Furthermore, through the two mushrooms, we found policies for design effects in multimodal interface and integration. Yasuhiro Minami, Minako Sawaki, Kohji Dohsaka, Ryuichiro Higashinaka, Kentaro Ishizuka, Hideki Isozaki, Tatsushi Matsubayashi, Masato Miyoshi, Atsushi Nakamura, Takanobu Oba, Hiroshi Sawada, Takeshi Yamada, Eisaku Maeda |
ICMI | 12 |
| 2007 | Acquiring classification rules by using adaptive resonance theoryabstractWe propose two on-line classification methods, ARTMAPEDand ARTMAPAW, which are based on adaptive resonance theory. ARTMAPED, classifies cases on the basis of Euclidean distance and it incorporates category merging as a generalization technique. ARTMAPAWis the modification of ARTMAPEDto consider the importance of each attribute. The importance of attributes is updated through generalizing and specializing classification rules. Experimental results show that ARTMAPAWacquires better classification rules with fewer categories than ARTMAPED, fuzzy ARTMAP and C4.5. Hiroaki Ueda, Yo Nasu, Takeshi Yamada, Kenichi Takahashi, Tetsuhiro Miyahara |
SMC | 3 |
| 2007 | Modeling user behavior in recommender systems based on maximum entropyabstractWe propose a model for user purchase behavior in online stores that provide recommendation services. We model the purchase probability given recommendations for each user based on the maximum entropy principle using features that deal with recommendations and user interests. The proposed model enable us to measure the effect of recommendations on user purchase behavior, and the effect can be used to evaluate recommender systems. We show the validity of our model using the log data of an online cartoon distribution service, and measure the recommendation effects for evaluating the recommender system. Tomoharu Iwata, Kazumi Saito, Takeshi Yamada |
WWW | 3 |
| 2006 | Learning Systems of Concepts with an Infinite Relational Model
Charles Kemp, Josh Tenenbaum, Thomas L. Griffiths 0001, Takeshi Yamada, Naonori Ueda |
AAAI | 4 |
| 2006 | Word intelligibility estimation of noise-reduced speech
Takeshi Yamada, Masakazu Kumakura, Nobuhiko Kitawaki |
INTERSPEECH | 1 |
| 2006 | Recommendation method for extending subscription periodsabstractOnline stores providing subscription services need to extend user subscription periods as long as possible to increase their profits. Conventional recommendation methods recommend items that best coincide with user's interests to maximize the purchase probability, which does not necessarily contribute to extend subscription periods. We present a novel recommendation method for subscription services that maximizes the probability of the subscription period being extended. Our method finds frequent purchase patterns in the long subscription period users, and recommends items for a new user to simulate the found patterns. Using survival analysis techniques, we efficiently extract information from the log data for finding the patterns. Furthermore, we infer user's interests from purchase histories based on maximum entropy models, and use the interests to improve the recommendations. Since a longer subscription period is the result of greater user satisfaction, our method benefits users as well as online stores. We evaluate our method using the real log data of an online cartoon distribution service for cell-phone in Japan. Tomoharu Iwata, Kazumi Saito, Takeshi Yamada |
KDD | 3 |
| 2006 | Performance Estimation of Speech Recognition System Under Noise Conditions Using Objective Quality Measures and Artificial VoiceabstractIt is essential to ensure quality of service (QoS) when offering a speech recognition service for use in noisy environments. This means that the recognition performance in the target noise environment must be investigated. One approach is to estimate the recognition performance from a distortion value, which represents the difference between noisy speech and its original clean version. Previously, estimation methods using the segmental signal-to-noise ratio (SNRseg), the cepstral distance (CD), and the perceptual evaluation of speech quality (PESQ) have been proposed. However, their estimation accuracy has not been verified for the case when a noise reduction algorithm is adopted as a preprocessing stage in speech recognition. We, therefore, evaluated the effectiveness of these distortion measures by experiments using the AURORA-2J connected digit recognition task and four different noise reduction algorithms. The results showed that in each case the distortion measure correlates well with the word accuracy when the estimators used are optimized for each individual noise reduction algorithm. In addition, it was confirmed that when a single estimator, optimized for all the noise reduction algorithms, is used, the PESQ method gives a more accurate estimate than SNRseg and CD. Furthermore, we have proposed the use of artificial voice of several seconds duration instead of a large amount of real speech and confirmed that a relatively accurate estimate can be obtained by using the artificial voice Takeshi Yamada, Masakazu Kumakura, Nobuhiko Kitawaki |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Integration of n-best recognition results obtained by multiple noise reduction algorithmsabstractDuring the last decade, a number of noise reduction algo-rithms were proposed for realizing noise robust speech recognition. However, their effectiveness strongly de-pends on noise conditions. One way for solving this problem is to select an optimal algorithm every time be-fore or after recognition process. This paper proposes a new method for integrating N-best recognition results ob-tained by multiple noise reduction algorithms. The pro-posed method selects the best recognition result by us-ing a confidence measure based on a frame-normalized log likelihood score. To evaluate the performance of the proposed method, recognition experiments were per-formed on the AURORA-2J connected digit recognition task. These results confirmed that the proposed method is very effective in the high and middle SNR conditions. 1. Takeshi Yamada, Jiro Okada, Nobuhiko Kitawaki |
INTERSPEECH | 1 |
| 2003 | Estimation of the number of sound sources using support vector machines and its application to sound source separationabstractA method of estimating the number of sound sources in a reverberant sound field is proposed in this paper. It is known that the eigenvalue distribution of the spatial correlation matrix calculated from a multiple microphone input reflects information on the number of sources. However, in a reverberant sound field, the feature of the number of sources in the eigenvalue distribution is degraded by the room reverberation. In this paper, support vector machines is applied to classify the eigenvalue distributions which are not clearly separable. The proposed method is then applied to the source separation system and is evaluated via automatic speech recognition. Kiyoshi Yamamoto, Futoshi Asano, Willem F. G. van Rooijen, Eric Y. L. Ling, Takeshi Yamada, Nobuhiko Kitawaki |
ICASSP (5) | 5 |
| 2003 | Cross-Entropy Directed Embedding of Network Data
Takeshi Yamada, Kazumi Saito, Naonori Ueda |
ICML | 1 |
| 2003 | Control of locomotion and head configuration of 3D snake robot (SMA)abstractIn this paper, we propose a winding control technique using a physical index of horizontal constraint force for a 3D snake-like robot, and it is shown that a winding motion of the robot can be realized with small joint torque. If it approaches a target point, it is necessary to raise the head and to work like a manipulator. Therefore, a control method of the head configuration using a criterion function, which can be used in both redundant and insufficient number of link cases is proposed. In order to show the validity of the proposed methods, we constructed a snake like robot called SMA. Using the experimental system, we show that the winding pattern with which the robot can avoid singular postures is generated automatically, and head position and head orientation can be controlled properly. Masaki Yamakita, Minoru Hashimoto, Takeshi Yamada |
ICRA | 3 |
| 2003 | Integration of noise reduction algorithms for Aurora2 taskabstractTo achieve high recognition performance for a wide variety of noise and for a wide range of signal-to-noise ratios, this paper presents the integration of four noise reduction algorithms: spectral subtraction with smoothing of time direction, temporal domain SVD-based speech enhancement, GMM-based speech estimation and KLT-based comb-filtering. Recognition results on the Aurora2 task show that the effectiveness of these algorithms and their combinations strongly depends on noise conditions, and excessive noise reduction tends to degrade recognition performance in multicondition training. Takeshi Yamada, Jiro Okada, Kazuya Takeda, Norihide Kitaoka, Masakiyo Fujimoto, Shingo Kuroiwa, Kazumasa Yamamoto, Takanobu Nishiura, Mitsunori Mizumachi, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2002 | Design and collection of acoustic sound data for hands-free speech recognition and sound scene understandingabstractThe sound data for open evaluation is necessary for studies such as sound source localization, sound retrieval, sound recognition and hands-free speech recognition in real acoustic environments. This paper reports on our project for acoustic data collection. There are many kinds of sound scenes in real environments. The sound scene is specified by sound sources and room acoustics. The number of combinations of the sound sources, source positions and rooms is huge in real acoustic environments. We assumed that the sound in the environments can be simulated by convolution of the isolated sound sources and impulse responses. As an isolated sound source, hundred kinds of environment sounds and speech sounds are collected. The impulse responses are collected in various acoustic environments. Additionally we collected sounds from a moving source. In this paper, progress of our sound scene database collection project and application to environment sound recognition and hands-free speech recognition are described. Satoshi Nakamura 0001, Kazuo Hiyane, Futoshi Asano, Yutaka Kaneda, Takeshi Yamada, Takanobu Nishiura, Tetsunori Kobayashi, Shiro Ise, Hiroshi Saruwatari |
ICME (2) | 5 |
| 2002 | Suitable design of adaptive beamformer based on average speech spectrum for noisy speech recognitionabstractRecognition of distant-talking speech is indispensable for self-moving robots or teleconference systems. However, background noise and room reverberations seriously degrade the sound capture quality in real acoustic environments. A microphone array is an ideal candidate as an effective method for capturing distant-talking speech. AMNOR (Adaptive Microphone-array for NOise Reduction) was proposed an adaptive beamformer for capturing the desired distant signals in noisy environments by Kaneda et al. Although AMNOR has proven itself effective, it could be further improved if we knew the spectrum characteristics of desired distant signals in advance. Therefore, in this paper we regard speech as a desired distant signal and design AMNOR based on the average speech spectrum for distant-talking speech capture and recognition. As a result of evaluation experiments in real acoustic environments, we could confirm that the ASR (Automatic Speech Recognition) performance was improved 5 ~ 10% by AMNOR based on average speech spectrum in noisy environments. Takanobu Nishiura, Satoshi Nakamura 0001, Yuka Okada, Takeshi Yamada, Kiyohiro Shikano |
INTERSPEECH | 4 |
| 2002 | Distant-talking speech recognition based on a 3-D Viterbi search using a microphone arrayabstractThis paper focuses on microphone arrays to realize distant-talking speech recognition in real environments. In distant-talking situations, users can speak at arbitrary positions while moving. Therefore, it,is very important for high quality speech acquisition using microphone arrays to localize a talker accurately. However, it is very difficult to localize a moving talker in noisy and reverberant environments. The talker localization errors result in performance degradation of speech recognition. One way to solve this problem is to integrate the speech recognition process and the talker localization into a unified framework. This paper proposes a new speech recognition algorithm based on a three-dimensional (3-D) Viterbi search. The 3-D Viterbi method extracts a direction-time sequence of parameter vectors by steering a beam to every direction in every frame, then finds the most likely path in a 3-D trellis space composed of talker directions, input frames and HMM states. This means that speech recognition and talker localization are performed simultaneously within a statistical framework. To evaluate the performance of the 3-D Viterbi method, recognition experiments for real environment data were carried out. The results confirmed that the 3-D Viterbi method drastically improves the recognition performance for the moving talker case as well as for the fixed-position talker case. Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Localization of multiple sound sources based on a CSP analysis with a microphone arrayabstractAccurate localization of multiple sound sources is indispensable for the microphone array-based high quality sound capture. For single sound source localization, the CSP (cross-power spectrum phase analysis) method has been proposed. The CSP method localizes a sound source as a crossing point of sound directions estimated using different microphone pairs. However, when localizing multiple sound sources, the CSP method has a problem that the localization accuracy is degraded due to cross-correlation among different sound sources. To solve this problem, this paper proposes a new method which suppresses the undesired cross-correlation by synchronous addition of CSP coefficients derived from multiple microphone pairs. Experiment results in a real room showed that the proposed method improves the localization accuracy when increasing the number of the synchronous addition. Takanobu Nishiura, Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 2 |
| 2000 | Characteristics of spoken language required for objective quality evaluation of echo cancellers
Nobuhiko Kitawaki, Futoshi Asano, Takeshi Yamada |
INTERSPEECH | 3 |
| 2000 | Acoustical Sound Database in Real Environments for Sound Scene Understanding and Hands-Free Speech Recognition
Satoshi Nakamura 0001, Kazuo Hiyane, Futoshi Asano, Takanobu Nishiura, Takeshi Yamada |
LREC | 5 |
| 2000 | Speech enhancement based on the subspace methodabstractA method of speech enhancement using microphone-array signal processing based on the subspace method is proposed and evaluated. The method consists of the following two stages corresponding to the different types of noise. In the first stage, less-directional ambient noise is reduced by eliminating the noise-dominant subspace. It is realized by weighting the eigenvalues of the spatial correlation matrix. This is based on the fact that the energy of less-directional noise spreads over all eigenvalues while that of directional components is concentrated on a few dominant eigenvalues. In the second stage, the spectrum of the target source is extracted from the mixture of spectra of the multiple directional components remaining in the modified spatial correlation matrix by using a minimum variance beamformer. Finally, the proposed method is evaluated in both a simulated model environment and a real environment. Futoshi Asano, Satoru Hayamizu, Takeshi Yamada, Satoshi Nakamura 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 1999 | Simultaneous recognition of multiple sound sources based on 3-d n-best search using microphone arrayabstractThe recognition of distant talking speech in a noisy and reverberant environments is key issue in any speech recognition system. A so-called hands-free speech recognition system plays an important role in the natural and friendly human-machine interface. Considering the practical use of a speech recognition system, we realize that such a system has to deal, also, with the case of the presence of multiple sound sources, including multiple talkers, as well as other noise sources. This paper proposes a novel method which recognizes multiple talkers simultaneously in real environments by extending the 3-D Viterbi search to a 3-D N-best search algorithm. While the 3-D Viterbi method finds the most likely path in the 3-D trellis space, the proposed method considers multiple hypotheses for each direction in every frame. Combinations of the direction sequence and the phoneme sequence of multiple sources are included in the N-best list. The paper investigates the performance of the proposed method through experiments using real utterances of multiple talkers. Panikos Heracleous, Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
EUROSPEECH | 2 |
| 1999 | Data collection in real acoustical environments for sound scene understanding and hands-free speech recognitionabstractThis paper describes a sound scene database necessary for studies such as sound source localization, sound retrieval, sound recognition and hands-free speech recognition in real acoustical environments.This paper reports on a project for collection of the sound scene data supported by Real World Computing Partnership(RWCP).There are many kinds of sound scenes in real environments.The sound scene is denoted by sound sources and room acoustics.The numb e r o f c o m bination of the sound sources, source positions and rooms is huge in real acoustical environments.Two approaches are taken to build the sound scene database in the early stage of the project.The rst approach is to collect isolated sound sources of many kinds of non-speech sounds and speech sounds.The second approach is to collect impulse responses in various acoustical environments.The sound in the environments can be simulated by convolution of the isolated sound sources and impulse responses.In a later stage, the sound scene data in real acoustical environments is planned to be collected using a three dimensional microphone array.In this paper, the plan and progress of our sound scene database project are described.1. Satoshi Nakamura 0001, Kazuo Hiyane, Futoshi Asano, Takeshi Yamada, Takashi Endo |
EUROSPEECH | 4 |
| 1998 | Hands-free speech recognition based on 3-D Viterbi search using a microphone arrayabstractA microphone array is a promising solution for realizing hands-free speech recognition in real environments. Accurate talker localization is very important for speech recognition using a microphone array. However localization of a moving talker is difficult in noisy reverberant environments. Talker localization errors degrade the performance of speech recognition. To solve the problem, this paper proposes a new speech recognition algorithm which considers multiple talker direction hypotheses simultaneously. The proposed algorithm performs a Viterbi search in 3-dimensional trellis space composed of talker directions, input frames, and HMM states. As a result, a locus of the talker and a phoneme sequence of the speech are obtained by finding an optimal path with the highest likelihood. To evaluate the performance of the proposed algorithm, speech recognition experiments are carried out on simulated data and real environment data. These results show that the proposed algorithm works well even if the talker moves. Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 1 |
| 1998 | An effect of adaptive beamforming on hands-free speech recognition based on 3-d viterbi searchabstractICSLP1998: the 5th International Conference on Spoken Language Processing, November 30 - December 4, 1998, Sydney, Australia. Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICSLP | 1 |
| 1998 | Genetic Algorithms, Path Relinking, and the Flowshop Sequencing ProblemabstractIn a previous paper, a simple genetic algorithm (GA) was developed for finding (approximately) the minimum makespan of the n-job, m-machine permutation flowshop sequencing problem (PFSP). The performance of the algorithm was comparable to that of a naive neighborhood search technique and a proven simulated annealing algorithm. However, recent results have demonstrated the superiority of a tabu search method in solving the PFSP. In this paper, we reconsider the implementation of a GA for this problem and show that by taking into account the features of the landscape generated by the operators used, we are able to improve its performance significantly. Colin R. Reeves, Takeshi Yamada |
Evol. Comput. | 2 |
| 1997 | Microphone array design measures for hands-free speech recognitionabstractEUROSPEECH1997: the 5th European Conference on Speech Communication and Technology , September 22-25, 1997, Rhodes, Greece. Masaaki Inoue, Satoshi Nakamura 0001, Takeshi Yamada, Kiyohiro Shikano |
EUROSPEECH | 3 |
| 1996 | Robust speech recognition with speaker localization by a microphone arrayabstractThis paper proposes robust speech recognition with Speaker Localization by a Arrayed Microphone (SLAM) to realize hands-free speech interface in noisy environments.In order to localize a speaker direction accurately in low SNR conditions, a speaker localization algorithm based on extracting a pitch harmonics is introduced.To evaluate the performance of the proposed system, speech recognition experiments are carried out both in computer simulation and real environments.These results show that the proposed system attains the much higher speech recognition performance than that of a single microphone not only in computer simulation but also in real environments. Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICSLP | 1 |
| 1996 | Scheduling by Genetic Local Search with Multi-Step Crossover
Takeshi Yamada, Ryohei Nakano |
PPSN | 1 |
| 1994 | Optimal Population Size under Constant Computation Cost
Ryohei Nakano, Yuval Davidor, Takeshi Yamada |
PPSN | 3 |
| 1992 | A Genetic Algorithm Applicable to Large-Scale Job-Shop Problems
Takeshi Yamada, Ryohei Nakano |
PPSN | 1 |