Tomoharu Iwata

dblp:29/5953 · DBLP profile ↗
← Back
151ranked-venue papers
47as first author
49since 2021 · last 2026
0000-0003-4425-1971ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 117 · 35 first-author · 42 since 2021Databases, data management, data science and information retrieval · 46 · 15 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 9 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Deep Koopman-layered model with universal property based on toeplitz matrices
Yuka Hashimoto, Tomoharu Iwata
Neurocomputing2
2026 Meta-learning representations for learning from multiple annotators
Atsutoshi Kumagai, Tomoharu Iwata, Taishi Nishiyama, Yasutoshi Ida, Yasuhiro Fujiwara
Neurocomputing2
2026 Data-Driven Projection Generation for Efficiently Solving Heterogeneous Quadratic Programming Problems
Tomoharu Iwata, Futoshi Futami
Mach. Learn.1
2026 Aggregated Multi-output Gaussian Processes with Knowledge Transfer Across Domains
Yusuke Tanaka 0002, Toshiyuki Tanaka 0003, Tomoharu Iwata, Takeshi Kurashima, Maya Okawa, Yasunori Akagi, Hiroyuki Toda
Mach. Learn.3
2025 Meta-learning from Heterogeneous Tensors for Few-shot Tensor Completion
abstract
We propose neural network-based models for tensor completion in few observation settings. The proposed model can meta-learn inductive bias from multiple heterogeneous tensors without shared modes. Although many tensor completion methods have been proposed, the existing methods cannot leverage knowledge across heterogeneous tensors, and their performance is low when only a small number of elements are observed. The proposed model encodes each element of a given tensor by considering information about other elements while reflecting the tensor structure via a self-attention mechanism. The missing values are predicted by tensor-specific linear projection from the encoded vectors. The proposed model is shared across different tensors, and it is meta-learned such that the expected tensor completion performance is improved using multiple tensors. By experiments using synthetic and real-world tensors, we demonstrate that the proposed method achieves better performance than the existing meta-learning and tensor completion methods.
Tomoharu Iwata, Atsutoshi Kumagai
AISTATS1
2025 Meta-learning Task-specific Regularization Weights for Few-shot Linear Regression
abstract
We propose a few-shot learning method for linear regression, which learns how to choose regularization weights from multiple tasks with different feature spaces, and uses the knowledge for unseen tasks. Linear regression is ubiquitous in a wide variety of fields. Although regularization weight tuning is crucial to performance, it is difficult when only a small amount of training data are available. In the proposed method, task-specific regularization weights are generated using a neural network-based model by taking a task-specific training dataset as input, where our model is shared across all tasks. For each task, linear coefficients are optimized by minimizing the squared loss with an L2 regularizer using the generated regularization weights and the training dataset. Our model is meta-learned by minimizing the expected test error of linear regression with the task-specific coefficients using various training datasets. In our experiments using synthetic and real-world datasets, we demonstrate the effectiveness of the proposed method on few-shot regression tasks compared with existing methods.
Tomoharu Iwata, Atsutoshi Kumagai, Yasutoshi Ida
AISTATS1
2025 Importance-weighted Positive-unlabeled Learning for Distribution Shift Adaptation
abstract
Positive and unlabeled (PU) learning is a fundamental task in many applications, which trains a binary classifier from only PU data. Existing PU learning methods typically assume that training and test distributions are identical. However, this assumption is often violated due to distribution shifts, and identifying shift types such as covariate and concept shifts is generally difficult. In this paper, we propose a distribution shift adaptation method for PU learning without assuming shift types by using a few PU data in the test distribution and PU data in the training distribution. Our method is based on the importance weighting, which learns the classifier in a principled manner by minimizing the importance-weighted training risk that approximates the test risk. Although existing methods require positive and negative data in both distributions for the importance weighting without assuming shift types, we theoretically show that it can be performed with only PU data in both distributions. Based on this finding, our neural network-based classifiers can be effectively trained by iterating the importance weight estimation and classifier learning. We show that our method outperforms various existing methods with seven real-world datasets.
Atsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama, Yasuhiro Fujiwara
AISTATS2
2025 Energy-consistent Neural Operators for Hamiltonian and Dissipative Partial Differential Equations
abstract
The operator learning has received significant attention in recent years, with the aim of learning a mapping between function spaces. Prior works have proposed deep neural networks (DNNs) for learning such a mapping, enabling the learning of solution operators of partial differential equations (PDEs). However, these works still struggle to learn dynamics that obeys the laws of physics. This paper proposes Energy-consistent Neural Operators (ENOs), a general framework for learning solution operators of PDEs that follows the energy conservation or dissipation law from observed solution trajectories. We introduce a novel penalty function inspired by the energy-based theory of physics for training, in which the functional derivative is calculated making full use of automatic differentiation, allowing one to bias the outputs of the DNN-based solution operators to obey appropriate energetic behavior without explicit PDEs. Experiments on multiple systems show that ENO outperforms existing DNN models in predicting solutions from data, especially in super-resolution settings.
Yusuke Tanaka 0002, Takaharu Yaguchi, Tomoharu Iwata, Naonori Ueda
AISTATS3
2025 Hyperbolic PHATE: Visualizing Continuous Hierarchy of Latent Differentiation Structures
abstract
This paper proposes a method for embedding diffusion potentials into a hyperbolic space in order to visualize the differentiation structure consisting of diffusion and branching inherent in high-dimensional data. In recent years, the rapid development of single-cell sequencing in the field of biological information processing has made it possible to observe the evolution of gene expression levels in a snapshot-like manner as cells grow from birth to each organ or tissue. Visualization of such high-dimensional (gene pattern dimension) data is expected to provide important insights into the mechanisms of cell differentiation. Therefore, in the visualization of such data, there is a need for a system that emphasizes the "diffusion" structure that gradually shifts with time and the "branching" structure that broadly branches off into individual organs and tissues. Conventionally, the diffusion map and its extension PHATE have been developed as visualization methods specializing in diffusion structures, and hyperbolic embedding has been used as a method specializing in branching structures. However, methods that attempt to explicitly capture diffusion and branching structures simultaneously have not yet received much attention. In this paper, we focus on diffusion mapping (and its extension, PHATE), which specializes in diffusion structures, and hyperbolic embedding, which specializes in branching structures, and propose a visualization method that combines the advantages of both in order to better capture differentiaion structures consisting of diffusion and branching. As a symbolic example, we demonstrate our method using gene-cell expression data in the context of single cell analysis.
Masahiro Nakano, Hiroki Sakuma, Ryo Nishikimi, Kenji Komiya, Tomoharu Iwata, Kunio Kashino
ICASSP5
2025 Positive-Unlabeled Diffusion Models for Preventing Sensitive Data Generation
abstract
Diffusion models are powerful generative models but often generate sensitive data that are unwanted by users, mainly because the unlabeled training data frequently contain such sensitive data. Since labeling all sensitive data in the large-scale unlabeled training data is impractical, we address this problem by using a small amount of labeled sensitive data. In this paper, we propose positive-unlabeled diffusion models, which prevent the generation of sensitive data using unlabeled and sensitive data. Our approach can approximate the evidence lower bound (ELBO) for normal (negative) data using only unlabeled and sensitive (positive) data. Therefore, even without labeled normal data, we can maximize the ELBO for normal data and minimize it for labeled sensitive data, ensuring the generation of only normal data. Through experiments across various datasets and settings, we demonstrated that our approach can prevent the generation of sensitive images without compromising image quality.
Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai, Yuuki Yamanaka, Tomoya Yamashita
ICLR2
2025 Learning to Generate Projections for Reducing Dimensionality of Heterogeneous Linear Programming Problems
abstract
We propose a data-driven method for reducing the dimensionality of linear programming problems (LPs) by generating instance-specific projection matrices using a neural network-based model. Once the model is trained using multiple LPs by maximizing the expected objective value, we can efficiently find high-quality feasible solutions of newly given LPs. Our method can shorten the computational time of any LP solvers due to its solver-agnostic nature, it can provide feasible solutions by relying on projection that reduces the number of variables, and it can handle LPs of different sizes using neural networks with permutation equivariance and invariance. We also provide a theoretical analysis of the generalization bound for learning a neural network to generate projection matrices that reduce the size of LPs. Our experimental results demonstrate that our method can obtain solutions with higher quality than the existing methods, while its computational time is significantly shorter than solving the original LPs.
Tomoharu Iwata, Shinsaku Sakaue
ICML1
2025 K2IE: Kernel Method-based Kernel Intensity Estimators for Inhomogeneous Poisson Processes
abstract
Kernel method-based intensity estimators, formulated within reproducing kernel Hilbert spaces (RKHSs), and classical kernel intensity estimators (KIEs) have been among the most easy-to-implement and feasible methods for estimating the intensity functions of inhomogeneous Poisson processes. While both approaches share the term "kernel", they are founded on distinct theoretical principles, each with its own strengths and limitations. In this paper, we propose a novel regularized kernel method for Poisson processes based on the least squares loss and show that the resulting intensity estimator involves a specialized variant of the representer theorem: it has the dual coefficient of unity and coincides with classical KIEs. This result provides new theoretical insights into the connection between classical KIEs and kernel method-based intensity estimators, while enabling us to develop an efficient KIE by leveraging advanced techniques from RKHS theory. We refer to the proposed model as the kernel method-based kernel intensity estimator (K$^2$IE). Through experiments on synthetic datasets, we show that K$^2$IE achieves comparable predictive performance while significantly surpassing the state-of-the-art kernel method-based estimator in computational efficiency.
Hideaki Kim, Tomoharu Iwata, Akinori Fujino
ICML2
2025 Positive-unlabeled AUC Maximization under Covariate Shift
abstract
Maximizing the area under the receiver operating characteristic curve (AUC) is a standard approach to imbalanced binary classification tasks. Existing AUC maximization methods typically assume that training and test distributions are identical. However, this assumption is often violated due to a covariate shift, where the input distribution can vary but the conditional distribution of the class label given the input remains unchanged. The importance weighting is a common approach to the covariate shift, which minimizes the test risk with importance-weighted training data. However, it cannot maximize the AUC. In this paper, to achieve this, we theoretically derive two estimators of the test AUC risk under the covariate shift by using positive and unlabeled (PU) data in the training distribution and unlabeled data in the test distribution. Our first estimator is calculated from importance-weighted PU data in the training distribution, and the second one is calculated from importance-weighted positive data in the training distribution and unlabeled data in the test distribution. We train classifiers by minimizing a weighted sum of the two AUC risk estimators that approximates the test AUC risk. Unlike the existing importance weighting, our method does not require negative labels and class-priors. We show the effectiveness of our method with six real-world datasets.
Atsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama, Kazuki Adachi, Yasuhiro Fujiwara
ICML2
2025 Fast Proximal Gradient Methods with Node Pruning for Tree-Structured Sparse Regularization
Yasutoshi Ida, Sekitoshi Kanai, Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara
ECML/PKDD (5)4
2025 Meta-learning of Class Knowledge in Zero-shot Learning
abstract
Zero-shot learning is a promising approach to generalizing a model to categories unseen during training, and various methods have been proposed. However, they assume that class knowledge, the semantic information of the classes, is available as prior knowledge and thus fails to support domains whose class knowledge is unavailable. We propose a meta-learning method that allows us to apply the zero-shot learning method even if class knowledge is unavailable. We assume multiple zero-shot learning tasks where some classes are missing for each task, but they appear in other tasks. Our method simultaneously estimates the appropriate class knowledge and classification model using a meta-learning approach that extracts task features. We can use the trained models to perform zero-shot classification on unseen tasks without class knowledge. In experiments on datasets where true class knowledge is available, class knowledge is unavailable, and class knowledge is provided but imprecise, we show that the proposed method performs better than existing zero-shot learning methods.
Yuta Nambu, Masahiro Kohjima, Tomoharu Iwata, Ryuji Yamamoto
SDM3
2024 Zero-Shot Task Adaptation with Relevant Feature Information
abstract
We propose a method to learn prediction models such as classifiers for unseen target tasks where labeled and unlabeled data are absent but a few relevant input features for solving the tasks are given. Although machine learning requires data for training, data are often difficult to collect in practice. On the other hand, for many applications, a few relevant features would be more easily obtained. Although zero-shot learning or zero-shot domain adaptation use external knowledge to adapt to unseen classes or tasks without data, relevant features have not been used in existing studies. The proposed method improves the generalization performance on the target tasks, where there are no data but a few relevant features are given, by meta-learning from labeled data in related tasks. In the meta-learning phase, it is essential to simulate test phases on target tasks where prediction model learning is required without data. To this end, our neural network-based prediction model is meta-learned such that it correctly responds to perturbations of the relevant features on randomly generated synthetic data. By this modeling, the prediction model can explicitly learn the discriminability of the relevant features without real target data. When unlabeled training data are available in the target tasks, the proposed method can incorporate such data to boost the performance in a unified framework. Our experiments demonstrate that the proposed method outperforms various existing methods with four real-world datasets.
Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara
AAAI2
2024 Information-theoretic Analysis of Bayesian Test Data Sensitivity
abstract
Bayesian inference is often used to quantify uncertainty. Several recent analyses have rigorously decomposed uncertainty in prediction by Bayesian inference into two types: the inherent randomness in the data generation process and the variability due to lack of data respectively. Existing studies have analyzed these uncertainties from an information-theoretic perspective, assuming the model is well-specified and treating the model parameters as latent variables. However, such information-theoretic uncertainty analysis fails to account for a widely believed property of uncertainty known as sensitivity between test and training data. This means that if the test data is similar to the training data in some sense, the uncertainty will be smaller. In this study, we study such sensitivity using a new decomposition of uncertainty. Our analysis successfully defines such sensitivity using information-theoretic quantities. Furthermore, we extend the existing analysis of Bayesian meta-learning and show the novel sensitivities among tasks for the first time.
Futoshi Futami, Tomoharu Iwata
AISTATS2
2024 Warped Diffusion for Latent Differentiation Inference
abstract
This paper proposes a Bayesian nonparametric diffusion model with a black-box warping function represented by a Gaussian process to infer potential diffusion structures latent in observed data, such as differentiation mechanisms of living cells and phylogenetic evolution processes of media information. In general, the task of inferring latent differentiation structures is very difficult to handle due to two interrelated settings. One is that the conversion mechanism between hidden structure and often higher dimensional observations is unknown (and is a complex mechanism). The other is that the topology of the hidden diffuse structure itself is unknown. Therefore, in this paper, we propose a BNP-based strategy as a natural way to deal with these two challenging settings simultaneously. Specifically, as an extension of the Gaussian process latent variable model, we propose a model in which the black box transformation from latent variable space to observed data space is represented by a Gaussian process, and introduce a BNP diffusion model for the latent variable space. We show its application to the visualization of the diffusion structure of media information and to the task of inferring cell differentiation structure from single-cell gene expression levels.
Masahiro Nakano, Hiroki Sakuma, Ryo Nishikimi, Ryohei Shibue, Tomoharu Iwata, Kunio Kashino
AISTATS6
2024 Explanation-based Training with Differentiable Insertion/Deletion Metric-aware Regularizers
abstract
The quality of explanations for the predictions made by complex machine learning predictors is often measured using insertion and deletion metrics, which assess the faithfulness of the explanations, i.e., how accurately the explanations reflect the predictor’s behavior. To improve the faithfulness, we propose insertion/deletion metric-aware explanation-based optimization (ID-ExpO), which optimizes differentiable predictors to improve both the insertion and deletion scores of the explanations while maintaining their predictive accuracy. Because the original insertion and deletion metrics are non-differentiable with respect to the explanations and directly unavailable for gradient-based optimization, we extend the metrics so that they are differentiable and use them to formalize insertion and deletion metric-based regularizers. Our experimental results on image and tabular datasets show that the deep neural network-based predictors that are fine-tuned using ID-ExpO enable popular post-hoc explainers to produce more faithful and easier-to-interpret explanations while maintaining high predictive accuracy. The code is available at https://github.com/yuyay/idexpo.
Yuya Yoshikawa, Tomoharu Iwata
AISTATS2
2024 Symplectic Neural Gaussian Processes for Meta-learning Hamiltonian Dynamics
Tomoharu Iwata, Yusuke Tanaka 0002
IJCAI1
2024 Fast Iterative Hard Thresholding Methods with Pruning Gradient Computations
abstract
We accelerate the iterative hard thresholding (IHT) method, which finds \(k\) important elements from a parameter vector in a linear regression model. Although the plain IHT repeatedly updates the parameter vector during the optimization, computing gradients is the main bottleneck. Our method safely prunes unnecessary gradient computations to reduce the processing time.The main idea is to efficiently construct a candidate set, which contains \(k\) important elements in the parameter vector, for each iteration. Specifically, before computing the gradients, we prune unnecessary elements in the parameter vector for the candidate set by utilizing upper bounds on absolute values of the parameters. Our method guarantees the same optimization results as the plain IHT because our pruning is safe. Experiments show that our method is up to 73 times faster than the plain IHT without degrading accuracy.
Yasutoshi Ida, Sekitoshi Kanai, Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara
NeurIPS4
2024 AUC Maximization under Positive Distribution Shift
abstract
Maximizing the area under the receiver operating characteristic curve (AUC) is a popular approach to imbalanced binary classification problems. Existing AUC maximization methods usually assume that training and test distributions are identical. However, this assumption is often violated in practice due to {\it a positive distribution shift}, where the negative-conditional density does not change but the positive-conditional density can vary. This shift often occurs in imbalanced classification since positive data are often more diverse and time-varying than negative data. To deal with this shift, we theoretically show that the AUC on the test distribution can be expressed by using the positive and marginal training densities and the marginal test density. Based on this result, we can maximize the AUC on the test distribution by using positive and unlabeled data in the training distribution and unlabeled data in the test distribution. The proposed method requires only positive labels in the training distribution as supervision. Moreover, the derived AUC has a simple form and thus is easy to implement. The effectiveness of the proposed method is shown with four real-world datasets.
Atsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama, Yasuhiro Fujiwara
NeurIPS2
2024 Meta-learning to calibrate Gaussian processes with deep kernels for regression uncertainty estimation
Tomoharu Iwata, Atsutoshi Kumagai
Neurocomputing1
2024 Meta-learning for heterogeneous treatment effect estimation with closed-form solvers
Tomoharu Iwata, Yoichi Chikahara
Mach. Learn.1
2024 Marked point process variational autoencoder with applications to unsorted spiking activities
abstract
Spike train modeling across large neural populations is a powerful tool for understanding how neurons code information in a coordinated manner. Recent studies have employed marked point processes in neural population modeling. The marked point process is a stochastic process that generates a sequence of events with marks. Spike train models based on such processes use the waveform features of spikes as marks and express the generative structure of the unsorted spikes without applying spike sorting. In such modeling, the goal is to estimate the joint mark intensity that describes how observed covariates or hidden states (e.g., animal behaviors, animal internal states, and experimental conditions) influence unsorted spikes. A major issue with this approach is that existing joint mark intensity models are not designed to capture high-dimensional and highly nonlinear observations. To address this limitation, we propose a new joint mark intensity model based on a variational autoencoder, capable of representing the dependency structure of unsorted spikes on observed covariates or hidden states in a data-driven manner. Our model defines the joint mark intensity as a latent variable model, where a neural network decoder transforms a shared latent variable into states and marks. With our model, we derive a new log-likelihood lower bound by exploiting the variational evidence lower bound and upper bound (e.g., the χ upper bound) and use this new lower bound for parameter estimation. To demonstrate the strength of this approach, we integrate our model into a state space model with a nonlinear embedding to capture the hidden state dynamics underlying the observed covariates and unsorted spikes. This enables us to reconstruct covariates from unsorted spikes, known as neural decoding. Our model achieves superior performance in prediction and decoding tasks for synthetic data and the spiking activities of place cells.
Ryohei Shibue, Tomoharu Iwata
PLoS Comput. Biol.2
2023 Meta-learning for Robust Anomaly Detection
abstract
We propose a meta-learning method to improve the anomaly detection performance on unseen target tasks that have only unlabeled data. Existing meta-learning methods for anomaly detection have shown remarkable performance but require labeled data in target tasks. Although they can treat unlabeled data as normal assuming anomalies in the unlabeled data are negligible, this assumption is often violated in practice. As a result, the methods have low performance. Our method meta-learns with related tasks that have labeled and unlabeled data such that the expected test anomaly detection performance is directly improved when the anomaly detector is adapted to given unlabeled data. Our method is based on autoencoders (AEs), which are widely used neural network-based anomaly detectors. We model anomalous attributes for each unlabeled instance in the reconstruction loss of the AE, which are used to prevent the anomalies from being reconstructed; they can remove the effect of the anomalies. We formulate adaptation to the unlabeled data as a learning problem of the last layer of the AE and the anomalous attributes. This formulation enables the optimum solution to be obtained with a closed-form alternate update formula, which is preferable to efficiently maximize the expected test anomaly detection performance. The effectiveness of our method is experimentally shown with four real-world datasets.
Atsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Yasuhiro Fujiwara
AISTATS2
2023 Personal History Affects Reference Points: A Case Study of Codeforces
abstract
Humans make decisions based on their internal value function, and its shape is known to be distorted and biased around a point, which the research community of behavior economics refers to as the reference point. People intensify activities that come to lie within the reach of their reference point, and abstain from acts that would incur losses once they've crossed the point. However, the impact of past experiences on decision making around the reference point has not been well studied. By analyzing a long series of user-level decisions gathered from a competitive programming website, we find that history has a clear impact on user's decision making around the reference point. Past experiences can strengthen, and sometimes weaken, the decision bias around the reference point. Experiences of past difficulties can strengthen the tendency towards loss aversion after achieving the reference point. When a person crosses a reference point for the first time, the cognitive decision bias is significant. However, repeating this crossing gradually weakens the effect. We also show the value of our insights in the task of predicting user behavior. Prediction models incorporating our insights may be used for motivating people to remain more active.
Takeshi Kurashima, Tomoharu Iwata, Tomu Tominaga, Shuhei Yamamoto, Hiroyuki Toda, Kazuhisa Takemura
ICWSM2
2023 Meta-learning representations for clustering with infinite Gaussian mixture models
Tomoharu Iwata
Neurocomputing1
2023 Gaussian Process Regression With Interpretable Sample-Wise Feature Weights
abstract
Gaussian process regression (GPR) is a fundamental model used in machine learning (ML). Due to its accurate prediction with uncertainty and versatility in handling various data structures via kernels, GPR has been successfully used in various applications. However, in GPR, how the features of an input contribute to its prediction cannot be interpreted. Here, we propose GPR with local explanation, which reveals the feature contributions to the prediction of each sample while maintaining the predictive performance of GPR. In the proposed model, both the prediction and explanation for each sample are performed using an easy-to-interpret locally linear model. The weight vector of the locally linear model is assumed to be generated from multivariate Gaussian process priors. The hyperparameters of the proposed models are estimated by maximizing the marginal likelihood. For a new test sample, the proposed model can predict the values of its target variable and weight vector, as well as their uncertainties, in a closed form. Experimental results on various benchmark datasets verify that the proposed model can achieve predictive performance comparable to those of GPR and superior to that of existing interpretable models and can achieve higher interpretability than them, both quantitatively and qualitatively.
Yuya Yoshikawa, Tomoharu Iwata
IEEE Trans. Neural Networks Learn. Syst.2
2022 Predictive variational Bayesian inference as risk-seeking optimization
abstract
Since the Bayesian inference works poorly under model misspecification, various solutions have been explored to counteract the shortcomings. Recently proposed predictive Bayes (PB) that directly optimizes the Kullback Leibler divergence between the empirical distribution and the approximate predictive distribution shows excellent performances not only under model misspecification but also for over-parametrized models. However, its behavior and superiority are still unclear, which limits the applications of PB. Specifically, the superiority of PB has been shown only in terms of the predictive test log-likelihood and the performance in the sense of parameter estimation has not been investigated yet. Also, it is not clear why PB is superior with misspecified and over-parameterized models. In this paper, we clarify these ambiguities by studying PB in the framework of risk-seeking optimization. To achieve this, first, we provide a consistency theory for PB and then present intuition of robustness of PB to model misspecification using a response function theory. Thereafter, we theoretically and numerically show that PB has an implicit regularization effect that leads to flat local minima in over-parametrized models.
Futoshi Futami, Tomoharu Iwata, Naonori Ueda, Issei Sato, Masashi Sugiyama
AISTATS2
2022 Tight Integration Of Neural- And Clustering-Based Diarization Through Deep Unfolding Of Infinite Gaussian Mixture Model
abstract
Speaker diarization has been investigated extensively as an important central task for meeting analysis. Recent trend shows that integration of end-to-end neural (EEND)- and clustering-based diarization is a promising approach to handle realistic conversational data containing overlapped speech with an arbitrarily large number of speakers, and achieved state-of-the-art results on various tasks. However, the approaches proposed so far have not realized tight integration yet, because the clustering employed therein was not optimal in any sense for clustering the speaker embeddings estimated by the EEND module. To address this problem, this paper introduces a trainable clustering algorithm into the integration framework, by deep-unfolding a non-parametric Bayesian model called the infinite Gaussian mixture model (iGMM). Specifically, the speaker embeddings are optimized during training such that it better fits iGMM clustering, based on a novel clustering loss based on Adjusted Rand Index (ARI). Experimental results based on CALLHOME data show that the proposed approach outperforms the conventional approach in terms of diarization error rate (DER), especially by substantially reducing speaker confusion errors, that indeed reflects the effectiveness of the proposed iGMM integration.
Keisuke Kinoshita, Marc Delcroix, Tomoharu Iwata
ICASSP3
2022 Transfer Anomaly Detection for Maximizing the Partial AUC
abstract
The partial area under the receiver operating characteristic curve (pAUC) is a useful performance measurement for binary classification that summarizes true positive rates (TPRs) within a specific range of false positive rates (FPRs). Obtaining anomaly detectors that achieve high pAUCs is important in practice. Although several methods have been proposed for maximizing the pAUC, existing methods require both anomalous and normal instances for training. However, anomalous instances are difficult to collect due to their rarity. In this paper, we propose a method to maximize the pAUC on target tasks, in which only normal instances are available. To alleviate the absence of anomalies, the proposed method transfers useful knowledge of both anomalous and normal instances in multiple source tasks to the target tasks. Our model first embeds each instance into a task-specific latent space created from normal instances in a task using neural networks. Then, the anomaly score for each instance is obtained by a FPR range-specific linear projection, which is modeled by a neural network that tasks a FPR range as input. By this modeling, we can achieve high pAUCs within any FPR range given target normal instances without re-training. Our model is trained by maximizing the expected test pAUC within any FPR range given normal training instances using anomalous and normal instances in source tasks. We experimentally demonstrate that the proposed method outperforms various existing methods.
Atsutoshi Kumagai, Tomoharu Iwata, Taishi Nishiyama, Yasuhiro Fujiwara
IJCNN2
2022 Predicting Opinion Dynamics via Sociologically-Informed Neural Networks
abstract
Opinion formation and propagation are crucial phenomena in social networks and have been extensively studied across several disciplines. Traditionally, theoretical models of opinion dynamics have been proposed to describe the interactions between individuals (i.e., social interaction) and their impact on the evolution of collective opinions. Although these models can incorporate sociological and psychological knowledge on the mechanisms of social interaction, they demand extensive calibration with real data to make reliable predictions, requiring much time and effort. Recently, the widespread use of social media platforms provides new paradigms to learn deep learning models from a large volume of social media data. However, these methods ignore any scientific knowledge about the mechanism of social interaction. In this work, we present the first hybrid method called Sociologically-Informed Neural Network (SINN), which integrates theoretical models and social media data by transporting the concepts of physics-informed neural networks (PINNs) from natural science (i.e., physics) into social science (i.e., sociology and social psychology). In particular, we recast theoretical models as ordinary differential equations (ODEs). Then we train a neural network that simultaneously approximates the data and conforms to the ODEs that represent the social scientific knowledge. In addition, we extend PINNs by integrating matrix factorization and a language model to incorporate rich side information (e.g., user profiles) and structural knowledge (e.g., cluster structure of the social interaction network). Moreover, we develop an end-to-end training procedure for SINN, which involves Gumbel-Softmax approximation to include stochastic mechanisms of social interaction. Extensive experiments on real-world and synthetic datasets show SINN outperforms six baseline methods in predicting opinion dynamics.
Maya Okawa, Tomoharu Iwata
KDD2
2022 Learning Optimal Priors for Task-Invariant Representations in Variational Autoencoders
abstract
The variational autoencoder (VAE) is a powerful latent variable model for unsupervised representation learning. However, it does not work well in case of insufficient data points. To improve the performance in such situations, the conditional VAE (CVAE) is widely used, which aims to share task-invariant knowledge with multiple tasks through the task-invariant latent variable. In the CVAE, the posterior of the latent variable given the data point and task is regularized by the task-invariant prior, which is modeled by the standard Gaussian distribution. Although this regularization encourages independence between the latent variable and task, the latent variable remains dependent on the task. To reduce this task-dependency, the previous work introduced an additional regularizer. However, its learned representation does not work well on the target tasks. In this study, we theoretically investigate why the CVAE cannot sufficiently reduce the task-dependency and show that the simple standard Gaussian prior is one of the causes. Based on this, we propose a theoretical optimal prior for reducing the task-dependency. In addition, we theoretically show that unlike the previous work, our learned representation works well on the target tasks. Experiments on various datasets show that our approach obtains better task-invariant representations, which improves the performances of various downstream applications such as density estimation and classification.
Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai, Sekitoshi Kanai, Masanori Yamada, Yuki Yamanaka, Hisashi Kashima
KDD2
2022 Symplectic Spectrum Gaussian Processes: Learning Hamiltonians from Noisy and Sparse Data
abstract
Hamiltonian mechanics is a well-established theory for modeling the time evolution of systems with conserved quantities (called Hamiltonian), such as the total energy of the system. Recent works have parameterized the Hamiltonian by machine learning models (e.g., neural networks), allowing Hamiltonian dynamics to be obtained from state trajectories without explicit mathematical modeling. However, the performance of existing models is limited as we can observe only noisy and sparse trajectories in practice. This paper proposes a probabilistic model that can learn the dynamics of conservative or dissipative systems from noisy and sparse data. We introduce a Gaussian process that incorporates the symplectic geometric structure of Hamiltonian systems, which is used as a prior distribution for estimating Hamiltonian systems with additive dissipation. We then present its spectral representation, Symplectic Spectrum Gaussian Processes (SSGPs), for which we newly derive random Fourier features with symplectic structures. This allows us to construct an efficient variational inference algorithm for training the models while simulating the dynamics via ordinary differential equation solvers. Experiments on several physical systems show that SSGP offers excellent performance in predicting dynamics that follow the energy conservation or dissipation law from noisy and sparse data.
Yusuke Tanaka 0002, Tomoharu Iwata, Naonori Ueda
NeurIPS2
2022 Sharing Knowledge for Meta-learning with Feature Descriptions
abstract
Language is an important tool for humans to share knowledge. We propose a meta-learning method that shares knowledge across supervised learning tasks using feature descriptions written in natural language, which have not been used in the existing meta-learning methods. The proposed method improves the predictive performance on unseen tasks with a limited number of labeled data by meta-learning from various tasks. With the feature descriptions, we can find relationships across tasks even when their feature spaces are different. The feature descriptions are encoded using a language model pretrained with a large corpus, which enables us to incorporate human knowledge stored in the corpus into meta-learning. In our experiments, we demonstrate that the proposed method achieves better predictive performance than the existing meta-learning methods using a wide variety of real-world datasets provided by the statistical office of the EU and Japan.
Tomoharu Iwata, Atsutoshi Kumagai
NeurIPS1
2022 Few-shot Learning for Feature Selection with Hilbert-Schmidt Independence Criterion
abstract
We propose a few-shot learning method for feature selection that can select relevant features given a small number of labeled instances. Existing methods require many labeled instances for accurate feature selection. However, sufficient instances are often unavailable. We use labeled instances in multiple related tasks to alleviate the lack of labeled instances in a target task. To measure the dependency between each feature and label, we use the Hilbert-Schmidt Independence Criterion, which is a kernel-based independence measure. By modeling the kernel functions with neural networks that take a few labeled instances in a task as input, we can encode the task-specific information to the kernels such that the kernels are appropriate for the task. Feature selection with such kernels is performed by using iterative optimization methods, in which each update step is obtained as a closed-form. This formulation enables us to directly and efficiently minimize the expected test error on features selected by a small number of labeled instances. We experimentally demonstrate that the proposed method outperforms existing feature selection methods.
Atsutoshi Kumagai, Tomoharu Iwata, Yasutoshi Ida, Yasuhiro Fujiwara
NeurIPS2
2022 Dynamic mode decomposition via convolutional autoencoders for dynamics modeling in videos
abstract
Extracting the underlying dynamics of objects in image sequences is one of the challenging problems in computer vision . Besides, dynamic mode decomposition (DMD) has recently attracted attention as a method for obtaining modal representations of nonlinear dynamics from general multivariate time-series data without explicit prior information about the dynamics. In this paper, we propose a convolutional autoencoder (CAE)-based DMD (CAE-DMD) to perform accurate modeling of underlying dynamics in videos. We develop a modified CAE model that encodes images to latent vectors and incorporated DMD on the latent vectors to extract DMD modes. These modes are split into background and foreground modes for foreground modeling in videos, or used for video classification tasks . And the latent vectors are mapped so as to recover the input image sequences through a decoder. We perform the network training in an end-to-end manner, i.e., by minimizing the mean square error between the original and reconstructed images. As a result, we obtain accurate extraction of underlying dynamic information in the videos. We empirically investigate the performance of CAE-DMD in two applications background foreground extraction and video classification on synthetic and publicly available datasets.
Israr Ul Haq, Tomoharu Iwata, Yoshinobu Kawahara
Comput. Vis. Image Underst.2
2022 Few-shot learning for spatial regression via neural embedding-based Gaussian processes
Tomoharu Iwata, Yusuke Tanaka 0002
Mach. Learn.1
2022 Context-aware spatio-temporal event prediction via convolutional Hawkes processes
abstract
Massive spatio-temporal event data sets are now available that cover events such as disease outbreaks, armed conflicts and crimes. Predicting such events and revealing the underlying triggering patterns are a crucial task for many applications, ranging from disease control to global politics. Traditional event prediction models based on Hawkes processes capture the spatio-temporal relationships between events, but cannot incorporate complex and heterogeneous external features, including population distribution, weather and terrain. This paper proposes an event prediction method that effectively utilizes the rich external information present in sets of unstructured data (e.g., map images, satellite images and weather map). Specifically, we extend a convolutional neural network (CNN) by combining it with continuous kernel convolution; and design the conditional intensity of Hawkes process based on the extended neural network model that accepts images as its input. Our approach of using the continuous convolution kernel provides a flexible way to discover the complex effect of external factors on the triggering process, as well as yielding tractable optimization algorithms. We use real-world event data from different domains (i.e., disease outbreaks, armed conflicts and protests) to demonstrate that the proposed method has better prediction performance than existing methods.
Maya Okawa, Tomoharu Iwata, Yusuke Tanaka 0002, Takeshi Kurashima, Hiroyuki Toda, Hisashi Kashima
Mach. Learn.2
2022 Probabilistic Pedestrian Models for Estimating Unobserved Road Populations
abstract
We propose a probabilistic model of pedestrian behavior for estimating the population at each road given the observed populations at a limited number of roads and a set of routes. Our proposed model has latent variables called route populations, which represent the pedestrian population who starts each route at each time period, and assumes that pedestrians stochastically walk through roads along the route. The pedestrian depends on the road’s congestion. The proposed model incorporates this dependence in a probabilistic framework, where the time length to pass a road is assumed to follow a Gaussian distribution that depends on road congestion. With the reproducing property of Gaussian distribution, we analytically derive the transition probability between roads on each route. The parameters of the proposed method, including the dependence between congestion and speed, congestion for each road, and the latent variables on the route population, are estimated from the given data by minimizing the error between the observed and estimated road populations based on gradient-based optimization methods. In experiments with simulated and real-world data sets, we demonstrate that the proposed model can estimate road populations more accurately than existing methods. We also confirm that it effectively estimates route populations, and the estimated route populations are useful for reproducing real-world population dynamics when used as inputs of a crowd simulator based on a multi-agent system.
Tomoharu Iwata, Hitoshi Shimizu, Naoki Marumo
IEEE Trans. Intell. Transp. Syst.1
2021 Skew-symmetrically perturbed gradient flow for convex optimization
abstract
Recently, many methods for optimization and sampling have been developed by designing continuous dynamics followed by discretization. The dynamics that have been used for optimization have their corresponding underlying functionals to be minimized. On the other hand, a wider class of dynamics have been studied for sampling, which is not necessarily limited to functional minimization. For example, dynamics perturbed with skew-symmetric matrices, which cannot be seen as minimization of functionals, have been widely used to reduce asymptotic variance. Following this success in sampling, exploring such perturbed dynamics in the context of optimization can open a new avenue to optimization algorithm design. In this work, we introduce a perturbation technique for sampling into optimization for strongly convex functions. We show that perturbation applied to the gradient flow yields rapid convergence in optimization for strongly convex functions. Based on this continuous dynamics, we propose an optimization algorithm for strongly convex functions with a novel discretization framework that combines the Euler method with the leapfrog method which is used in the Hamilton Monte Carlo method. Our numerical experiments show that the perturbation technique is useful for optimization.
Futoshi Futami, Tomoharu Iwata, Naonori Ueda, Ikko Yamane
ACML2
2021 Context-aware Neural Machine Translation with Mini-batch Embedding
abstract
It is crucial to provide an inter-sentence context in Neural Machine Translation (NMT) models for higher-quality translation.With the aim of using a simple approach to incorporate inter-sentence information, we propose minibatch embedding (MBE) as a way to represent the features of sentences in a mini-batch.We construct a mini-batch by choosing sentences from the same document, and thus the MBE is expected to have contextual information across sentences.Here, we incorporate MBE in an NMT model, and our experiments show that the proposed method consistently outperforms the translation capabilities of strong baselines and improves writing style or terminology to fit the document's context. 1
Makoto Morishita, Jun Suzuki 0001, Tomoharu Iwata, Masaaki Nagata
EACL3
2021 Semi-supervised Anomaly Detection on Attributed Graphs
abstract
We propose a simple yet effective method for detecting anomalous instances on an attribute graph with label information of a small number of instances. Although standard anomaly detection methods usually assume that instances are independent and identically distributed, in many real-world applications, instances are often explicitly connected, resulting in so-called attributed graphs. The proposed method embeds nodes (instances) on the attributed graph in a latent space by taking into account their attributes as well as the graph structure on the basis of graph convolutional networks (GCNs). To learn node embeddings specialized for anomaly detection, in which there is a class imbalance due to the rarity of anomalies, the parameters of a GCN are trained to minimize the volume of a hypersphere that encloses the node embeddings of normal instances while embedding anomalous ones outside the hypersphere. This enables us to detect anomalies by simply calculating the distances between the node embeddings and hypersphere center. The proposed method can effectively propagate label information on a small amount of nodes to unlabeled ones by taking into account the node's attributes, graph structure, and class imbalance. In experiments with five real-world attributed graph datasets, we demonstrate that the proposed method outperforms various existing anomaly detection methods.
Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara
IJCNN2
2021 Fault Detection of ICT systems with Deep Learning Model for Missing Data
Kengo Tajiri, Tomoharu Iwata, Yoichi Matsuo, Keishiro Watanabe
IM2
2021 Dynamic Hawkes Processes for Discovering Time-evolving Communities' States behind Diffusion Processes
abstract
Sequences of events including infectious disease outbreaks, social network activities, and crimes are ubiquitous and the data on such events carry essential information about the underlying diffusion processes between communities (e.g., regions, online user groups). Modeling diffusion processes and predicting future events are crucial in many applications including epidemic control, viral marketing, and predictive policing. Hawkes processes offer a central tool for modeling the diffusion processes, in which the influence from the past events is described by the triggering kernel. However, the triggering kernel parameters, which govern how each community is influenced by the past events, are assumed to be static over time. In the real world, the diffusion processes depend not only on the influences from the past, but also the current (time-evolving) states of the communities, e.g., people's awareness of the disease and people's current interests. In this paper, we propose a novel Hawkes process model that is able to capture the underlying dynamics of community states behind the diffusion processes and predict the occurrences of events based on the dynamics. Specifically, we model the latent dynamic function that encodes these hidden dynamics by a mixture of neural networks. Then we design the triggering kernel using the latent dynamic function and its integral. The proposed method, termed DHP (Dynamic Hawkes Processes), offers a flexible way to learn complex representations of the time-evolving communities' states, while at the same time it allows to computing the exact likelihood, which makes parameter learning tractable. Extensive experiments on four real-world event datasets show that DHP outperforms five widely adopted methods for event prediction.
Maya Okawa, Tomoharu Iwata, Yusuke Tanaka 0002, Hiroyuki Toda, Takeshi Kurashima, Hisashi Kashima
KDD2
2021 Loss function based second-order Jensen inequality and its application to particle variational inference
abstract
Bayesian model averaging, obtained as the expectation of a likelihood function by a posterior distribution, has been widely used for prediction, evaluation of uncertainty, and model selection. Various approaches have been developed to efficiently capture the information in the posterior distribution; one such approach is the optimization of a set of models simultaneously with interaction to ensure the diversity of the individual models in the same way as ensemble learning. A representative approach is particle variational inference (PVI), which uses an ensemble of models as an empirical approximation for the posterior distribution. PVI iteratively updates each model with a repulsion force to ensure the diversity of the optimized models. However, despite its promising performance, a theoretical understanding of this repulsion and its association with the generalization ability remains unclear. In this paper, we tackle this problem in light of PAC-Bayesian analysis. First, we provide a new second-order Jensen inequality, which has the repulsion term based on the loss function. Thanks to the repulsion term, it is tighter than the standard Jensen inequality. Then, we derive a novel generalization error bound and show that it can be reduced by enhancing the diversity of models. Finally, we derive a new PVI that optimizes the generalization error bound directly. Numerical experiments demonstrate that the performance of the proposed PVI compares favorably with existing methods in the experiment.
Futoshi Futami, Tomoharu Iwata, Naonori Ueda, Issei Sato, Masashi Sugiyama
NeurIPS2
2021 Meta-Learning for Relative Density-Ratio Estimation
abstract
The ratio of two probability densities, called a density-ratio, is a vital quantity in machine learning. In particular, a relative density-ratio, which is a bounded extension of the density-ratio, has received much attention due to its stability and has been used in various applications such as outlier detection and dataset comparison. Existing methods for (relative) density-ratio estimation (DRE) require many instances from both densities. However, sufficient instances are often unavailable in practice. In this paper, we propose a meta-learning method for relative DRE, which estimates the relative density-ratio from a few instances by using knowledge in related datasets. Specifically, given two datasets that consist of a few instances, our model extracts the datasets' information by using neural networks and uses it to obtain instance embeddings appropriate for the relative DRE. We model the relative density-ratio by a linear model on the embedded space, whose global optimum solution can be obtained as a closed-form solution. The closed-form solution enables fast and effective adaptation to a few instances, and its differentiability enables us to train our model such that the expected test error for relative DRE can be explicitly minimized after adapting to a few instances. We empirically demonstrate the effectiveness of the proposed method by using three problems: relative DRE, dataset comparison, and outlier detection.
Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara
NeurIPS2
2021 Time-delayed collective flow diffusion models for inferring latent people flow from aggregated data at limited locations
abstract
The rapid adoption of wireless sensor devices has made it easier to record location information of people in a variety of spaces (e.g., exhibition halls). Location information is often aggregated due to privacy and/or cost concerns. The aggregated data we use as input consist of the numbers of incoming and outgoing people at each location and at each time step. Since the aggregated data lack tracking information of individuals, determining the flow of people between locations is not straightforward. In this article, we address the problem of inferring latent people flows, that is, transition populations between locations, from just aggregated population data gathered from observed locations. Existing models assume that everyone is always in one of the observed locations at every time step; this, however, is an unrealistic assumption, because we do not always have a large enough number of sensor devices to cover the large-scale spaces targeted. To overcome this drawback, we propose a probabilistic model with flow conservation constraints that incorporate travel duration distributions between observed locations. To handle noisy settings, we adopt noisy observation models for the numbers of incoming and outgoing people, where the noise is regarded as a factor that may disturb flow conservation, e.g., people may appear in or disappear from the predefined space of interest. We develop an approximate expectation-maximization (EM) algorithm that simultaneously estimates transition populations and model parameters. Our experiments demonstrate the effectiveness of the proposed model on real-world datasets of pedestrian data in exhibition halls, bike trip data and taxi trip data in New York City.
Yusuke Tanaka 0002, Tomoharu Iwata, Takeshi Kurashima, Hiroyuki Toda, Naonori Ueda, Toshiyuki Tanaka 0003
Artif. Intell.2
2020 Semi-Supervised Learning for Maximizing the Partial AUC
abstract
The partial area under a receiver operating characteristic curve (pAUC) is a performance measurement for binary classification problems that summarizes the true positive rate with the specific range of the false positive rate. Obtaining classifiers that achieve high pAUC is important in a wide variety of applications, such as cancer screening and spam filtering. Although many methods have been proposed for maximizing the pAUC, existing methods require many labeled data for training. In this paper, we propose a semi-supervised learning method for maximizing the pAUC, which trains a classifier with a small amount of labeled data and a large amount of unlabeled data. To exploit the unlabeled data, we derive two approximations of the pAUC: the first is calculated from positive and unlabeled data, and the second is calculated from negative and unlabeled data. A classifier is trained by maximizing the weighted sum of the two approximations of the pAUC and the pAUC that is calculated from positive and negative data. With experiments using various datasets, we demonstrate that the proposed method achieves higher test pAUCs than existing methods.
Tomoharu Iwata, Akinori Fujino, Naonori Ueda
AAAI1
2020 Co-Occurrence Estimation from Aggregated Data with Auxiliary Information
Tomoharu Iwata, Naoki Marumo
AAAI1
2020 Disentangled Representations for Sequence Data using Information Bottleneck Principle
abstract
We propose the factorizing variational autoencoder (FAVAE), a generative model for learning dis- entangled representations from sequential data via the information bottleneck principle without supervision. Real-world data are often generated by a few explanatory factors of variation, and disentangled representation learning obtains these factors from the data. We focus on the disen- tangled representation of sequential data which can be useful in a wide range of applications, such as video, speech, and stock markets. Factors in sequential data are categorized into dynamic and static ones: dynamic factors are time dependent, and static factors are time independent. Previous models disentangle between static and dynamic factors and between dynamic factors with different time dependencies by explicitly modeling the priors of latent variables. However, these models cannot disentangle representations between dynamic factors with the same time dependency, such as disentangling “picking up” and “throwing” in robotic tasks. On the other hand, FAVAE can disentangle multiple dynamic factors via the information bottleneck principle where it does not require modeling priors. We conducted experiments to show that FAVAE can extract disentangled dynamic factors on synthetic, video, and speech datasets.
Masanori Yamada, Heecheol Kim 0002, Kosuke Miyoshi, Tomoharu Iwata, Hiroshi Yamakawa
ACML4
2020 Frame-Level Phoneme-Invariant Speaker Embedding for Text-Independent Speaker Recognition on Extremely Short Utterances
abstract
This paper investigates a phoneme-invariant speaker embedding approach for speaker recognition on extremely short utterances. Intuitively, phonemes are nuisance information for text-independent speaker recognition task since the contents of the speech are usually mismatched between enrolling and testing time. However, many studies have shown that incorporating phoneme information is quite effective to improve the performance of the speaker recognition system. One reasonable explanation for this counter-intuitive result is that the pooling mechanism of segment-based speaker embedding can focus on the specific phonemes which contain rich speaker information, and phoneme information may help this. From this insight, we hypothesize that the pooling mechanism and phoneme-aware training are harmful to extract the speaker embeddings from extremely short utterances. To verify this hypothesis, an adversarial framework is introduced to remove phoneme-variability from the frame-wise speaker embeddings. The experimental results on the Librispeech corpus confirm that our frame-wise, phoneme-adversarial approach outperforms the conventional segment-wise, phoneme-aware approach for short utterances of less than about 1.4 seconds.
Naohiro Tawara, Atsunori Ogawa, Tomoharu Iwata, Marc Delcroix, Tetsuji Ogawa
ICASSP3
2020 Fast Deterministic CUR Matrix Decomposition with Accuracy Assurance
abstract
The deterministic CUR matrix decomposition is a low-rank approximation method to analyze a data matrix. It has attracted considerable attention due to its high interpretability, which results from the fact that the decomposed matrices consist of subsets of the original columns and rows of the data matrix. The subset is obtained by optimizing an objective function with sparsity-inducing norms via coordinate descent. However, the existing algorithms for optimization incur high computation costs. This is because coordinate descent iteratively updates all the parameters in the objective until convergence. This paper proposes a fast deterministic CUR matrix decomposition. Our algorithm safely skips unnecessary updates by efficiently evaluating the optimality conditions for the parameters to be zeros. In addition, we preferentially update the parameters that must be nonzeros. Theoretically, our approach guarantees the same result as the original approach. Experiments demonstrate that our algorithm speeds up the deterministic CUR while achieving the same accuracy.
Yasutoshi Ida, Sekitoshi Kanai, Yasuhiro Fujiwara, Tomoharu Iwata, Koh Takeuchi 0001, Hisashi Kashima
ICML4
2020 Reinforcement Learning in Latent Action Sequence Space
abstract
One problem in real-world applications of reinforcement learning is the high dimensionality of the action search spaces, which comes from the combination of actions over time. To reduce the dimensionality of action sequence search spaces, macro actions have been studied, which are sequences of primitive actions to solve tasks. However, previous studies relied on humans to define macro actions or assumed macro actions to be repetitions of the same primitive actions. We propose encoded action sequence reinforcement learning (EASRL), a reinforcement learning method that learns flexible sequences of actions in a latent space for a high-dimensional action sequence search space. With EASRL, encoder and decoder networks are trained with demonstration data by using variational autoencoders for mapping macro actions into the latent space. Then, we learn a policy network in the latent space, which is a distribution over encoded macro actions given a state. By learning in the latent space, we can reduce the dimensionality of the action sequence search space and handle various patterns of action sequences. We experimentally demonstrate that the proposed method outperforms other reinforcement learning methods on tasks that require an extensive amount of search.
Heecheol Kim 0002, Masanori Yamada, Kosuke Miyoshi, Tomoharu Iwata, Hiroshi Yamakawa
IROS4
2020 Meta-learning from Tasks with Heterogeneous Attribute Spaces
abstract
We propose a heterogeneous meta-learning method that trains a model on tasks with various attribute spaces, such that it can solve unseen tasks whose attribute spaces are different from the training tasks given a few labeled instances. Although many meta-learning methods have been proposed, they assume that all training and target tasks share the same attribute space, and they are inapplicable when attribute sizes are different across tasks. Our model infers latent representations of each attribute and each response from a few labeled instances using an inference network. Then, responses of unlabeled instances are predicted with the inferred representations using a prediction network. The attribute and response representations enable us to make predictions based on the task-specific properties of attributes and responses even when attribute and response sizes are different across tasks. In our experiments with synthetic datasets and 59 datasets in OpenML, we demonstrate that our proposed method can predict the responses given a few labeled instances in new tasks after being trained with tasks with heterogeneous attribute spaces.
Tomoharu Iwata, Atsutoshi Kumagai
NeurIPS1
2020 Deep Reinforcement Learning for Pedestrian Guidance
Hitoshi Shimizu, Takanori Hara 0002, Tomoharu Iwata
PRIMA3
2020 Transfer Metric Learning for Unseen Domains
abstract
Abstract We propose a transfer metric learning method to infer domain-specific data embeddings for unseen domains, from which no data are given in the training phase, by using knowledge transferred from related domains. When training and test distributions are different, the standard metric learning cannot infer appropriate data embeddings. The proposed method can infer appropriate data embeddings for the unseen domains by using latent domain vectors, which are latent representations of domains and control the property of data embeddings for each domain. This latent domain vector is inferred by using a neural network that takes the set of feature vectors in the domain as an input. The neural network is trained without the unseen domains. The proposed method can instantly infer data embeddings for the unseen domains without (re)-training once the sets of feature vectors in the domains are given. To accumulate knowledge in advance, the proposed method uses labeled and unlabeled data in multiple source domains. Labeled data, i.e., data with label information such as class labels or pair (similar/dissimilar) constraints, are used for learning data embeddings in such a way that similar data points are close and dissimilar data points are separated in the embedding space. Although unlabeled data do not have labels, they have geometric information that characterizes domains. The proposed method incorporates this information in a natural way on the basis of a probabilistic framework. The conditional distributions of the latent domain vectors, the embedded data, and the observed data are parameterized by neural networks and are optimized by maximizing the variational lower bound using stochastic gradient descent. The effectiveness of the proposed method was demonstrated through experiments using three clustering tasks.
Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara
Data Sci. Eng.2
2020 Anomaly detection with inexact labels
Tomoharu Iwata, Machiko Toyoda, Shotaro Tora, Naonori Ueda
Mach. Learn.1
2019 Refining Coarse-Grained Spatial Data Using Auxiliary Spatial Data Sets with Various Granularities
abstract
We propose a probabilistic model for refining coarse-grained spatial data by utilizing auxiliary spatial data sets. Existing methods require that the spatial granularities of the auxiliary data sets are the same as the desired granularity of target data. The proposed model can effectively make use of auxiliary data sets with various granularities by hierarchically incorporating Gaussian processes. With the proposed model, a distribution for each auxiliary data set on the continuous space is modeled using a Gaussian process, where the representation of uncertainty considers the levels of granularity. The finegrained target data are modeled by another Gaussian process that considers both the spatial correlation and the auxiliary data sets with their uncertainty. We integrate the Gaussian process with a spatial aggregation process that transforms the fine-grained target data into the coarse-grained target data, by which we can infer the fine-grained target Gaussian process from the coarse-grained data. Our model is designed such that the inference of model parameters based on the exact marginal likelihood is possible, in which the variables of finegrained target and auxiliary data are analytically integrated out. Our experiments on real-world spatial data sets demonstrate the effectiveness of the proposed model.
Yusuke Tanaka 0002, Tomoharu Iwata, Toshiyuki Tanaka 0003, Takeshi Kurashima, Maya Okawa, Hiroyuki Toda
AAAI2
2019 Neural Collective Graphical Models for Estimating Spatio-Temporal Population Flow from Aggregated Data
abstract
We propose a probabilistic model for estimating population flow, which is defined as populations of the transition between areas over time, given aggregated spatio-temporal population data. Since there is no information about individual trajectories in the aggregated data, it is not straightforward to estimate population flow. With the proposed method, we utilize a collective graphical model with which we can learn individual transition models from the aggregated data by analytically marginalizing the individual locations. Learning a spatio-temporal collective graphical model only from the aggregated data is an ill-posed problem since the number of parameters to be estimated exceeds the number of observations. The proposed method reduces the effective number of parameters by modeling the transition probabilities with a neural network that takes the locations of the origin and the destination areas and the time of day as inputs. By this modeling, we can automatically learn nonlinear spatio-temporal relationships flexibly among transitions, locations, and times. With four real-world population data sets in Japan and China, we demonstrate that the proposed method can estimate the transition population more accurately than existing methods.
Tomoharu Iwata, Hitoshi Shimizu
AAAI1
2019 Unsupervised Domain Adaptation by Matching Distributions Based on the Maximum Mean Discrepancy via Unilateral Transformations
abstract
We propose a simple yet effective method for unsupervised domain adaptation. When training and test distributions are different, standard supervised learning methods perform poorly. Semi-supervised domain adaptation methods have been developed for the case where labeled data in the target domain are available. However, the target data are often unlabeled in practice. Therefore, unsupervised domain adaptation, which does not require labels for target data, is receiving a lot of attention. The proposed method minimizes the discrepancy between the source and target distributions of input features by transforming the feature space of the source domain. Since such unilateral transformations transfer knowledge in the source domain to the target one without reducing dimensionality, the proposed method can effectively perform domain adaptation without losing information to be transfered. With the proposed method, it is assumed that the transformed features and the original features differ by a small residual to preserve the relationship between features and labels. This transformation is learned by aligning the higher-order moments of the source and target feature distributions based on the maximum mean discrepancy, which enables to compare two distributions without density estimation. Once the transformation is found, we learn supervised models by using the transformed source data and their labels. We use two real-world datasets to demonstrate experimentally that the proposed method achieves better classification performance than existing methods for unsupervised domain adaptation.
Atsutoshi Kumagai, Tomoharu Iwata
AAAI2
2019 Variational Autoencoder with Implicit Optimal Priors
abstract
The variational autoencoder (VAE) is a powerful generative model that can estimate the probability of a data point by using latent variables. In the VAE, the posterior of the latent variable given the data point is regularized by the prior of the latent variable using Kullback Leibler (KL) divergence. Although the standard Gaussian distribution is usually used for the prior, this simple prior incurs over-regularization. As a sophisticated prior, the aggregated posterior has been introduced, which is the expectation of the posterior over the data distribution. This prior is optimal for the VAE in terms of maximizing the training objective function. However, KL divergence with the aggregated posterior cannot be calculated in a closed form, which prevents us from using this optimal prior. With the proposed method, we introduce the density ratio trick to estimate this KL divergence without modeling the aggregated posterior explicitly. Since the density ratio trick does not work well in high dimensions, we rewrite this KL divergence that contains the high-dimensional density ratio into the sum of the analytically calculable term and the lowdimensional density ratio term, to which the density ratio trick is applied. Experiments on various datasets show that the VAE with this implicit optimal prior achieves high density estimation performance.
Hiroshi Takahashi, Tomoharu Iwata, Yuki Yamanaka, Masanori Yamada, Satoshi Yagi
AAAI2
2019 Unsupervised Multilingual Word Embedding with Limited Resources using Neural Language Models
abstract
Recently, a variety of unsupervised methods have been proposed that map pre-trained word embeddings of different languages into the same space without any parallel data.These methods aim to find a linear transformation based on the assumption that monolingual word embeddings are approximately isomorphic between languages.However, it has been demonstrated that this assumption holds true only on specific conditions, and with limited resources, the performance of these methods decreases drastically.To overcome this problem, we propose a new unsupervised multilingual embedding method that does not rely on such assumption and performs well under resource-poor scenarios, namely when only a small amount of monolingual data (i.e., 50k sentences) are available, or when the domains of monolingual data are different across languages.Our proposed model, which we call 'Multilingual Neural Language Models', shares some of the network parameters among multiple languages, and encodes sentences of multiple languages into the same space.The model jointly learns word embeddings of different languages in the same space, and generates multilingual embeddings without any parallel data or pre-training.Our experiments on word alignment tasks have demonstrated that, on the low-resource condition, our model substantially outperforms existing unsupervised and even supervised methods trained with 500 bilingual pairs of words.Our model also outperforms unsupervised methods given different-domain corpora across languages.Our code is publicly available 1 .
Takashi Wada 0001, Tomoharu Iwata, Yuji Matsumoto 0001
ACL (1)2
2019 A Unified Framework for Feature-based Domain Adaptation of Neural Network Language Models
abstract
An important task for language models is the adaptation of general-domain models to specific target domains. For neural network-based language models, feature-based domain adaptation has been a popular method in previous research. Conventional methods use an adaptation feature providing context information that is calculated from a topic model. However, such a topic model needs to be trained separately from the language model. To unify the language and context model training, we present an approach that combines an extractor network and a domain adaptation layer. The extractor network learns a context representation from a fixed-size window of past words and provides the context information for the adaptation layer. The benefit of our method is that the extractor network can be trained jointly with the language model in a single training step. Our proposed method showed superior performance over conventional domain adaptation with topic features on a dataset of TED talks with respect to perplexity and word error rate after 100-best rescoring.
Michael Hentschel, Marc Delcroix, Atsunori Ogawa, Tomoharu Iwata, Tomohiro Nakatani
ICASSP4
2019 Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders
abstract
We introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speech (TTS) encoder decoder architectures. These autoencoders learn features from speech only and text only datasets by switching the encoders and decoders used in the ASR and TTS models. Simultaneously, they aim to encode features to be compatible with ASR and TTS models by a multi-task loss. Additionally, we anticipate that TTS joint training can also improve the ASR performance because both ASR and TTS models learn transformations between speech and text. The experimental result we obtained with our semi-supervised end-to-end ASR/TTS training revealed reductions from a model initially trained with a small paired subset of the LibriSpeech corpus in the character error rate from 10.4% to 8.4% and word error rate from 20.6% to 18.0% by retraining the model with a large unpaired subset of the corpus.
Shigeki Karita, Shinji Watanabe 0001, Tomoharu Iwata, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani
ICASSP3
2019 Transfer Metric Learning for Unseen Domains
abstract
We propose a transfer metric learning method to infer domain-specific data embeddings for unseen domains, from which no data are given in the training phase, by using knowledge transferred from related domains. When training and test distributions are different, the standard metric learning cannot infer appropriate data embeddings. The proposed method can infer appropriate data embeddings for the unseen domains by using latent domain vectors, which are latent representations of domains and control the property of data embeddings for each domain. This latent domain vector is inferred by using a neural network that takes the set of feature vectors in the domain as an input. The neural network is trained without the unseen domains. The proposed method can instantly infer data embeddings for the unseen domains without (re)-training once the sets of feature vectors in the domains are given. To accumulate knowledge in advance, the proposed method uses labeled and unlabeled data in multiple source domains. Labeled data, i.e., data with label information such as class labels or pair (similar/dissimilar) constraints, are used for learning data embeddings in such a way that similar data points are close and dissimilar data points are separated in the embedding space. Although unlabeled data do not have labels, they have geometric information that characterizes domains. The proposed method incorporates this information in a natural way on the basis of a probabilistic framework. The conditional distributions of the latent domain vectors, the embedded data, and the observed data are parametrized by neural networks and are optimized by maximizing the variational lower bound. The effectiveness of the proposed method was demonstrated through experiments using three clustering tasks.
Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara
ICDM2
2019 Deep Mixture Point Processes: Spatio-temporal Event Prediction with Rich Contextual Information
abstract
Predicting when and where events will occur in cities, like taxi pick-ups, crimes, and vehicle collisions, is a challenging and important problem with many applications in fields such as urban planning, transportation optimization and location-based marketing. Though many point processes have been proposed to model events in a continuous spatio-temporal space, none of them allow for the consideration of the rich contextual factors that affect event occurrence, such as weather, social activities, geographical characteristics, and traffic. In this paper, we propose DMPP (Deep Mixture Point Processes), a point process model for predicting spatio-temporal events with the use of rich contextual information; a key advance is its incorporation of the heterogeneous and high-dimensional context available in image and text data. Specifically, we design the intensity of our point process model as a mixture of kernels, where the mixture weights are modeled by a deep neural network. This formulation allows us to automatically learn the complex nonlinear effects of the contextual factors on event occurrence. At the same time, this formulation makes analytical integration over the intensity, which is required for point process estimation, tractable. We use real-world data sets from different domains to demonstrate that DMPP has better predictive performance than existing methods.
Maya Okawa, Tomoharu Iwata, Takeshi Kurashima, Yusuke Tanaka 0002, Hiroyuki Toda, Naonori Ueda
KDD2
2019 Spatially Aggregated Gaussian Processes with Multivariate Areal Outputs
abstract
We propose a probabilistic model for inferring the multivariate function from multiple areal data sets with various granularities. Here, the areal data are observed not at location points but at regions. Existing regression-based models can only utilize the sufficiently fine-grained auxiliary data sets on the same domain (e.g., a city). With the proposed model, the functions for respective areal data sets are assumed to be a multivariate dependent Gaussian process (GP) that is modeled as a linear mixing of independent latent GPs. Sharing of latent GPs across multiple areal data sets allows us to effectively estimate the spatial correlation for each areal data set; moreover it can easily be extended to transfer learning across multiple domains. To handle the multivariate areal data, we design an observation model with a spatial aggregation process for each areal data set, which is an integral of the mixed GP over the corresponding region. By deriving the posterior GP, we can predict the data value at any location point by considering the spatial correlations and the dependences between areal data sets, simultaneously. Our experiments on real-world data sets demonstrate that our model can 1) accurately refine coarse-grained areal data, and 2) offer performance improvements by using the areal data sets from multiple domains.
Yusuke Tanaka 0002, Toshiyuki Tanaka 0003, Tomoharu Iwata, Takeshi Kurashima, Maya Okawa, Yasunori Akagi, Hiroyuki Toda
NeurIPS3
2019 Transfer Anomaly Detection by Inferring Latent Domain Representations
abstract
We propose a method to improve the anomaly detection performance on target domains by transferring knowledge on related domains. Although anomaly labels are valuable to learn anomaly detectors, they are difficult to obtain due to their rarity. To alleviate this problem, existing methods use anomalous and normal instances in the related domains as well as target normal instances. These methods require training on each target domain. However, this requirement can be problematic in some situations due to the high computational cost of training. The proposed method can infer the anomaly detectors for target domains without re-training by introducing the concept of latent domain vectors, which are latent representations of the domains and are used for inferring the anomaly detectors. The latent domain vector for each domain is inferred from the set of normal instances in the domain. The anomaly score function for each domain is modeled on the basis of autoencoders, and its domain-specific property is controlled by the latent domain vector. The anomaly score function for each domain is trained so that the scores of normal instances become low and the scores of anomalies become higher than those of the normal instances, while considering the uncertainty of the latent domain vectors. When target normal instances can be used during training, the proposed method can also use them for training in a unified framework. The effectiveness of the proposed method is demonstrated through experiments using one synthetic and four real-world datasets. Especially, the proposed method without re-training outperforms existing methods with target specific training.
Atsutoshi Kumagai, Tomoharu Iwata, Yasuhiro Fujiwara
NeurIPS2
2019 Autoencoding Binary Classifiers for Supervised Anomaly Detection
Yuki Yamanaka, Tomoharu Iwata, Hiroshi Takahashi, Masanori Yamada, Sekitoshi Kanai
PRICAI (2)2
2018 Few-shot learning of neural networks from scratch by pseudo example optimization
Akisato Kimura, Zoubin Ghahramani, Koh Takeuchi 0001, Tomoharu Iwata, Naonori Ueda
BMVC4
2018 Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific Representations
abstract
Training recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to a target domain, have been proposed for low-resource language modeling. While there are the commonalities and discrepancies between the source and target domains in terms of the statistics of words and their contexts, these methods for domain adaptation make the commonalities and discrepancies jumbled. We propose novel domain adaptation techniques for RNNLM by introducing domain-shared and domain-specific word embedding and contextual features. This explicit modeling of the commonalities and discrepancies would improve the language modeling performance. Experimental comparisons using multiparty conversation data as the target domain augmented by lecture data from the source domain demonstrate that the proposed domain adaptation method exhibits improvements in the perplexity and word error rate over the long short-term memory based language model (LSTMLM) trained using the source and target domain data.
Tsuyoshi Morioka, Naohiro Tawara, Tetsuji Ogawa, Atsunori Ogawa, Tomoharu Iwata, Tetsunori Kobayashi
ICASSP5
2018 Student-t Variational Autoencoder for Robust Density Estimation
abstract
We propose a robust multivariate density estimator based on the variational autoencoder (VAE). The VAE is a powerful deep generative model, and used for multivariate density estimation. With the original VAE, the distribution of observed continuous variables is assumed to be a Gaussian, where its mean and variance are modeled by deep neural networks taking latent variables as their inputs. This distribution is called the decoder. However, the training of VAE often becomes unstable. One reason is that the decoder of VAE is sensitive to the error between the data point and its estimated mean when its estimated variance is almost zero. We solve this instability problem by making the decoder robust to the error using a Bayesian approach to the variance estimation: we set a prior for the variance of the Gaussian decoder, and marginalize it out analytically, which leads to proposing the Student-t VAE. Numerical experiments with various datasets show that training of the Student-t VAE is robust, and the Student-t VAE achieves high density estimation performance.
Hiroshi Takahashi, Tomoharu Iwata, Yuki Yamanaka, Masanori Yamada, Satoshi Yagi
IJCAI2
2018 Estimating Latent People Flow without Tracking Individuals
abstract
Analyzing people flows is important for better navigation and location-based advertising. Since the location information of people is often aggregated for protecting privacy, it is not straightforward to estimate transition populations between locations from aggregated data. Here, aggregated data are incoming and outgoing people counts at each location; they do not contain tracking information of individuals. This paper proposes a probabilistic model for estimating unobserved transition populations between locations from only aggregated data. With the proposed model, temporal dynamics of people flows are assumed to be probabilistic diffusion processes over a network, where nodes are locations and edges are paths between locations. By maximizing the likelihood with flow conservation constraints that incorporate travel duration distributions between locations, our model can robustly estimate transition populations between locations. The statistically significant improvement of our model is demonstrated using real-world datasets of pedestrian data in exhibition halls, bike trip data and taxi trip data in New York City.
Yusuke Tanaka 0002, Tomoharu Iwata, Takeshi Kurashima, Hiroyuki Toda, Naonori Ueda
IJCAI2
2018 Semi-Supervised End-to-End Speech Recognition
Shigeki Karita, Shinji Watanabe 0001, Tomoharu Iwata, Atsunori Ogawa, Marc Delcroix
INTERSPEECH3
2018 Learning Dynamics of Decision Boundaries without Additional Labeled Data
abstract
We propose a method for learning the dynamics of the decision boundary to maintain classification performance without additional labeled data. In various applications, such as spam-mail classification, the decision boundary dynamically changes over time. Accordingly, the performance of classifiers deteriorates quickly unless the classifiers are retrained using additional labeled data. However, continuously preparing such data is quite expensive or impossible. The proposed method alleviates this deterioration in performance by using newly obtained unlabeled data, which are easy to prepare, as well as labeled data collected beforehand. With the proposed method, the dynamics of the decision boundary is modeled by Gaussian processes. To exploit information on the decision boundaries from unlabeled data, the low-density separation criterion, i.e., the decision boundary should not cross high-density regions, but instead lie in low-density regions, is assumed with the proposed method. We incorporate this criterion into our framework in a principled manner by introducing the entropy posterior regularization to the posterior of the classifier parameters on the basis of the generic regularized Bayesian framework. We developed an efficient inference algorithm for the model based on variational Bayesian inference. The effectiveness of the proposed method was demonstrated through experiments using two synthetic and four real-world data sets.
Atsutoshi Kumagai, Tomoharu Iwata
KDD2
2018 On Reducing Dimensionality of Labeled Data Efficiently
Guoxi Zhang, Tomoharu Iwata, Hisashi Kashima
PAKDD (3)2
2018 Improving Route Traffic Estimation by Considering Staying Population
Hitoshi Shimizu, Tatsushi Matsubayashi, Yusuke Tanaka 0002, Tomoharu Iwata, Naonori Ueda, Hiroshi Sawada
PRIMA4
2018 Topic Models for Unsupervised Cluster Matching
abstract
We propose topic models for unsupervised cluster matching, which is the task of finding matching between clusters in different domains without correspondence information. For example, the proposed model finds correspondence between document clusters in English and German without alignment information, such as dictionaries and parallel sentences/documents. The proposed model assumes that documents in all languages have a common latent topic structure, and there are potentially infinite number of topic proportion vectors in a latent topic space that is shared by all languages. Each document is generated using one of the topic proportion vectors and language-specific word distributions. By inferring a topic proportion vector used for each document, we can allocate documents in different languages into common clusters, where each cluster is associated with a topic proportion vector. Documents assigned into the same cluster are considered to be matched. We develop an efficient inference procedure for the proposed model based on collapsed Gibbs sampling. The effectiveness of the proposed model is demonstrated with real data sets including multilingual corpora of Wikipedia and product reviews.
Tomoharu Iwata, Tsutomu Hirao, Naonori Ueda
IEEE Trans. Knowl. Data Eng.1
2017 Read the Silence: Well-Timed Recommendation via Admixture Marked Point Processes
abstract
Everything has its time, which is also true in the point-of-interest (POI) recommendation task. A truly intelligent recommender system, even if you don't visit any sites or remain silent, should draw hints of your next destination from the ``silence", and revise its recommendations as needed. In this paper, we construct a well-timed POI recommender system that updates its recommendations in accordance with the silence, the temporal period in which no visits are made. To achieve this, we propose a novel probabilistic model to predict the joint probabilities of the user visiting POIs and their time-points, by using the admixture or mixed-membership structure to extend marked point processes. With the admixture structure, the proposed model obtains a low dimensional representation for each user, leading to robust recommendation against sparse observations. We also develop an efficient and easy-to-implement estimation algorithm for the proposed model based on collapsed Gibbs and slice sampling. We apply the proposed model to synthetic and real-world check-in data, and show that it performs well in the well-timed recommendation task.
Hideaki Kim, Tomoharu Iwata, Yasuhiro Fujiwara, Naonori Ueda
AAAI2
2017 Learning Non-Linear Dynamics of Decision Boundaries for Maintaining Classification Performance
abstract
We propose a method that involves a probabilistic model for learning future classifiers for tasks in which decision boundaries nonlinearly change over time. In certain applications, such as spam-mail classification, the decision boundary dynamically changes over time. Accordingly, the performance of the classifiers will deteriorate quickly unless the classifiers are updated using additional data. However, collecting such data can be expensive or impossible. The proposed model alleviates this deterioration in performance without additional data by modeling the non-linear dynamics of the decision boundary using Gaussian processes. The method also involves our developed learning algorithm for our model based on empirical variational Bayesian inference by which uncertainty of dynamics can be incorporated for future classification. The effectiveness of the proposed method was demonstrated through experiments using synthetic and real-world data sets.
Atsutoshi Kumagai, Tomoharu Iwata
AAAI2
2017 Localized Lasso for High-Dimensional Regression
abstract
We introduce the localized Lasso, which learns models that both are interpretable and have a high predictive power in problems with high dimensionality d and small sample size n. More specifically, we consider a function defined by local sparse models, one at each data point. We introduce sample-wise network regularization to borrow strength across the models, and sample-wise exclusive group sparsity (a.k.a., l12 norm) to introduce diversity into the choice of feature sets in the local models. The local models are interpretable in terms of similarity of their sparsity patterns. The cost function is convex, and thus has a globally optimal solution. Moreover, we propose a simple yet efficient iterative least-squares based optimization procedure for the localized Lasso, which does not need a tuning parameter, and is guaranteed to converge to a globally optimal solution. The solution is empirically shown to outperform alternatives for both simulated and genomic personalized/precision medicine data.
Makoto Yamada, Koh Takeuchi 0001, Tomoharu Iwata, John Shawe-Taylor, Samuel Kaski
AISTATS3
2017 Latent Dimensionality Estimation for Probabilistic Canonical Correlation Analysis Using Normalized Maximum Likelihood Code-Length
abstract
Discovering hidden common factors from multiple different but related datasets is an important task in data mining. Probabilistic canonical correlation analysis (PCCA) is successfully used for this task, where private factors, which represent independent factors that have influence on a dataset, are modeled as well as common factors. We propose a method for estimating the latent dimensionality of PCCA, which represents the numbers of common and private factors. The dimensionality estimation is indispensable for both generalization ability and interpretability. The proposed method applies the minimum description length criterion using normalized maximum likelihood coding to PCCA in a theoretically justified manner, where PCCA is transformed to a regular model by latent variable completion. We demonstrate that the proposed method surpasses conventional methods in terms of dimensionality estimation performance.
Tomohiko Nakmaura, Tomoharu Iwata, Kenji Yamanishi
DSAA2
2017 SVD-Based Screening for the Graphical Lasso
abstract
The graphical lasso is the most popular approach to estimating the inverse covariance matrix of high-dimension data. It iteratively estimates each row and column of the matrix in a round-robin style until convergence. However, the graphical lasso is infeasible due to its high computation cost for large size of datasets. This paper proposes Sting, a fast approach to the graphical lasso. In order to reduce the computation cost, it efficiently identifies blocks in the estimated matrix that have nonzero elements before entering the iterations by exploiting the singular value decomposition of data matrix. In addition, it selectively updates elements of the estimated matrix expected to have nonzero values. Theoretically, it guarantees to converge to the same result as the original algorithm of the graphical lasso. Experiments show that our approach is faster than existing approaches.
Yasuhiro Fujiwara, Naoki Marumo, Mathieu Blondel, Koh Takeuchi 0001, Hideaki Kim, Tomoharu Iwata, Naonori Ueda
IJCAI6
2017 Learning Latest Classifiers without Additional Labeled Data
abstract
In various applications such as spam mail classification, the performance of classifiers deteriorates over time. Although retraining classifiers using labeled data helps to maintain the performance, continuously preparing labeled data is quite expensive. In this paper, we propose a method to learn classifiers by using newly obtained unlabeled data, which are easy to prepare, as well as labeled data collected beforehand. A major reason for the performance deterioration is the emergence of new features that do not appear in the training phase. Another major reason is the change of the distribution between the training and test phases. The proposed method learns the latest classifiers that overcome both problems. With the proposed method, the conditional distribution of new features given existing features is learned using the unlabeled data. In addition, the proposed method estimates the density ratio between training and test distributions by using the labeled and unlabeled data. We approximate the classification error of a classifier, which exploits new features as well as existing features, at the test phase by incorporating both the conditional distribution of new features and the densityratio, simultaneously. By minimizing the approximated error while integrating out new feature values, we obtain a classifier that exploits new features and fits on the test phase. The effectiveness of the proposed method is demonstrated with experiments using synthetic and real-world data sets.
Atsutoshi Kumagai, Tomoharu Iwata
IJCAI2
2017 Structurally Regularized Non-negative Tensor Factorization for Spatio-Temporal Pattern Discoveries
Koh Takeuchi 0001, Yoshinobu Kawahara, Tomoharu Iwata
ECML/PKDD (1)3
2017 Robust Multi-view Topic Modeling by Incorporating Detecting Anomalies
Guoxi Zhang, Tomoharu Iwata, Hisashi Kashima
ECML/PKDD (2)2
2017 Scaling Locally Linear Embedding
abstract
Locally Linear Embedding (LLE) is a popular approach to dimensionality reduction as it can effectively represent nonlinear structures of high-dimensional data. For dimensionality reduction, it computes a nearest neighbor graph from a given dataset where edge weights are obtained by applying the Lagrange multiplier method, and it then computes eigenvectors of the LLE kernel where the edge weights are used to obtain the kernel. Although LLE is used in many applications, its computation cost is significantly high. This is because, in obtaining edge weights, its computation cost is cubic in the number of edges to each data point. In addition, the computation cost in obtaining the eigenvectors of the LLE kernel is cubic in the number of data points. Our approach, Ripple, is based on two ideas: (1) it incrementally updates the edge weights by exploiting the Woodbury formula and (2) it efficiently computes eigenvectors of the LLE kernel by exploiting the LU decomposition-based inverse power method. Experiments show that Ripple is significantly faster than the original approach of LLE by guaranteeing the same results of dimensionality reduction.
Yasuhiro Fujiwara, Naoki Marumo, Mathieu Blondel, Koh Takeuchi 0001, Hideaki Kim, Tomoharu Iwata, Naonori Ueda
SIGMOD Conference6
2017 Robust unsupervised cluster matching for network data
Tomoharu Iwata, Katsuhiko Ishiguro
Data Min. Knowl. Discov.1
2017 Unsupervised group matching with application to cross-lingual topic matching without alignment information
Tomoharu Iwata, Motonobu Kanagawa, Tsutomu Hirao, Kenji Fukumizu
Data Min. Knowl. Discov.1
2016 Learning Future Classifiers without Additional Data
abstract
We propose probabilistic models for predicting future classifiers given labeled data with timestamps collected until the current time. In some applications, the decision boundary changes over time. For example, in spam mail classification, spammers continuously create new spam mails to overcome spam filters, and therefore, the decision boundary that classifies spam or non-spam can vary. Existing methods require additional labeled and/or unlabeled data to learn a time-evolving decision boundary. However, collecting these data can be expensive or impossible. By incorporating time-series models to capture the dynamics of a decision boundary, the proposed model can predict future classifiers without additional data. We developed two learning algorithms for the proposed model on the basis of variational Bayesian inference. The effectiveness of the proposed method is demonstrated with experiments using synthetic and real-world data sets.
Atsutoshi Kumagai, Tomoharu Iwata
AAAI2
2016 Identifying Key Observers to Find Popular Information in Advance
Takuya Konishi, Tomoharu Iwata, Kohei Hayashi, Ken-ichi Kawarabayashi
IJCAI2
2016 Multi-view Anomaly Detection via Robust Probabilistic Latent Variable Models
abstract
We propose probabilistic latent variable models for multi-view anomaly detection, which is the task of finding instances that have inconsistent views given multi-view data. With the proposed model, all views of a non-anomalous instance are assumed to be generated from a single latent vector. On the other hand, an anomalous instance is assumed to have multiple latent vectors, and its different views are generated from different latent vectors. By inferring the number of latent vectors used for each instance with Dirichlet process priors, we obtain multi-view anomaly scores. The proposed model can be seen as a robust extension of probabilistic canonical correlation analysis for noisy multi-view data. We present Bayesian inference procedures for the proposed model based on a stochastic EM algorithm. The effectiveness of the proposed model is demonstrated in terms of performance when detecting multi-view anomalies.
Tomoharu Iwata, Makoto Yamada
NIPS1
2016 Inferring Latent Triggers of Purchases with Consideration of Social Effects and Media Advertisements
abstract
This paper proposes a method for inferring from single-source data the factors that trigger purchases. Here, single-source data are the histories of item purchases and media advertisement views for each individual. We assume a sequence of purchase events to be a stochastic process incorporating the following three factors: (a) user preference, (b) social effects received from other users, and (c) media advertising effects. As our user-purchase model incorporates the latent relationships between users and advertisers, it can infer the latent triggers of purchases. Experiments on real single-source data show that our model can (a) achieve high prediction accuracy for purchases, (b) discover the key information, i.e., popular items, influential users, and influential advertisers, (c) estimate the relative impact of the three factors on purchases, and (d) find user segments according to the estimated factors.
Yusuke Tanaka 0002, Takeshi Kurashima, Yasuhiro Fujiwara, Tomoharu Iwata, Hiroshi Sawada
WSDM4
2016 Probabilistic latent variable models for unsupervised many-to-many object matching
Tomoharu Iwata, Tsutomu Hirao, Naonori Ueda
Inf. Process. Manag.1
2016 Unsupervised Many-to-Many Object Matching for Relational Data
abstract
We propose a method for unsupervised many-to-many object matching from multiple networks, which is the task of finding correspondences between groups of nodes in different networks. For example, the proposed method can discover shared word groups from multi-lingual document-word networks without cross-language alignment information. We assume that multiple networks share groups, and each group has its own interaction pattern with other groups. Using infinite relational models with this assumption, objects in different networks are clustered into common groups depending on their interaction patterns, discovering a matching. The effectiveness of the proposed method is experimentally demonstrated by using synthetic and real relational data sets, which include applications to cross-domain recommendation without shared user/item identifiers and multi-lingual word clustering.
Tomoharu Iwata, James Robert Lloyd, Zoubin Ghahramani
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Non-Linear Regression for Bag-of-Words Data via Gaussian Process Latent Variable Set Model
abstract
Gaussian process (GP) regression is a widely used method for non-linear prediction.The performance of the GP regression depends on whether it can properly capture the covariance structure of target variables, which is represented by kernels between input data.However, when the input is represented as a set of features, e.g. bag-of-words, it is difficult to calculate desirable kernel values because the co-occurrence of different but relevant words cannot be reflected in the kernel calculation.To overcome this problem, we propose a Gaussian process latent variable set model (GP-LVSM), which is a non-linear regression model effective for bag-of-words data.With the GP-LVSM, a latent vector is associated with each word, and each document is represented as a distribution of the latent vectors for words appearing in the document. We efficiently represent the distributions by using the framework of kernel embeddings of distributions that can hold high-order moment information of distributions without need for explicit density estimation.By learning latent vectors so as to maximize the posterior probability, kernels that reflect relations between words are obtained, and also words are visualized in a low-dimensional space.In experiments using 25 item review datasets, we demonstrate the effectiveness of the GP-LVSM in prediction and visualization.
Yuya Yoshikawa, Tomoharu Iwata, Hiroshi Sawada
AAAI2
2015 Cross-domain recommendation without shared users or items by sharing latent vector distributions
abstract
We propose a cross-domain recommendation method for predicting the ratings of items in different domains, where neither users nor items are shared across domains. The proposed method is based on matrix factorization, which learns a latent vector for each user and each item. Matrix factorization techniques for a single-domain fail in the cross-domain recommendation task because the learned latent vectors are not aligned over different domains. The proposed method assumes that latent vectors in different domains are generated from a common Gaussian distribution with a full covariance matrix. By inferring the mean and covariance of the common Gaussian from given cross-domain rating matrices, the latent factors are aligned, which enables us to predict ratings in different domains. Experiments conducted on rating datasets from a wide variety of domains, e.g., movie, books and electronics, demonstrate that the proposed method achieves higher performance for predicting cross-domain ratings than existing methods.
Tomoharu Iwata, Koh Takeuchi 0001
AISTATS1
2015 Multiscale recurrent neural network based language model
Tsuyoshi Morioka, Tomoharu Iwata, Takaaki Hori, Tetsunori Kobayashi
INTERSPEECH2
2015 Cross-Domain Matching for Bag-of-Words Data via Kernel Embeddings of Latent Distributions
abstract
We propose a kernel-based method for finding matching between instances across different domains, such as multilingual documents and images with annotations. Each instance is assumed to be represented as a multiset of features, e.g., a bag-of-words representation for documents. The major difficulty in finding cross-domain relationships is that the similarity between instances in different domains cannot be directly measured. To overcome this difficulty, the proposed method embeds all the features of different domains in a shared latent space, and regards each instance as a distribution of its own features in the shared latent space. To represent the distributions efficiently and nonparametrically, we employ the framework of the kernel embeddings of distributions. The embedding is estimated so as to minimize the difference between distributions of paired instances while keeping unpaired instances apart. In our experiments, we show that the proposed method can achieve high performance on finding correspondence between multi-lingual Wikipedia articles, between documents and tags, and between images and tags.
Yuya Yoshikawa, Tomoharu Iwata, Hiroshi Sawada, Takeshi Yamada
NIPS2
2015 Higher Order Fused Regularization for Supervised Learning with Grouped Parameters
Koh Takeuchi 0001, Yoshinobu Kawahara, Tomoharu Iwata
ECML/PKDD (1)3
2014 Probabilistic latent network visualization: inferring and embedding diffusion networks
abstract
The diffusion of information, rumors, and diseases are assumed to be probabilistic processes over some network structure. An event starts at one node of the network, and then spreads to the edges of the network. In most cases, the underlying network structure that generates the diffusion process is unobserved, and we only observe the times at which each node is altered/influenced by the process. This paper proposes a probabilistic model for inferring the diffusion network, which we call Probabilistic Latent Network Visualization (PLNV); it is based on cascade data, a record of observed times of node influence. An important characteristic of our approach is to infer the network by embedding it into a low-dimensional visualization space. We assume that each node in the network has latent coordinates in the visualization space, and diffusion is more likely to occur between nodes that are placed close together. Our model uses maximum a posteriori estimation to learn the latent coordinates of nodes that best explain the observed cascade data. The latent coordinates of nodes in the visualization space can 1) enable the system to suggest network layouts most suitable for browsing, and 2) lead to high accuracy in inferring the underlying network when analyzing the diffusion process of new or rare information, rumors, and disease.
Takeshi Kurashima, Tomoharu Iwata, Noriko Takaya, Hiroshi Sawada
KDD2
2014 Latent Support Measure Machines for Bag-of-Words Data Classification
Yuya Yoshikawa, Tomoharu Iwata, Hiroshi Sawada
NIPS2
2014 Generating structure of latent variable models for nested data
Masakazu Ishihata, Tomoharu Iwata
UAI2
2013 Unsupervised Cluster Matching via Probabilistic Latent Variable Models
abstract
We propose a probabilistic latent variable model for unsupervised cluster matching, which is the task of finding correspondences between clusters of objects in different domains. Existing object matching methods find one-to-one matching. The proposed model finds many-to-many matching, and can handle multiple domains with different numbers of objects. The proposed model assumes that there are an infinite number of latent vectors that are shared by all domains, and that each object is generated using one of the latent vectors and a domain-specific linear projection. By inferring a latent vector to be used for generating each object, objects in different domains are clustered in shared groups, and thus we can find matching between clusters in an unsupervised manner. We present efficient inference procedures for the proposed model based on a stochastic EM algorithm. The effectiveness of the proposed model is demonstrated with experiments using synthetic and real data sets.
Tomoharu Iwata, Tsutomu Hirao, Naonori Ueda
AAAI1
2013 Active Learning for Interactive Visualization
abstract
Many automatic visualization methods have been proposed. However, a visualization that is automatically generated might be different to how a user wants to arrange the objects in visualization space. By allowing users to re-locate objects in the embedding space of the visualization, they can adjust the visualization to their preference. We propose an active learning framework for interactive visualization which selects objects for the user to re-locate so that they can obtain their desired visualization by re-locating as few as possible. The framework is based on an information theoretic criterion, which favors objects that reduce the uncertainty of the visualization. We present a concrete application of the proposed framework to the Laplacian eigenmap visualization method. We demonstrate experimentally that the proposed framework yields the desired visualization with fewer user interactions than existing methods.
Tomoharu Iwata, Neil Houlsby, Zoubin Ghahramani
AISTATS1
2013 A Probabilistic Model for Diversifying Recommendation Lists
Yutaka Kabutoya, Tomoharu Iwata, Hiroyuki Toda, Hiroyuki Kitagawa
APWeb2
2013 Clustering-based anomaly detection in multi-view data
abstract
This paper proposes a simple yet effective anomaly detection method for multi-view data. The proposed approach detects anomalies by comparing the neighborhoods in different views. Specifically, clustering is performed separately in the different views and affinity vectors are derived for each object from the clustering results. Then, the anomalies are detected by comparing affinity vectors in the multiple views. An advantage of the proposed method over existing methods is that the tuning parameters can be determined effectively from the given data. Through experiments on synthetic and benchmark datasets, we show that the proposed method outperforms existing methods.
Alejandro Marcos Alvarez, Makoto Yamada, Akisato Kimura, Tomoharu Iwata
CIKM4
2013 A Probabilistic Behavior Model for Discovering Unrecognized Knowledge
abstract
Discovering interesting behavior patterns and profiles of users as they interact with E-commerce (EC) sites is an important task for site managers. We propose a probabilistic behavior model for extracting latent classes of items that impact the users' item selections but cannot be inferred from the current knowledge of the managers. The proposed model assumes that the current knowledge is represented by categories of items that are defined in the EC site, and a user selects items depending on both of their categories and latent classes. By estimating latent classes, each of which shows items accessed by users with common interests, we can find interesting factors for explaining user behavior. We evaluate our proposed model using item-access log data observed in an EC site. The results show that our model can accurately predict users' item selection, and actually discover latent classes of items having similar latent characteristic such as "colored design" and "impression" by using item categories such as "coat" and "hat" as the current knowledge of the managers.
Takeshi Kurashima, Tomoharu Iwata, Noriko Takaya, Hiroshi Sawada
ICDM2
2013 Discovering latent influence in online social activities via shared cascade poisson processes
abstract
Many people share their activities with others through online communities. These shared activities have an impact on other users' activities. For example, users are likely to become interested in items that are adopted (e.g. liked, bought and shared) by their friends. In this paper, we propose a probabilistic model for discovering latent influence from sequences of item adoption events. An inhomogeneous Poisson process is used for modeling a sequence, in which adoption by a user triggers the subsequent adoption of the same item by other users. For modeling adoption of multiple items, we employ multiple inhomogeneous Poisson processes, which share parameters, such as influence for each user and relations between users. The proposed model can be used for finding influential users, discovering relations between users and predicting item popularity in the future. We present an efficient Bayesian inference procedure of the proposed model based on the stochastic EM algorithm. The effectiveness of the proposed model is demonstrated by using real data sets in a social bookmark sharing service.
Tomoharu Iwata, Amar Shah 0001, Zoubin Ghahramani
KDD1
2013 Warped Mixtures for Nonparametric Cluster Shapes
Tomoharu Iwata, David Duvenaud, Zoubin Ghahramani
UAI1
2013 Geo topic model: joint modeling of user's activity area and interests for location recommendation
abstract
This paper proposes a method that analyzes the location log data of multiple users to recommend locations to be visited. The method uses our new topic model, called Geo Topic Model, that can jointly estimate both the user's interests and activity area hosting the user's home, office and other personal places. By explicitly modeling geographical features of locations and users, the user's interests in other features of locations, which we call latent topics, can be inferred effectively. The topic interests estimated by our model 1) lead to high accuracy in predicting visit behavior as driven by personal interests, 2) make possible the generation of recommendations when the user is in an unfamiliar area (e.g. sightseeing), and 3) enable the recommender system to suggest an interpretable representation of the user profile that can be customized by the user. Experiments are conducted using real location logs of landmark and restaurant visits to evaluate the recommendation performance of the proposed method in terms of the accuracy of predicting visit selections. We also show that our model can estimate latent features of locations such as art, nature and atmosphere as latent topics, and describe each user's preference based on them.
Takeshi Kurashima, Tomoharu Iwata, Takahide Hoshide, Noriko Takaya, Ko Fujimura
WSDM2
2013 Topic model for analyzing purchase data with price information
Tomoharu Iwata, Hiroshi Sawada
Data Min. Knowl. Discov.1
2013 Travel route recommendation using geotagged photos
Takeshi Kurashima, Tomoharu Iwata, Go Irie, Ko Fujimura
Knowl. Inf. Syst.2
2013 Influence relation estimation based on lexical entrainment in conversation
Tomoharu Iwata, Shinji Watanabe 0001
Speech Commun.1
2013 Modeling Noisy Annotated Data with Application to Social Annotation
abstract
We propose a probabilistic topic model for analyzing and extracting content-related annotations from noisy annotated discrete data such as webpages stored using social bookmarking services. With these services, because users can attach annotations freely, some annotations do not describe the semantics of the content, thus they are noisy, i.e., not content related. The extraction of content-related annotations can be used as a prepossessing step in machine learning tasks such as text classification and image recognition, or can improve information retrieval performance. The proposed model is a generative model for content and annotations, in which the annotations are assumed to originate either from topics that generated the content or from a general distribution unrelated to the content. We demonstrate the effectiveness of the proposed method by using synthetic data and real social annotation data for text and images.
Tomoharu Iwata, Takeshi Yamada, Naonori Ueda
IEEE Trans. Knowl. Data Eng.1
2012 Handling uncertain observations in unsupervised topic-mixture language model adaptation
abstract
We propose an extension to the recent approaches in topic-mixture modeling such as Latent Dirichlet Allocation and Topic Tracking Model for the purpose of unsupervised adaptation in speech recognition. Instead of using the 1-best input given by the speech recognizer, the proposed model takes confusion network as an input to alleviate recognition errors. We incorporate a selection variable which helps reweight the recognition output, thus creating a more accurate latent topic estimate. Compared to adapting based on just one recognition hypothesis, the proposed model show WER improvements on two different tasks.
Ekapol Chuangsuwanich, Shinji Watanabe 0001, Takaaki Hori, Tomoharu Iwata, James R. Glass
ICASSP4
2012 Effect of dialog acts on word use in polylogue
abstract
In this work we examine the effect of dialog acts on word use, in context of the influence of interlocutors in a polylogue on each other. The basic idea of this work is the extension of the cache model and the influence model by dialog act information. The cache model covers the re-usage of words and the influence model calculates the influence of interlocutors in a polylogue on each other. Both approaches could be used to improve the word prediction accuracy in a word generative model. We start to examine the usage of dialog acts to improve our word generative model in terms of perplexity. For the usage of dialog acts, a knowledge about the future dialog act is required. Therefore, we examine how dialog act miss-prediction influences the resulting performance. Further on, we introduce a new approach to generate artificial dialog acts which guarantees the knowledge about the following dialog act. Our final experiments present the improvements in terms of perplexity using our new approach in AMI, NIST and NTT meeting corpora.
Roland Roller, Shinji Watanabe 0001, Tomoharu Iwata
ICASSP3
2012 Creating Stories: Social Curation of Twitter Messages
Kevin Duh, Tsutomu Hirao, Akisato Kimura, Katsuhiko Ishiguro, Tomoharu Iwata, Ching-man Au Yeung
ICWSM5
2012 Fast mining and forecasting of complex time-stamped events
abstract
Given huge collections of time-evolving events such as web-click logs, which consist of multiple attributes (e.g., URL, userID, times- tamp), how do we find patterns and trends? How do we go about capturing daily patterns and forecasting future events? We need two properties: (a) effectiveness, that is, the patterns should help us understand the data, discover groups, and enable forecasting, and (b) scalability, that is, the method should be linear with the data size. We introduce TriMine, which performs three-way mining for all three attributes, namely, URLs, users, and time. Specifically TriMine discovers hidden topics, groups of URLs, and groups of users, simultaneously. Thanks to its concise but effective summarization, it makes it possible to accomplish the most challenging and important task, namely, to forecast future events. Extensive experiments on real datasets demonstrate that TriMine discovers meaningful topics and makes long-range forecasts, which are notoriously difficult to achieve. In fact, TriMine consistently outperforms the best state-of-the-art existing methods in terms of accuracy and execution speed (up to 74x faster).
Yasuko Matsubara, Yasushi Sakurai, Christos Faloutsos, Tomoharu Iwata, Masatoshi Yoshikawa
KDD4
2012 Bidirectional Semi-supervised Learning with Graphs
Tomoharu Iwata, Kevin Duh
ECML/PKDD (2)1
2012 A Topic Model for Recommending Movies via Linked Open Data
abstract
We propose an algorithm for recommending both well-watched old movies and unwatched new ones. To recommend both old favourites and new releases, hybrids of collaborative and content-based filtering are the most suitable methods. However, hybrid movie recommenders have two issues. First, it is necessary to acquire content-descriptive metadata, which is not always easily available. Second, the metadata, once acquired, may be noisy, which can damage recommendation accuracy. In our algorithm, we address the first issue by automatically drawing movie metadata from Linked Open Data, and the second by modeling the relevance of the collected metadata to the transaction history before using the relationship between them to make recommendations. We experimentally demonstrate that our method can effectively collect metadata from LOD, and that our method outperforms conventional hybrid methods found in the literature in both well-watched and unwatched movie recommendation using the noisy collected movie metadata.
Yutaka Kabutoya, Róbert Sumi, Tomoharu Iwata, Toshio Uchiyama, Tadasu Uchiyama
Web Intelligence3
2012 Sequential Modeling of Topic Dynamics with Multiple Timescales
abstract
We propose an online topic model for sequentially analyzing the time evolution of topics in document collections. Topics naturally evolve with multiple timescales. For example, some words may be used consistently over one hundred years, while other words emerge and disappear over periods of a few days. Thus, in the proposed model, current topic-specific distributions over words are assumed to be generated based on the multiscale word distributions of the previous epoch. Considering both the long- and short-timescale dependency yields a more robust model. We derive efficient online inference procedures based on a stochastic EM algorithm, in which the model is sequentially updated using newly obtained data; this means that past data are not required to make the inference. We demonstrate the effectiveness of the proposed method in terms of predictive performance and computational efficiency by examining collections of real documents with timestamps.
Tomoharu Iwata, Takeshi Yamada, Yasushi Sakurai, Naonori Ueda
ACM Trans. Knowl. Discov. Data1
2011 Transfer Learning for Multiple-Domain Sentiment Analysis - Identifying Domain Dependent/Independent Word Polarity
abstract
Sentiment analysis is the task of determining the attitude (positive or negative) of documents. While the polarity of words in the documents is informative for this task, polarity of some words cannot be determined without domain knowledge. Detecting word polarity thus poses a challenge for multiple-domain sentiment analysis. Previous approaches tackle this problem with transfer learning techniques, but they cannot handle multiple source domains and multiple target domains. This paper proposes a novel Bayesian probabilistic model to handle multiple source and multiple target domains. In this model, each word is associated with three factors: Domain label, domain dependence/independence and word polarity. We derive an efficient algorithm using Gibbs sampling for inferring the parameters of the model, from both labeled and unlabeled texts. Using real data, we demonstrate the effectiveness of our model in a document polarity classification task compared with a method not considering the differences between domains. Moreover our method can also tell whether each word's polarity is domain-dependent or domain-independent. This feature allows us to construct a word polarity dictionary for each domain.
Yasuhisa Yoshida, Tsutomu Hirao, Tomoharu Iwata, Masaaki Nagata, Yuji Matsumoto 0001
AAAI3
2011 Extracting multi-dimensional relations: a generative model of groups of entities in a corpus
abstract
Extracting relations among different entities from various data sources has been an important topic in data mining. While many methods focus only on a single type of relations, real world entities maintain relations that contain much richer information. We propose a hierarchical Bayesian model for extracting multi-dimensional relations among entities from a text corpus. Using data from Wikipedia, we show that our model can accurately predict the relevance of an entity given the topic of the document as well as the set of entities that are already mentioned in that document.
Ching-man Au Yeung, Tomoharu Iwata
CIKM2
2011 Fashion Coordinates Recommender System Using Photographs from Fashion Magazines
abstract
Fashion magazines contain a number of photographs of fashion models, and their clothing coordinates serve as useful references. In this paper, we propose a recommender system for clothing coordinates using full-body photographs from fashion magazines. The task is that, given a photograph of a fashion item (e.g. tops) as a query, to recommend a photograph of other fashion items (e.g. bottoms) that is appropriate to the query. With the proposed method, we use a probabilistic topic model for learning information about coordinates from visual features in each fashion item region. We demonstrate the effectiveness of the proposed method using real photographs from a fashion magazine and two fashion style sharing services with the task of making top (bottom) recommendations given bottom (top) photographs. 1
Tomoharu Iwata, Shinji Watanabe 0001, Hiroshi Sawada
IJCAI1
2011 Learning Influences from Word Use in Polylogue
abstract
We propose a probabilistic model for estimating influences among speakers from conversation data with multiple people. In conversations, people tend to mimic their companions ’ be-havior depending on their level of trust. With the proposed model, we assume that the word use of a speaker depends on the word use of previous speakers as well as their own earlier word use and the general word distribution. The influences can be ef-ficiently estimated by using the expectation maximization (EM) algorithm. Experiments on two meeting data sets in Japanese and in English demonstrate the effectiveness of the proposed method. Index Terms: conversation analysis, influence, latent variable model
Tomoharu Iwata, Shinji Watanabe 0001
INTERSPEECH1
2011 Alignment Inference and Bayesian Adaptation for Machine Translation
Kevin Duh, Katsuhito Sudoh, Tomoharu Iwata, Hajime Tsukada
MTSummit3
2011 Strength of social influence in trust networks in product review sites
abstract
Some popular product review sites such as Epinions allow users to establish a trust network among themselves, indicating who they trust in providing product reviews and ratings. While trust relations have been found to be useful in generating personalised recommendations, the relations between trust and product ratings has so far been overlooked. In this paper, we examine large datasets collected from Epinions and Ciao, two popular product review sites. We discover that in general users who trust each other tend to have smaller differences in their ratings as time passes, giving support to the theories of homophily and social influence. However, we also discover that this does not hold true across all trusted users. A trust relation does not guarantee that two users have similar preferences, implying that personalised recommendations based on trust relations do not necessarily produce more accurate predictions. We propose a method to estimate the strengths of trust relations so as to estimate the true influence among the trusted users. Our method extends the popular matrix factorisation technique for collaborative filtering, which allow us to generate more accurate rating predictions at the same time. We also show that the estimated strengths of trust relations correlate with the similarity among the users. Our work contributes to the understanding of the interplay between trust relations and product ratings, and suggests that trust networks may serve as a more general socialising venue than only an indication of similarity in user preferences.
Ching-man Au Yeung, Tomoharu Iwata
WSDM2
2011 Topic tracking language model for speech recognition
Shinji Watanabe 0001, Tomoharu Iwata, Takaaki Hori, Atsushi Sako, Yasuo Ariki
Comput. Speech Lang.2
2011 Improving Classifier Performance Using Data with Different Taxonomies
abstract
We propose a framework for improving classifier performance by effectively using auxiliary samples. The auxiliary samples are labeled not in terms of the target taxonomy according to which we wish to classify samples, but according to classification schemes or taxonomies that are different from the target taxonomy. Our method finds a classifier by minimizing a weighted error over the target and auxiliary samples. The weights are defined so that the weighted error approximates the expected error when samples are classified into the target taxonomy. Experiments using synthetic and text data show that our method significantly improves the classifier performance in most cases compared to conventional data augmentation methods.
Tomoharu Iwata, Toshiyuki Tanaka 0003, Takeshi Yamada, Naonori Ueda
IEEE Trans. Knowl. Data Eng.1
2010 Travel route recommendation using geotags in photo sharing sites
abstract
The ability to create geotagged photos enables people to share their personal experiences as tourists at specific locations and times. Assuming that the collection of each photographer's geotagged photos is a sequence of visited locations, photo-sharing sites are important sources for gathering the location histories of tourists. By following their location sequences, we can find representative and diverse travel routes that link key landmarks. In this paper, we propose a travel route recommendation method that makes use of the photographers' histories as held by Flickr. Recommendations are performed by our photographer behavior model, which estimates the probability of a photographer visiting a landmark. We incorporate user preference and present location information into the probabilistic behavior model by combining topic models and Markov models. We demonstrate the effectiveness of the proposed method using a real-life dataset holding information from 71,718 photographers taken in the United States in terms of the prediction accuracy of travel behavior.
Takeshi Kurashima, Tomoharu Iwata, Go Irie, Ko Fujimura
CIKM2
2010 Statistical Learning-based Approach for Automatic Generation System of Multiple-choice Cloze Questions
Tomoko Kojiri, Takuya Goto, Toyohide Watanabe, Tomoharu Iwata, Takeshi Yamada
ICCE4
2010 Effective Question Recommendation Based on Multiple Features for Question Answering Communities
Yutaka Kabutoya, Tomoharu Iwata, Hisako Shiohara, Ko Fujimura
ICWSM2
2010 Online multiscale dynamic topic models
abstract
We propose an online topic model for sequentially analyzing the time evolution of topics in document collections. Topics naturally evolve with multiple timescales. For example, some words may be used consistently over one hundred years, while other words emerge and disappear over periods of a few days. Thus, in the proposed model, current topic-specific distributions over words are assumed to be generated based on the multiscale word distributions of the previous epoch. Considering both the long-timescale dependency as well as the short-timescale dependency yields a more robust model. We derive efficient online inference procedures based on a stochastic EM algorithm, in which the model is sequentially updated using newly obtained data; this means that past data are not required to make the inference. We demonstrate the effectiveness of the proposed method in terms of predictive performance and computational efficiency by examining collections of real documents with timestamps.
Tomoharu Iwata, Takeshi Yamada, Yasushi Sakurai, Naonori Ueda
KDD1
2010 Dynamic Infinite Relational Model for Time-varying Relational Data Analysis
abstract
We propose a new probabilistic model for analyzing dynamic evolutions of relational data, such as additions, deletions and split & merge, of relation clusters like communities in social networks. Our proposed model abstracts observed time-varying object-object relationships into relationships between object clusters. We extend the infinite Hidden Markov model to follow dynamic and time-sensitive changes in the structure of the relational data and to estimate a number of clusters simultaneously. We show the usefulness of the model through experiments with synthetic and real-world data sets.
Katsuhiko Ishiguro, Tomoharu Iwata, Naonori Ueda, Josh Tenenbaum
NIPS2
2010 Application of topic tracking model to language model adaptation and meeting analysis
abstract
In a real environment, acoustic and language features often vary depending on the speakers, speaking styles and topic changes. This paper focuses on changes in the language environment, and applies a topic tracking model to language model adaptation for speech recognition and topic word extraction for meeting analysis. The topic tracking model can adaptively track changes in topics based on current text information and previously estimated topic models in an online manner. The effectiveness of the proposed method is shown experimentally by the improvement in speech recognition performance achieved with the Corpus of Spontaneous Japanese and by providing appropriate topic information in an automatic meeting analyzer.
Shinji Watanabe 0001, Tomoharu Iwata, Takaaki Hori, Atsushi Sako, Yasuo Ariki
SLT2
2010 Modeling Multiple Users' Purchase over a Single Account for Collaborative Filtering
Yutaka Kabutoya, Tomoharu Iwata, Ko Fujimura
WISE2
2009 Topic Tracking Model for Analyzing Consumer Purchase Behavior
Tomoharu Iwata, Shinji Watanabe 0001, Takeshi Yamada, Naonori Ueda
IJCAI1
2009 Modeling Social Annotation Data with Content Relevance using a Topic Model
abstract
We propose a probabilistic topic model for analyzing and extracting content-related annotations from noisy annotated discrete data such as web pages stored in social bookmarking services. In these services, since users can attach annotations freely, some annotations do not describe the semantics of the content, thus they are noisy, i.e. not content-related. The extraction of content-related annotations can be used as a preprocessing step in machine learning tasks such as text classification and image recognition, or can improve information retrieval performance. The proposed model is a generative model for content and annotations, in which the annotations are assumed to originate either from topics that generated the content or from a general distribution unrelated to the content. We demonstrate the effectiveness of the proposed method by using synthetic data and real social annotation data for text and images.
Tomoharu Iwata, Takeshi Yamada, Naonori Ueda
NIPS1
2008 Probabilistic latent semantic visualization: topic model for visualizing documents
abstract
We propose a visualization method based on a topic model for discrete data such as documents. Unlike conventional visualization methods based on pairwise distances such as multi-dimensional scaling, we consider a mapping from the visualization space into the space of documents as a generative process of documents. In the model, both documents and topics are assumed to have latent coordinates in a two- or three-dimensional Euclidean space, or visualization space. The topic proportions of a document are determined by the distances between the document and the topics in the visualization space, and each word is drawn from one of the topics according to its topic proportions. A visualization, i.e. latent coordinates of documents, can be obtained by fitting the model to a given set of documents using the EM algorithm, resulting in documents with similar topics being embedded close together. We demonstrate the effectiveness of the proposed model by visualizing document and movie data sets, and quantitatively compare it with conventional visualization methods.
Tomoharu Iwata, Takeshi Yamada, Naonori Ueda
KDD1
2008 English Grammar Learning System Based on Knowledge Network of Fill-in-the-Blank Exercises
Takuya Goto, Tomoko Kojiri, Toyohide Watanabe, Takeshi Yamada, Tomoharu Iwata
KES (3)5
2008 Recommendation Algorithm for Learning Materials That Maximizes Expected Test Scores
Tomoharu Iwata, Tomoko Kojiri, Takeshi Yamada, Toyohide Watanabe
PRICAI1
2008 Recommendation Method for Improving Customer Lifetime Value
abstract
It is important for online stores to improve customer lifetime value (LTV) if they are to increase their profits. Conventional recommendation methods suggest items that best coincide with user's interests to maximize the purchase probability, and this does not necessarily help improve LTV. We present a novel recommendation method that maximizes the probability of the LTV being improved, which can apply to both measured and subscription services. Our method finds frequent purchase patterns among high-LTV users and recommends items for a new user that simulate the found patterns. Using survival analysis techniques, we efficiently find the patterns from log data. Furthermore, we infer a user's interests from the purchase history based on maximum entropy models and use the interests to improve recommendation. Since a higher LTV is the result of greater user satisfaction, our method benefits users as well as online stores. We evaluate our method using two sets of real log data for measured and subscription services.
Tomoharu Iwata, Kazumi Saito, Takeshi Yamada
IEEE Trans. Knowl. Data Eng.1
2007 Modeling user behavior in recommender systems based on maximum entropy
abstract
We propose a model for user purchase behavior in online stores that provide recommendation services. We model the purchase probability given recommendations for each user based on the maximum entropy principle using features that deal with recommendations and user interests. The proposed model enable us to measure the effect of recommendations on user purchase behavior, and the effect can be used to evaluate recommender systems. We show the validity of our model using the log data of an online cartoon distribution service, and measure the recommendation effects for evaluating the recommender system.
Tomoharu Iwata, Kazumi Saito, Takeshi Yamada
WWW1
2007 Parametric Embedding for Class Visualization
abstract
We propose a new method, parametric embedding (PE), that embeds objects with the class structure into a low-dimensional visualization space. PE takes as input a set of class conditional probabilities for given data points and tries to preserve the structure in an embedding space by minimizing a sum of Kullback-Leibler divergences, under the assumption that samples are generated by a gaussian mixture with equal covariances in the embedding space. PE has many potential uses depending on the source of the input data, providing insight into the classifier's behavior in supervised, semisupervised, and unsupervised settings. The PE algorithm has a computational advantage over conventional embedding methods based on pairwise object relations since its complexity scales with the product of the number of objects and the number of classes. We demonstrate PE by visualizing supervised categorization of Web pages, semisupervised categorization of digits, and the relations of words and latent topics found by an unsupervised algorithm, latent Dirichlet allocation.
Tomoharu Iwata, Kazumi Saito, Naonori Ueda, Sean Stromsten, Thomas L. Griffiths 0001, Josh Tenenbaum
Neural Comput.1
2006 Visual nonlinear discriminant analysis for classifier design
Tomoharu Iwata, Kazumi Saito, Naonori Ueda
ESANN1
2006 Recommendation method for extending subscription periods
abstract
Online stores providing subscription services need to extend user subscription periods as long as possible to increase their profits. Conventional recommendation methods recommend items that best coincide with user's interests to maximize the purchase probability, which does not necessarily contribute to extend subscription periods. We present a novel recommendation method for subscription services that maximizes the probability of the subscription period being extended. Our method finds frequent purchase patterns in the long subscription period users, and recommends items for a new user to simulate the found patterns. Using survival analysis techniques, we efficiently extract information from the log data for finding the patterns. Furthermore, we infer user's interests from purchase histories based on maximum entropy models, and use the interests to improve the recommendations. Since a longer subscription period is the result of greater user satisfaction, our method benefits users as well as online stores. We evaluate our method using the real log data of an online cartoon distribution service for cell-phone in Japan.
Tomoharu Iwata, Kazumi Saito, Takeshi Yamada
KDD1
2004 Visualisation of Anomaly Using Mixture Model
Tomoharu Iwata, Kazumi Saito
KES1
2004 Parametric Embedding for Class Visualization
abstract
In this paper, we propose a new method, Parametric Embedding (PE), for visualizing the posteriors estimated over a mixture model. PE simultane- ously embeds both objects and their classes in a low-dimensional space. PE takes as input a set of class posterior vectors for given data points, and tries to preserve the posterior structure in an embedding space by minimizing a sum of Kullback-Leibler divergences, under the assump- tion that samples are generated by a Gaussian mixture with equal covari- ances in the embedding space. PE has many potential uses depending on the source of the input data, providing insight into the classifier’s be- havior in supervised, semi-supervised and unsupervised settings. The PE algorithm has a computational advantage over conventional embedding methods based on pairwise object relations since its complexity scales with the product of the number of objects and the number of classes. We demonstrate PE by visualizing supervised categorization of web pages, semi-supervised categorization of digits, and the relations of words and latent topics found by an unsupervised algorithm, Latent Dirichlet Allo- cation.
Tomoharu Iwata, Kazumi Saito, Naonori Ueda, Sean Stromsten, Thomas L. Griffiths 0001, Josh Tenenbaum
NIPS1